# Frozen data and evaluation protocol

## Provenance and partitions
Use only supplied `data/raw/train.csv`, `test.csv`, `sample_submission.csv`; no further downloads. Raw data and archive remain outside worker-exposed paths. `prepare.py` validates exact columns, 699,635 labeled rows, 299,844 competition-test rows, distinct unique IDs, aligned sample format, and observed numeric target values {0,1}. Counts: 389,296 zeros, 310,339 ones. Positive probability means the observed label **1**; no inverted target mapping.

With scikit-learn 1.6.1 and seed 20261004, first split 70% training vs 30% holdout, then stratify holdout equally into validation and final. Preserve generated order; do not sort or regenerate. Training-pool size is 489,744; search-validation 104,945; final 104,946. From the pool select stratified 50,000 quick-train rows and then 10,000 disjoint quick-dev rows from its remainder, using the same seed. Quick rows are never from either official holdout.

`starter/data/` contains only labeled pool/quick rows and competition test/sample format. `private/` contains canonical training copy, separate validation/final label-free features and ID/label CSVs. Original sources remain in `data/raw/`. `private/manifest.json` records source/split file SHA-256, counts, mapping, ordered-ID hashes, and preparation versions. `eval/data_manifest.json` exposes only hashes/counts/schema, no held-out row IDs or individual labels. No private/raw link_dirs are configured. Preparation reruns verify hashes instead of silently rewriting frozen splits.

## Leakage rules
Workers must not read raw/private data, private predictions/logs or labels; must not alter evaluator, data, split manifests or config. Same-user workers could access them: the boundary is cooperative, not a permissions sandbox. The evaluator supplies the candidate only labeled training and label-free prediction CSV paths. It never supplies official/final targets. ID is an output key only, not a model feature. Competition-test rows may not be added to training or used to fit development preprocessing. All transformations/target encodings fit within the supplied training rows (cross-fit supervised encodings where necessary). Early stopping may use only an internal split of those training rows. The baseline uses no early stopping.

Quick training uses only quick-train; quick-dev labels are public for diagnosis, not fitting. Official validation/final training uses the same full frozen pool, never any held-out rows. Predict inputs may only be transformed with training-fitted preprocessing; no label-free holdout fitting or transductive feature generation.

## Execution and scoring
Budgets are distinct: quick 60s, validation 180s, final 240s. The whole candidate subprocess is bounded by that level's budget, stricter than a fit-only training cap. Timeout kills its process group; descendants are cleaned up after exit too. Environment caps common CPU libraries at 8 threads, candidate gets `--threads 8`, and its CPU affinity is limited to ≤8 available CPUs. GPU mask is preserved, identity checked, logical device 0 selected. Candidate must not bypass resource restrictions or spawn services.

Frozen `eval/eval.py` is the single scoring implementation. It independently loads labels, verifies frozen input hashes and schema, runs candidate `train.py`, rejects stale outputs, requires exact ordered unique prediction IDs and row count, and rejects nonfinite/out-of-range probabilities. It computes `sklearn.metrics.roc_auc_score` with label 1 positive. Result JSON contains score, n, level, cap, seed, measured whole-candidate/evaluator time, independently checked GPU identity, data/manifest/protocol/prediction hashes, entrypoint hash, and model hash if supplied. Optional candidate metadata is diagnostic, not scoring authority. Failure always writes `score:null`, an error, and exits nonzero. Metrics contain no individual held-out labels or labeled prediction table.

Quick predictions/logs go to its output directory. Official/final predictions and candidate logs stay in task `private/evaluations/`, not the worktree; model artifacts and aggregate metrics may go to harness run_dir. Neither predictions joined with labels nor official labels are copied to workers. Prediction hashes permit audit without exporting tables. AUC is a statistic over positive/negative pairs: per-row values whose mean is AUC are not invented. `per_item`, `score_std`, and `score_sem` are omitted, so the harness cannot infer a bogus paired-noise estimate. Uncertainty is unknown. Fixed seeds freeze data/method choices, not bitwise GPU arithmetic.

Protocol hash covers frozen evaluator, hardware guard, job wrapper, public data manifest, budget map and model seed. Individual file SHA-256 values identify data/predictions/artifact; models use a non-bitwise-deterministic GPU fit and may differ on rerun. The private/public manifests remain authoritative for exact rows and labels.

## Harness integration and holdout release
`job_enabled=true`, one GPU, no resumption, 360s job timeout, stall detector disabled, no harness early stopping, final_top_k=1. The queue assigns physical index 1 with `RSI_GPUS=1` and the caller's isolated `RSI_JOBS_DIR`; standalone checks inherit the UUID mask. Do not call CPU-only `run_monitored_eval`, which clears GPU visibility. The builder does not initialize/run a research experiment or a separate service. `rsi new-task` calls its own `verify_pack()` after builder completion and queues only quick/validation.

`eval/run.py --budget N --run-dir DIR` runs from candidate worktree, maps 60/180/240 to quick/validation/final exactly, and delegates to the same evaluator to write DIR/metrics.json. No separate training/scoring implementation. Ordinary evaluator supports --agent-dir, --seedset, --out, --python, --quiet. Copied evaluator + starter contents support quick without access to private files; official levels use the ORIGINAL task-pack evaluator and private path derived from its resolved file location.

Default `final_seedset=''` and `final_eval_root=false` prevent automatic final evaluation. A direct final call also fails before opening final data unless the original task.toml explicitly sets `final_seedset='final'`. Only the campaign owner at the end of the whole campaign may enable this; coordinate this pack-level gate with any already-initialized experiment's final_seedset, keep final_eval_root=false, and run final only for the chosen fixed method. Do not release at the end of an intermediate 8/32-node stage. Validation is adaptively selected and cannot be called independent confirmation. No evaluator ever submits to Kaggle.

Search policy: initial 8 sequential nodes / ≤16 implementation attempts, one hypothesis and ≤3 development trials each. Repeated seeds must be preplanned, not selected on official outcomes. Total expansions across the campaign ≤128 excluding root. The caller must configure these harness search/research limits and track continuation totals; this pack does not override framework configuration. A justified saturation result may stop early.
