# PRE-REGISTRATION — does retail attention condition the overnight drift? **Frozen 2026-09-16, BEFORE any result was computed.** > **ANSWERED 2026-09-16** by `2026-09-16-three-layer-overnight-results.md`, which ran this hypothesis as block B1 on a larger design (2,581 NQ nights, 2016→2026). **Refuted:** attention z ≥ 1.0 gives +$282/night, p = 0.117, and shuffling the attention series across dates reproduces the effect ($100.11 vs a $98.00 baseline). H2–H4 were not run separately; the quintile and cross-root work is superseded by that study's arms. This file stays on record as the frozen design. Successor to `2026-09-16-overnight-event-nights-prereg.md`, which found the pre-FOMC night reproduces (+$973 net on NQ, n=53, t=3.12) and that **no event arm explained the rest** — quiet nights still carried about 31% of NQ gross. That result is what motivates this one: if scheduled events do not explain the drift, perhaps *crowd attention* marks the nights that do. ## Data `attention/wikipedia/*.jsonl` (added 2026-09-16): daily en.wikipedia pageviews for 23 articles, 2016-01-01 → today, mapped to roots. Point-in-time by construction — a day's counts publish the next day at 09:00 UTC (05:00 ET), so a night starting 16:00 ET on day D may use days up to D-1 only. `features/context_daily.jsonl` carries `attn_{nq,gc,cl,btc}_z`: the z-score of day D-1's total against the prior 60 usable days. Overnight returns come from `/fp-data/glbx//1m`, 16:00 ET → 09:30 ET, net of $4 round turn, exactly as the predecessor study computed them. ## Hypotheses (frozen) - **H1.** NQ overnight drift is higher on nights following a high-attention day (`attn_nq_z >= 1.0`) than on other nights. - **H2.** The relationship is monotone across attention quintiles of `attn_nq_z`. - **H3.** The effect, if any, is not merely the event arms already tested: it survives excluding pre-FOMC nights and nights flagged `amc_heavy`, `cpi_pce_nfp`, `eia_wasde` or `policy_geo`. - **H4 (cross-root control).** The same test on GC using `attn_gc_z`, and on CL using `attn_cl_z`. ## Method 1. Join `features/context_daily.jsonl` to the overnight returns by date. Sample: 2016-01-01 → the archive's last complete night, all nights where both exist. 2. Primary metric: mean net dollars per contract per night by arm, with a two-sided t-test and a bootstrap 95% interval (10,000 resamples). 3. H2: quintile means with a Spearman rank correlation across quintiles. 4. H3: repeat H1 on the subset with every known event flag false. 5. Exclude quad-witching seams (`quad_witching` true) throughout. 6. **Train/confirm split:** 2016–2021 and 2022–2026. A result that flips sign between halves is reported as refuted, not averaged. 7. **Multiple testing:** four hypotheses × three roots = 12 tests. Bonferroni threshold p < 0.0042. ## Decision rules (frozen) - **Validated** only if H1 holds in both halves at the corrected threshold, H2 is monotone in sign, and H3 survives on the no-event subset. - **Suggestive** if it passes in the full sample but not in both halves. Suggestive means "do not trade it", exactly as in the previous studies. - **Refuted** otherwise, and written up as refuted. ## Stopping rule No re-binning of the z threshold, no swapping the window, no adding articles after seeing results. The 60-day window and the 1.0 threshold are frozen here. If the answer is inconclusive, it is reported inconclusive and the attention series stays context, not signal. ## Prior expectation Low. Wikipedia attention is a crowd-interest proxy graded C at best; the desk has already refuted the event-day premium, COT positioning and net liquidity. The value of running it is that it is cheap and that a negative result further narrows where the overnight drift can be coming from.