# RESULTS — the overnight drift under three layers: technical, event, behavioural **Run 2026-09-16 against the pre-registration frozen the same morning** (`2026-09-16-three-layer-overnight-prereg.md`). Script: `run_three_layer_overnight.py`. Machine-readable output: `2026-09-16-three-layer-overnight-results.json`. ## Verdict in one line **The overnight drift is real and almost entirely unconditionable.** Of 19 arms across three layers, exactly one passes every frozen test: the **pre-FOMC night on NQ**. The technical layer adds nothing, the behavioural layer adds nothing, and the wider event layer adds nothing that survives a placebo and a train/confirm split. ## Sample | Root | Nights | 2016-01-05 → 2026-09-11 | Mean net $/night | t | Total per contract | |---|---|---|---|---|---| | NQ | 2,581 | long, flat to flat | **$98.00** | 2.06 | $252,946 | | ES | 2,581 | " | $38.45 | 1.44 | $99,251 | | GC | 2,553 | " | **$107.33** | 2.22 | $274,018 | Net of the desk's round-turn costs ($14 NQ, $29 ES, $24 GC), quad-witching seams excluded. This extends the tape study's window back to 2016 and reproduces its finding: NQ and GC pay to hold the night; ES does not clear its own friction. ## The one thing that passed **N1 · pre-FOMC night · NQ — VALIDATED** | | | |---|---| | n | 53 nights | | Mean | **+$972.70** per contract per night | | t | 3.12 (p = 0.0018, below the 0.0033 Bonferroni threshold) | | Win rate | 69.8% | | Placebo | p = 0.009 — random 53-night samples rarely do this | | Train / confirm | +$700 (2016–21, n=16) → **+$1,091** (2022–26, n=37); signs agree | | Share | **20.4% of the total drift from 2.1% of the nights** | It also survives the multivariate test: with every other block in the regression, the pre-FOMC coefficient is **+$824 ± 311 (t 2.65)** and it is the only coefficient that is significant at even 0.05. This independently reproduces the pre-FOMC overnight result the desk found on 2026-08-25, on a longer sample, with a different pipeline, and with the event flags rebuilt from the new archive. ES shows the same sign (+$425, p = 0.005) but fails the corrected threshold; **GC shows nothing** (+$99, t 0.24), which is what you would expect if the effect is an equity-risk-premium story rather than a general "event" story. ## What failed, by layer **Technical (T1–T4): nothing.** The Lou–Polk–Skouras mirror does not appear — nights after a down day pay *less* on NQ ($74) than nights after an up day ($117), the opposite of the hypothesis and insignificant either way. Outsized prior moves, VIX buckets and weekend nights are all flat, and several flip sign between halves. GC's "after an up day" arm reaches p = 0.0027 but its placebo p is 0.24, meaning random samples of that size do this routinely: **refuted, not suggestive**. **Events beyond pre-FOMC: nothing that holds.** Mega-cap earnings after the close look large on NQ (+$612 heavy, +$646 light) but the heavy arm **flips sign between halves** (−$216 train → +$950 confirm) and neither clears the threshold. CPI/PCE/NFP, Fed speeches after 16:00, policy and geopolitics headlines, and EIA/WASDE are all indistinguishable from any random night. Fed speeches are the one arm with a *negative* point estimate on NQ (−$383), driven entirely by the confirm half, on 69 nights. **Behavioural: nothing.** This is the block the psychology library motivated, and it is the cleanest null in the study: | Hypothesis | NQ mean | p | Verdict | |---|---|---|---| | B1 attention-induced buying (pageviews z ≥ 1.0) | +$282 | 0.117 | refuted | | B2 disposition × attention (high attention after a down day) | +$380 | 0.180 | refuted | | B3 sensation seeking / chasing (after a >1.5σ up day) | +$34 | 0.847 | refuted | | B4 loss aversion (after 2+ down days) | +$96 | 0.465 | refuted | The decisive check is the placebo: **shuffling the attention series across dates produces a mean of $100.11 (t 0.70)** — statistically identical to the real series' effect on the baseline. Whatever the attention signal is measuring, it is not information about the next night. On GC the attention arms are *negative* (−$237, −$492) with a suspiciously good placebo score (p = 0.016, 0.002) and insignificant t-statistics, which is exactly what a spurious cell looks like when you test many of them. ## What this means for the desk 1. **Hold the night, stop trying to time it.** The drift pays; the conditioning does not. Every extra filter tested here would have removed nights without improving the average. 2. **Pre-FOMC on NQ is the exception, and it is small in count.** 53 nights in eleven years, about five a year. It is worth flagging in the playbook as a known, dated, event-conditioned long, with the caveat that its confirm-half strength (+$1,091) comes from the 2022–26 hiking-and-cutting cycle and may be regime-bound. 3. **The behavioural layer is now tested rather than assumed.** Attention-induced buying, the disposition effect, chasing and loss aversion are real findings about *retail equity investors*; none of them shows up in overnight index-futures returns at this resolution. They belong in the marketing and product work, where the evidence for them is direct, not in the trading model. 4. **The event archive earned its keep by producing a null.** It reproduced the one known effect and refused to manufacture others. That is the correct outcome for a tool whose job is to stop the desk from believing things. ## Caveats, stated plainly - **This is a long-only beta trade in a bull sample.** 2016–2026 contains one of the great equity runs. The drift is compensation for holding overnight risk; a regime that pays for that risk differently will change the number. - **No slippage beyond the fixed round turn**, and fills are assumed at the 16:00 and 09:30 prints. - **Event coverage is uneven.** Policy and geopolitics rows thin out before 2020, and the headline wire only starts in 2024 (GDELT is rate-limited, Alpaca's backfill is partial). A "policy/geo headline" arm therefore tests a sparser tape in the train half than the confirm half — one reason several of those arms flip sign. - **Attention is a proxy.** Wikipedia pageviews are not Google Trends, brokerage account opens or order flow. The null here is a null about *this* proxy. - **Pre-FOMC n = 53** against roughly 88 scheduled meetings in the window: the flag comes from the event archive's own FOMC rows, so coverage, not the market, sets the sample. - **Three roots × 19 arms is a lot of tests.** The Bonferroni correction was set for 15 on the primary root and applied there. Only one arm passed it, which is the point. ## What would change the verdict - A **longer or licensed news feed** (the wire back to 2016) would let the policy, geopolitics and earnings arms be tested on an even tape rather than a growing one. - **Order-flow or positioning data at the 16:00 close** (dealer gamma back-history, retail flow) would test a mechanism the current layers only proxy. - **Another eleven years.** At $98 a night with a t of 2.06, the baseline itself is only just significant; the conditioning arms need far more nights than they have.