# RESULTS — what the desk can actually publish **Run 2026-09-16 against the pre-registration frozen the same morning** (`2026-09-16-product-strategies-prereg.md`). Script: `run_product_strategies.py`. Machine-readable: `2026-09-16-product-strategies-results.json`. ## Verdict in one line **Twenty-nine arms run against a 27-test pre-registration, zero validated.** No root pays to be long during the day session after costs, no overnight timing refinement beats the plain hold, no calendar filter improves it, and the first 30 minutes after a gap is a coin flip. The pre-committed fallback in the prereg is therefore the answer: **the desk cannot sell intraday entries.** What it can sell is the overnight structure, a dated event calendar, and honest risk geometry. ## Deviation from the pre-registration, declared The prereg froze family B as "4 tests (NQ and GC)". Implementing the 2×2 grid on *both* roots yields three new cells per root (the fourth cell is family A), so **six** B arms ran and the suite totalled **29 inferential arms, not 27**. The Bonferroni threshold should therefore be 0.05/29 = **0.00172**, not the 0.00185 printed by the runner. No verdict is sensitive to the difference: the smallest p among the non-artifact arms is 0.0074 (family C), an order of magnitude above either threshold. The deviation is recorded here rather than absorbed silently. ## The correction that matters most Three arms came back with enormous t-statistics — ZN overnight (−$37.74/night, t −6.90), ZN day session (−$33.84/day, t −7.81), 6E overnight (−$31.08, t −2.03). None of them is an edge. Because the round turn is a *constant* subtracted from a low-variance series, charging $35.25 to a market that barely moves manufactures a huge t. Adding the cost back: | Root | Overnight gross $/night | Round turn | Net | Day-session gross $/day | Net | |---|---:|---:|---:|---:|---:| | NQ | **88.36** | 14.00 | 74.36 | 43.85 | 29.85 | | ES | 56.42 | 29.00 | 27.42 | 33.17 | 4.17 | | YM | 40.17 | 14.00 | 26.17 | 26.49 | 12.49 | | RTY | 46.92 | 14.00 | 32.92 | −15.46 | −29.46 | | GC | 69.32 | 24.00 | 45.32 | −9.37 | −33.37 | | CL | 30.48 | 24.00 | 6.48 | 0.80 | −23.20 | | ZN | **−2.49** | 35.25 | −37.74 | **1.41** | −33.84 | | 6E | −14.58 | 16.50 | −31.08 | 9.35 | −7.15 | ZN's gross drift is −$2.49 a night and +$1.41 a day: statistically indistinguishable from zero, and the entire "signal" is my own commission assumption. 6E's −$14.58 gross is the euro's 2010→2026 decline, which is beta, not alpha. **Any rule built on those two t-statistics would have been a rule about the cost line.** This is the kind of thing the suite exists to catch. ## Family A — overnight hold, eight roots Long the 16:00 ET close, flat the 09:30 open, 2010–2026, quad seams excluded. | Root | n | Net $/night | t | p | Win % | Halves | |---|---:|---:|---:|---:|---:|---| | NQ | 3,156 | **74.36** | 2.19 | 0.0285 | 56.2 | agree | | RTY | 1,782 | 32.92 | 1.76 | 0.0793 | 54.3 | agree | | ES | 3,156 | 27.42 | 1.43 | 0.1514 | 53.5 | **flip** | | YM | 3,160 | 26.17 | 1.82 | 0.0688 | 53.8 | agree | | GC | 3,126 | 45.32 | 1.22 | 0.2238 | 52.2 | **flip** | | CL | 3,143 | 6.48 | 0.30 | 0.7605 | 51.8 | **flip** | | ZN | 3,229 | −37.74 | −6.90 | ~0 | 43.9 | agree (friction, see above) | | 6E | 1,959 | −31.08 | −2.03 | 0.0421 | 46.1 | agree (beta, see above) | Only **NQ** is suggestive, and it fails the corrected threshold (needs p < 0.00172). On the longer 2010–2026 window the NQ overnight is *weaker* than on the 2016–2026 window the three-layer study used ($74 vs $98): the drift is concentrated in the post-2016 half, which is exactly what you would expect of a bull-sample risk premium and exactly why it should not be sold as a mechanical edge. GC's overnight **flips sign between halves** here. The three-layer study's $107/night on 2016+ is the confirm half of this longer sample. That is a material downgrade to gold's overnight story. ## Family B — overnight timing grid | Arm | n | Net $/night | t | p | |---|---:|---:|---:|---:| | NQ 16:00→09:30 (family A) | 3,156 | 74.36 | 2.19 | 0.0285 | | NQ 18:00→10:00 | 4,063 | 74.36 | 2.29 | 0.0219 | | NQ 16:00→10:00 | 3,156 | 68.93 | 1.82 | 0.0680 | | NQ 18:00→09:30 | 4,064 | 59.19 | 2.09 | 0.0370 | | GC 16:00→10:00 | 3,137 | 27.07 | 0.71 | 0.4804 | | GC 18:00→09:30 | 4,055 | 13.07 | 0.41 | 0.6793 | | GC 18:00→10:00 | 4,064 | −7.99 | −0.24 | 0.8083 | The plain hold and the 18:00→10:00 variant land on the **same $74.36** to the cent — a coincidence, but a telling one: the 16:00–18:00 stub and the 09:30–10:00 stub cancel. **There is nothing to optimise.** Every GC variant is dead, two of the three negative. Moving the entry to the Globex reopen buys nothing but 900 more nights of the same average. ## Family C — calendar filters | Arm | n | Net $/night | t | p | Halves | |---|---:|---:|---:|---:|---| | NQ turn-of-month nights | 783 | **−17.97** | −0.26 | 0.7919 | **flip** | | NQ non-turn-of-month | 2,373 | **104.82** | 2.68 | 0.0074 | agree | | NQ Nov–Apr | 1,515 | 81.86 | 1.63 | 0.1033 | agree | **The turn-of-month effect is inverted in this tape.** The four sessions that retail literature says carry the month's drift are the *only* subset with a negative overnight mean, and the remaining 2,373 nights carry all of it and more. The turn-of-month arm flips sign between halves, so this is refuted rather than a tradeable short — but it is a clean, publishable refutation of a widely repeated claim, and it belongs in the practitioner-noise tree. "Sell in May" gets no support either: Nov–Apr's $82 is inside the noise around the $74 baseline. ## Family D — the day session **No root pays to be long 09:30→16:00 after costs.** Best is NQ at +$29.85/day on t 0.75; five of eight roots are negative net. Against 3,985 sessions, NQ's day session is a coin flip with a small positive tilt that friction eats. Event-day arms: | Arm | n | Net $/day | t | p | |---|---:|---:|---:|---:| | NQ RTH, CPI/PCE/NFP mornings | 258 | −290.01 | −1.15 | 0.2482 | | NQ RTH, FOMC decision days | 53 | −387.40 | −0.59 | 0.5539 | Both negative, neither significant, the FOMC arm flips sign between halves. **Refuted.** Note the asymmetry with the overnight work: the pre-FOMC *night* is the one validated effect the desk owns (+$972.70, n=53), and the FOMC *day* that follows it is worth −$387 with no significance. The premium is paid for holding the risk into the decision, not for trading the decision. ### NQ hour profile (descriptive, no test attached) | Bucket | n | Mean gross $ | t | Mean absolute move $ | |---|---:|---:|---:|---:| | 09:30–10:00 | 4,114 | 15.72 | 0.95 | **616** | | 10:00–11:00 | 4,112 | −22.80 | −1.31 | **654** | | 11:00–12:00 | 4,096 | 30.72 | 2.21 | 495 | | 12:00–13:00 | 4,016 | −0.77 | −0.07 | **411** | | 13:00–14:00 | 3,981 | 18.29 | 1.33 | 411 | | 14:00–15:00 | 3,983 | 11.46 | 0.94 | 416 | | 15:00–16:00 | 3,984 | −4.93 | −0.33 | 492 | No hour has a directional edge worth trading (the 11:00 bucket's t 2.21 is one of seven descriptive buckets and is +$17 net of friction). What the profile *does* show is **where the range lives**: the first ninety minutes carry 50–60% more movement than midday, and midday is 33% narrower than the open. That is a real, stable, publishable fact about when a day trader is paid for attention and when they are paying commissions for noise. ## Family E — gap follow-through | Arm | n | Net $ per trade | t | p | |---|---:|---:|---:|---:| | Trade with the gap, 09:30→10:00, after AMC mega-cap earnings | 100 | −121.05 | −0.59 | 0.5551 | | Trade with the gap, 09:30→10:00, after a quiet night | 333 | −7.08 | −0.09 | 0.9292 | Refuted in both event classes. Gap continuation is not a strategy in the first 30 minutes, and the earnings-night version is the worse of the two. Combined with the already-refuted opening-range family, **the desk now has direct evidence against every simple open-driven entry it could publish.** ## Descriptive risk geometry for day traders Computed after the fact and labelled descriptive — no test attached, nothing here is a signal. NQ RTH, gross, 2010–2026: | Weekday | n | Mean $ | Mean absolute move $ | |---|---:|---:|---:| | Mon | 762 | 253.23 | 1,358 | | Tue | 835 | 35.69 | 1,354 | | Wed | 832 | 24.64 | 1,514 | | Thu | 817 | −79.69 | 1,553 | | Fri | 739 | −4.62 | 1,474 | On the 2020+ window where event flags exist (baseline mean absolute day move **$2,776**): - CPI/PCE/NFP mornings: mean absolute move **$2,880** — only **4% wider** than an ordinary day. - Quiet days: **$2,601** — **6% narrower**. **The folk claim that CPI and payrolls days are dramatically wilder than normal days is not supported.** They are barely wider, and their mean return is −$276. The risk on those days is directional surprise, not range expansion. Finally, the overnight/day split on NQ gross, 2010–2026: **$278,855 accrued overnight across 3,156 nights versus $174,740 across 3,985 day sessions.** Per unit of time held, the night pays roughly twice the day. That is the single most important structural fact in this study, and it is the opposite of how a day-trading product is usually sold. ## What goes into the product Written against the decision rules frozen in the prereg. **1. The Desk Read keeps exactly one mechanical, dated rule: the pre-FOMC night on NQ.** It is the only thing in the last four studies that passed a corrected threshold, a placebo and a split. Roughly five dates a year, published in advance, carrying the CFTC 4.41 statement and its own sample size. Nothing else here earns a "do this." **2. The overnight structure becomes context, not a call.** Publish the drift as a described property of the tape — where it lives (post-2016), how thin it is ($74/night on a t of 2.19, well short of the corrected threshold), that it is long-only beta in a bull sample, and that no filter tested improves it. It informs *sizing and holding* decisions; it is not a signal. **3. For members who day-trade, the deliverable is timing and risk geometry, not entries.** Concretely: - **The hours that pay attention:** 09:30–11:00 carries 50–60% more range than midday; 12:00–14:00 is the narrowest stretch of the session and the most expensive place to pay commissions. - **What a day actually costs:** the average absolute NQ day is ~$1,450 gross per contract over the full sample and ~$2,776 on the 2020+ tape. Position size against that number, not against a target. - **Event days:** CPI/PCE/NFP mornings are only ~4% wider than normal, but the mean is negative and the tails are directional. The stand-down argument for those days is about surprise risk, not volatility. - **What we have tested and cannot offer:** opening range, gap continuation, crown policies, wall and cage rules, flow direction, gamma-conditioned late-day momentum, and now RTH direction on eight roots. Publishing this list *is* the product — it is what separates the desk from the noise. **4. Two refutations go on the ledger as publishable content.** The inverted turn-of-month and the "event days aren't wilder" finding are both direct, evidenced contradictions of widely repeated retail claims, and both are exactly the kind of thing the practitioner-noise tree exists to carry. **5. Nothing goes to the ledger as an open question except the Monday tilt** (+$253 gross RTH mean, n=762), which was not pre-registered and must be frozen in its own prereg before it is tested. It is not to be mentioned in member-facing material until then. ## Caveats - **Long-only, bull sample.** 2010–2026 contains one of the great equity runs. Every positive number here is compensation for holding risk in a regime that paid for it. - **Friction is a fixed $4 commission plus two ticks.** It dominates the low-variance roots entirely (see the correction above) and it is an assumption, not a measurement. A member paying different costs gets different answers, and for ZN the answer changes sign. - **Fills assumed at the 16:00, 18:00, 09:30 and 10:00 prints.** No slippage beyond the round turn, no partial fills. - **Event flags only exist from 2020-01-03**, not 2016 as the prereg assumed. Family D's event arms and family E therefore have a train half of two years, not six, and their splits are correspondingly weak. Stated here rather than adjusted after the fact. - **The placebo test is degenerate for full-population arms.** Families A, B and the eight D root arms draw their placebo from their own population, so `placebo_p = 1.000` is uninformative there by construction. Those verdicts rest on p-value and sign agreement alone. The placebo is only meaningful for the conditional arms (C, the D event arms, E). - **RTY starts 2017 and 6E's coverage is partial** — 1,782 and 1,959 nights respectively against ~3,150 for the majors. - **Twenty-nine arms, one pre-registered threshold, zero passes.** That is the honest headline, and it is more useful to the business than a thirtieth arm would have been.