# Why the tape-graded record loses — diagnosis and redesign **Date:** 2026-09-15 · **Author:** FP Operator Admin session · **Data:** /fp-data (GLBX 1m, 2010–2026; gex/history QQQ per-strike, 730 days) · **Scripts:** `scripts-2026-09-15/study{,2,3,4,5,6}.mjs` (originally in a temporary job directory under ) ## The question . The tape-graded picks were losing while the model-graded picks were winning. Is the tape record bad luck, bad exits, or a bad rule? ## 1. The published pick family has no edge — replicated three ways Replication of production mechanics (5m close through the level, no-chase >0.5R, fill next 1m open, stop wins ties, mark at the session bell) on 8 roots (CL NQ ES GC ZN RTY YM 6E), 2026 YTD, 3,971 sessions: | Variant | n | gross R/trade | net of $4 RT + 1 tick each way | |---|---|---|---| | OR 5m, target 1.5R | 5,505 | −0.033 | **−0.217** | | OR 5m, run to bell / 10R | 5,505 | −0.001 | −0.185 | | OR 15m, target 1.5R | 5,812 | −0.041 | −0.171 | | OR 30m, target 1.5R | 5,507 | −0.013 | −0.112 | | OR 15m + RVOL ≥ 2, run to bell | 592 | +0.008 | −0.091 | Target choice is irrelevant: hit rates track the breakeven line at every R multiple (0.5R 62.5% vs 66.7% needed; 1R 41.9% vs 50%; 1.5R 29.3% vs 40%; 2R 21.2% vs 33.3%). Breakeven-stop at +0.5R does not rescue it. Median friction is 4–10% of 1R; the family's gross edge is smaller than its friction. Independent agreement: the desk's own pre-registration `practitioner-noise/prereg/RESULTS-2026-08-22.md` (pn-001, ~10,500 NQ+ES trades, 2010–2026) already recorded **ORB REFUTED as a system — no cell beats friction**. Mesfin 2026 (MNQ, 947 days, 14 signal families): max gross edge 0.07–1.50 pts vs 2.0 pts friction. ## 2. Three mechanical aggravators 1. **Entry buys the exhaustion.** Bar-close entry arrives after the impulse — Mesfin's expansion-bar continuation is t = −10.96 next bar. ORB delay +15 bars beats +1 bar (−0.82 → +2.82 pts); pullback entry is refuted hard (19.3% win, 80.7% stop-out). 2. **The level is the worst price in the range.** Published S/R levels coincide with limit-order-book depth peaks (Kavajecz & Odders-White 2004); liquidity demanders at focal prices lose ~$1bn/yr (Bhattacharya, Holden & Jacobsen 2012); price crosses a published level without reversing only **39.2%** of the time (Osler 2000) — below the 40% a 1.5R target needs. Take-profits rest *at* the focal price, stops rest 1–10 ticks *beyond* it: a trigger at the level buys into the wall. 3. **The grading window truncates winners only.** A loser realises the full −1R; a winner still open at the bell is marked mid-flight. . ## 3. What replicates on our own tape - **Late-session momentum (Baltussen et al. JFE 2021), 702 days, gamma from our own archive (prior day, no look-ahead):** does NOT replicate. NQ all days −$68/contract net; ES **−$68, t = −2.90** (significantly negative). Gamma *ordering* is right — most-negative-gamma tercile is the best cell (NQ +$17, ES −$34) and most-positive is the worst (NQ −$124, t −1.97) — but no cell is significantly positive. Consistent with Kurth et al. 2026: short-horizon trend died post-2008 on small-tick index futures. - **Overnight drift (16:00 ET close → 09:30 ET open), 2020–2026, ~1,660 nights each:** | Root | net $/contract/night | win | t | ann. SR | RTH-day control | |---|---|---|---|---|---| | NQ | **+167.79** | 57.0% | 2.27 | 0.88 | +41 (t 0.43) | | GC | **+162.64** | 55.2% | 2.21 | 0.86 | −43 (t −0.79) | | ES | +65.21 | 55.2% | 1.60 | 0.62 | +14 (t 0.27) | | CL | +4.39 | 51.5% | 0.11 | 0.04 | −14 (t −0.45) | The entire drift is overnight; the day session is flat to negative everywhere. Caveats: long-only exposure correlated with buy-and-hold, a 2020–2026 bull sample, gap tail risk, and t ≈ 2.3 does not clear the t > 3 multiple-testing hurdle (Harvey, Liu & Zhu 2016). ## 4. Redesign 1. **Grade only on the tape.** Model verdicts must never enter the record; stand-asides must never carry R (). 2. **Publish net-of-friction expectancy**, with a random-level control (a random level in the day's range is "respected" 56% of the time — Osler 2000) and n ≥ 100 per bucket before a family is publishable. 3. **Retire level-triggered intraday scalps as the graded product**, or re-test them only as: trigger offset *past* the focal price, entry delayed k bars, top-tercile participation selection. Each is one parameter, pre-registered. 4. **Move the tradable claim to horizons where friction is a small share of the move** — overnight and event-driven — and size them as exposures, not as "picks". 5. Pre-register every change in `practitioner-noise/prereg/` before it trades; no re-tuning inside the evaluation window (self-attribution bias ratchets confidence up after streaks — Daniel & Hirshleifer 2016).