# RVOL-conditioned opening-range breakout on futures — RESULTS **Date:** 2026-08-25 · **Pre-reg:** `2026-08-25-rvol-orb-prereg.md` (frozen 19:30 MDT, amended 19:45 MDT, both **before** any result was computed) · **Data:** `/fp-data` GLBX 1m, 8 instruments, 2010-06-07 → 2026-08-25 (NQ: 4,033 sessions, 3,913 trades) · **Scripts:** `fp-backtest/rvol_orb.py`, `fp-backtest/rvol_orb_battery.py` · **Cost model:** desk `COST_PTS` (NQ 1.0 pt RT) · **Raw output:** `2026-08-25-rvol-orb-results.txt` --- ## Question The asset-selection library (built the same day) has one finding every brief converged on: **selection dominates entry.** Zarattini–Barbon–Aziz (2024) run one 5-minute opening-range breakout on the whole US stock universe (Sharpe 0.48, worse than the index) and the *identical entry rule* on the top-20 names by opening-bar relative volume (Sharpe 2.81, alpha 35.8 %, beta 0.00), with per-trade PnL rising **monotonically** across RVOL bins from <1× to >30×. This desk has already refuted ORB on NQ/MNQ **unconditionally**. Those two results are not in conflict — Zarattini's claim is that the unconditional rule is *supposed* to be break-even and that all the edge lives in the selection layer. So: **does RVOL conditioning rescue ORB on futures?** ## VERDICT: **REFUTED** | Pre-registered criterion | Result | | |---|---|---| | 1. Primary cell t > 3.0 and positive | **t = −0.33**, mean **−1.11 pt** | **FAIL** | | 2. Monotonicity ρ > 0 across quintiles | ρ = +0.100 | pass on the letter, **fails on substance** (see below) | | 3. ≥ 4 of 7 instruments replicate at t > 2.0 | **0 / 7** | **FAIL** | Primary cell — NQ · 5m OR · Entry B (desk confirm rule) · Q5 · 1.0 pt friction: **n = 577, mean −1.114 pt/trade, t = −0.33, hit 33.3 %, total −642.8 pt.** --- ## The result that matters most: it is not a cost story | NQ · Entry B | all trades | top RVOL quintile | |---|---|---| | **gross, 0.0 pt friction** | +0.744 pt, t = 0.61 | **−0.114 pt, t = −0.03** | | 0.5 pt | +0.244, t = 0.20 | −0.614, t = −0.18 | | **1.0 pt (desk)** | −0.256, t = −0.21 | **−1.114, t = −0.33** | | 2.0 pt (Mesfin's MNQ assumption) | −1.256, t = −1.04 | −2.114, t = −0.63 | **Before a single tick of friction, the high-RVOL bucket is already worse than the unconditional set.** The unconditional ORB carries a small, statistically insignificant positive gross drift (+0.74 pt, t = 0.61) which friction erases — consistent with both Zarattini's "break-even unconditional" premise and the desk's existing refutation. But RVOL conditioning *subtracts* rather than adds. There is no cost level at which this rule becomes viable. ## Monotonicity: technically passed, substantively absent | bucket | n | mean net | t | hit % | RVOL range | |---|---|---|---|---|---| | Q1 | 605 | −1.396 | −0.47 | 24.3 | 0.02–0.78 | | Q2 | 870 | **+1.436** | 0.58 | 27.1 | 0.68–0.94 | | Q3 | 953 | −0.910 | −0.40 | 26.8 | 0.90–1.13 | | Q4 | 908 | +0.112 | 0.04 | 28.4 | 1.07–1.43 | | Q5 | 577 | −1.114 | −0.33 | 33.3 | 1.29–3.85 | ρ = +0.100 clears the letter of criterion 2, and I am recording that **the criterion was written too loosely**. A Spearman ρ from five points has no power; the actual pattern (+, −, +, −) is noise, and the best bucket is Q2, not Q5. Criterion 2 should have required a *strictly increasing* ordering, as Zarattini reports on stocks. Scored honestly, all three criteria fail. The one directionally-consistent signal: hit rate does rise with RVOL (24.3 % → 33.3 %). Higher opening volume genuinely does make a breakout more likely to *hold*. It does not make it profitable, because the winners shrink faster than the hit rate grows. ## Why the transfer fails — the durable finding **NQ's opening-range RVOL has almost no dispersion.** | percentile | p50 | p75 | p90 | p95 | p99 | max | |---|---|---|---|---|---|---| | NQ RVOL | 1.01 | 1.22 | 1.47 | 1.66 | 2.23 | **3.85** | Sessions above 5× RVOL in 16 years: **zero**. Above 3×: six (0.15 %). Zarattini's gradient runs from <1× to **>30×**, and the payoff he reports (+0.38R) lives in the **>30× bin**. That region does not exist on an index future. "Q5 on NQ" means RVOL ≈ 1.3–1.9 — which in his stock table is the roughly-break-even bin, not the paying one. The reason is structural, and it generalizes: **Zarattini's selection is cross-sectional.** On any given day, a 3,000-name universe contains *some* stock with a catalyst trading 30× its normal volume, and the method's job is to find it. A single index future offers no cross-section to select from — its opening volume is a stable, mean-reverting series around its own base. Conditioning NQ on its own volume history is not the same operation as picking the extreme name out of thousands, and it should not have been expected to behave the same way. This is the same lesson the library's own §3 warns about in the abstract ("A monthly, and C or D intraday" grade splits) applied to a *venue* transfer rather than a horizon transfer. ## Instability across years (NQ Q5) Positive-mean years: **6 of 17**. IS (<2020) mean −0.95, t = −0.60; OOS (≥2020) mean −1.44, t = −0.15. 2024 **+58.0 pt/trade** · 2025 **−30.6** · 2022 +22.2 · 2021 −20.7 · 2020 −15.2 · 2015 −8.8. This is Mesfin's documented failure mode in a more extreme form: single years dominate the pooled number in both directions. Any pooled statistic here is close to meaningless. ## Replication set — every instrument negative | symbol | n | mean net | t | hit % | p99 RVOL | |---|---|---|---|---|---| | NQ | 577 | −1.114 | −0.33 | 33.3 | 2.23 | | ES | 670 | −0.901 | −1.29 | 27.9 | 2.21 | | RTY | 341 | −0.114 | −0.12 | 34.9 | 2.05 | | YM | 632 | −9.911 | −1.78 | 30.2 | 2.22 | | CL | 752 | −0.046 | −1.80 | 28.3 | 3.35 | | GC | 704 | −0.566 | −1.49 | 30.0 | 4.04 | | ZN | 779 | −0.062 | −9.83 | 22.6 | 3.58 | | 6E | 464 | −0.001 | −8.73 | 20.9 | 4.37 | **0 of 7 replicate.** All eight are negative in their own top RVOL quintile. Note every instrument's p99 RVOL sits between 2 and 4.4 — none of them reach the region where the stock effect lives. ZN and 6E show large-magnitude t-stats because they lose a small, *consistent* amount (the friction) with low variance, not because the effect is large. ## Bug found and fixed mid-study (documented, not hidden) The first replication run produced impossible numbers — 6E "losses" of 0.52 (≈ 650× its median OR width) and t-stats of −36. Cause: `walkPlan`'s stop-slippage floor `max(0.5, 0.25 × excursion)` is **0.5 NQ points = 2 NQ ticks**. Copying the literal `0.5` to other instruments makes it 16 ticks on ZN and 10,000 ticks on 6E. Fixed by expressing the floor as **2 ticks per instrument** (`rvol_orb.py` `TICK` / `SLIP_FLOOR`). **NQ results are bit-identical before and after the fix**, so the primary test was never affected. Anyone porting `walkPlan` to a non-NQ instrument must re-express that constant — it is silently NQ-specific. ## What this changes on the desk **Nothing is shipped. ORB stays refuted**, now with a mechanism and cross-instrument evidence rather than a single-instrument result. Concretely: 1. The existing `refuted-adjacent` status for ORB in `practitioner-noise/` and the asset-selection catalog is **upheld and strengthened** — extend it from NQ/MNQ to all 8 desk instruments. 2. The library's headline "selection dominates entry" is **not transferable to single-instrument futures selection**. It is a cross-sectional result and needs a cross-section. 3. The productive next test is therefore the *actual* analogue: apply the selection principle **across instruments** — rank NQ/ES/RTY/YM/CL/GC/ZN/6E each morning by their own normalized opening activity and trade only the day's leader. That is catalog row **as-123** (futures lead-lag / trade the price-discovery leader, grade A, `fp_data_testable = true`) and it is the one direction this study actively motivates. 4. `00-taxonomy.md` §7 open question 4 (precision/recall of an RVOL ≥ 2 screen on futures) is **answered in the negative** for the ORB use. ## Honest caveats - One entry per session, no re-entry after a stop; a re-entry variant is untested. - Stop = OR width re-projected from the fill (desk semantics). A tighter or ATR-scaled stop is a different strategy and was deliberately not swept — sweeping it would have manufactured the snooping problem this pre-registration exists to prevent. - 320 cells were pre-counted; one was confirmatory. No deflated-Sharpe adjustment was needed because nothing came close to passing. - The gamma-conditioned version of this question — the one I actually wanted — remains **untestable** until the forward-only GEX archive reaches n ≈ 200 sessions (≈ 2027-06). It is not answered here. - Findings are hypotheses about a refutation, which is the cheapest kind to trust: this study removes a candidate, it does not add an edge.