HoldoutLabs
Scientific audits for trading strategies
Strategy Fragility Audit
Report 93ebca5673ec · 2026-09-25 20:13 UTC
holdout-audit 0.3.0 · rubric v1.1

Batch 2: Sector momentum rotation (12-1 month) on 9 sector SPDRs

Sector momentum rotation (12-1 month), as published, on 9 sector SPDRs. Benchmark: buy-and-hold SPY. Part of Indicator Audit Batch 2, 'Classic market rules' (14 audits, one sealed preregistration, Holm across the batch).

Executive summary

Fragility grade
F
5/16 points (31%)

Very fragile. The backtest shows no robust edge over its benchmark once selection, holdout and stability are accounted for. The grade measures how fragile the historical evidence is. It is not a forecast and not a recommendation.

Declared objective: beat the benchmark, declared by the client on 2026-09-25T20:11:24Z (declaration SHA-256 ff8cace63cad0597…).

Benchmark: buy-and-hold of the traded asset. This is the rubric v1.1 default. The strategy did not beat its benchmark under the rubric definition, so the grade is at most C.

Excess Sharpe vs benchmark (95% CI)
0.07 [-0.31, 0.45]
Excess return a year
0.7%
Absolute Sharpe
0.57
Deflated Sharpe (N=12)
0.491
PBO (CSCV)
0.988
Holdout excess Sharpe (in → out)
0.16 → -0.12
Max drawdown
48.1%
Buy-and-hold Sharpe, same period
0.52

Checks that failed: Selection-adjusted significance vs benchmark (Deflated Sharpe); Probability of backtest overfitting (CSCV); Multiple-testing haircut vs benchmark (Harvey-Liu); Beats the benchmark after data snooping (Hansen SPA); Holdout degradation (excess Sharpe).

Scorecard

CheckResultValueRule
Selection-adjusted significance vs benchmark (Deflated Sharpe)FAILDSR 0.491 (N = 12)PASS DSR >= 0.95; CAUTION >= 0.80; else FAIL (on excess returns)
Probability of backtest overfitting (CSCV)FAILPBO 0.988 (12870 splits)PASS PBO <= 0.20; CAUTION <= 0.50, or PBO > 0.50 with P(OOS excess SR < 0) <= 0.10; else FAIL
Multiple-testing haircut vs benchmark (Harvey-Liu)FAILadj. p 1.0000 (BHY)PASS one-sided adjusted p <= 0.05; CAUTION <= 0.10; else FAIL (excess-return t-statistic; evidence for the strategy only)
Beats the benchmark after data snooping (Hansen SPA)FAILSPA p 0.696; RC p 0.670PASS p <= 0.05; CAUTION <= 0.10; else FAIL
Holdout degradation (excess Sharpe)FAILexcess SR 0.16 -> -0.12PASS holdout and in-sample excess SR > 0 and holdout >= 50% of in-sample; CAUTION holdout excess SR > 0; else FAIL
Period and regime stability (vs benchmark)CAUTION52% of years aheadPASS >= 60% of years ahead of benchmark and excess SR > 0 in every volatility tercile; CAUTION >= 50%; else FAIL
Transaction-cost headroom (absolute)PASSbreak-even 203.9 bpsPASS break-even cost >= 20 bps per unit traded; CAUTION >= 5 bps; else FAIL (absolute)
Parameter plateau (absolute)PASSneighbours 89% of chosenPASS neighbour SR >= 70% of chosen SR; CAUTION >= 40%; else FAIL, or FAIL if chosen SR <= 0 (absolute)

How the grade is set (Holdout Labs Fragility Rubric v1.1, sealed before this report was produced). PASS = 2 points, CAUTION = 1, FAIL = 0; N/A checks are excluded. Share of available points: A ≥ 85%, B ≥ 70%, C ≥ 55%, D ≥ 40%, otherwise F. Hard caps: Objective cap (beat the benchmark): if the strategy does not beat its benchmark (SPA check FAIL, or annualised mean excess return <= 0), the grade is at most C. Objective cap (reduce drawdown): if the drawdown reduction is not robust (drawdown check FAIL) or the declared return-cost tolerance is breached (tolerance check FAIL), the grade is at most C. Selection cap: if the deflated check FAILS, the grade is at most C. Overfitting cap: if the PBO check FAILS, the grade is at most D. Disclosure cap: if the number of variants tried was not declared and no variant matrix was supplied, the grade is at most B.

What was audited

Source: Moskowitz, T. J. & Grinblatt, M. (1999). Do Industries Explain Momentum? Journal of Finance 54(4), 1249-1290; Faber, M. T. (2010). Relative Strength Strategies for Investing (SSRN 1585517). Rule: At each month end rank the nine original sector SPDRs by total return from 12 months ago to 1 month ago, and hold the top three equally weighted until the next month end. Audited variant: lookback=12|skip=1|top=3. Grid: 12 variants (see the preregistration).
Sample2000-02-01 to 2026-09-24 (6702 periods, 252 per year)
Variants tried (N)12 · variant matrix of 12 columns supplied
Base-case cost5 bps per unit traded (one way)
Holdout split2018-01-02
Compound annual return, strategy (net)9.3%
Buy-and-hold of the underlying, same periodcompound annual return 8.5%; annualised Sharpe 0.52; max drawdown 55.2%

Results in detail

Growth of 1 (log scale)holdout12510200020032006200920122015201820212024StrategyBuy and hold
Growth of 1 on a log scale, net of base-case costs, with buy-and-hold of the underlying in grey. The shaded area is the holdout period.

1. Sharpe ratio and its uncertainty

Rubric checks use the excess series (strategy minus buy-and-hold of the traded asset); the absolute series is shown for reference.

Excess over benchmarkAbsolute
Annualised Sharpe (sqrt-time scaling)0.0720.566
Annualised Sharpe (Lo 2002 autocorrelation-adjusted)0.0790.701
Standard error, annualised (non-normal)0.1940.195
Standard error per period: IID-normal / non-normal0.0122 / 0.01220.0122 / 0.0123
Skewness / kurtosis-0.35 / 9.38-0.15 / 12.70
Probabilistic Sharpe ratio vs 00.64530.9982
Minimum track record for 95% confidence518.3 years8.5 years
Annualised mean / volatility0.7% / 9.3%10.7% / 18.9%

2. The variant-count effect (Deflated Sharpe)

Deflated Sharpe ratio against number of variants tried00.20.40.60.811: 0.64512: 0.59825: 0.535510: 0.4991012: 0.4911220: 0.4682050: 0.43350100: 0.409100200: 0.388200500: 0.3625001000: 0.3441000variants tried (log scale)
The same backtest, judged as the best of N tries. The dark dot is the declared N = 12; the dashed line is the 0.95 pass mark. Computed on the excess Sharpe over the benchmark. The more variants tried, the higher the bar.
N triedNoise hurdle (annual excess SR)DSR
10.0000.6453
20.0240.5983
50.0550.5355
100.0730.4993
120.0770.4907
200.0880.4684
500.1050.4330
1000.1170.4094
2000.1280.3878
5000.1410.3619
10000.1500.3440

Dispersion of trial Sharpe ratios used: 0.046 (annualised standard deviation), estimated from the variant matrix.

3. Probability of backtest overfitting (CSCV)

Logit of the in-sample winner's out-of-sample rank-2-10median
PBO0.988
Splits (blocks)12870 (16)
Variants12
P(IS winner trails benchmark out of sample)0.708
Median OOS excess Sharpe of IS winner (annual)-0.09
Degradation slope (OOS on IS)-1.05

Histogram of logit ranks: values left of the dashed line are splits where the in-sample winner finished at or below the out-of-sample median (logit 0).

4. Multiple-testing haircut

t-statistic (excess SR x sqrt(years))0.37
p-value, single test0.35460
Bonferroni p (12 tests)1.00000
Sidak p0.99478
Holm p1.00000
BHY p1.00000
Haircut excess Sharpe (Bonferroni)0.000 (100% haircut)
Haircut excess Sharpe (BHY)0.000

5. Reality Check and SPA

Benchmarkbuy-and-hold of the traded asset
Variants in the test12
Best mean excess return (annualised)1.44%
White Reality Check p0.670
Hansen SPA p (consistent / lower / upper)0.696 / 0.691 / 0.696
Bootstrapstationary, 1000 draws, mean block 19

6. Holdout degradation

PeriodsAnnualised excess Sharpe
In-sample (before 2018-01-02)45080.157
Holdout2194-0.115
Holdout / in-sample-0.73
Consistency p-value0.422

7. Period and regime stability

Return minus benchmark return, by calendar year-20%0%20%2000: 5.9%2001: -2.9%2002: 12.0%2003: 0.8%2004: 6.9%2005: 10.2%2006: -2.1%2007: 4.5%2008: 4.4%2009: -2.8%2010: 0.2%2011: -3.9%2012: -10.6%2013: 3.6%2014: -2.7%2015: 0.9%2016: 2.6%2017: -1.4%2018: -4.7%2019: -3.3%2020: 1.1%2021: -5.1%2022: 31.9%2023: -22.1%2024: -11.4%2025: -7.0%2026: 4.0%200020032006200920122015201820212024
Strategy compounded return minus benchmark compounded return, by year: ahead in 14 of 27 years (absolute return positive in 22).
Volatility regime (trailing 21-day, lagged)PeriodsAnnualised mean excessAnnualised excess Sharpe
low vol22270.5%0.08
mid vol22272.0%0.25
high vol2227-0.4%-0.03

8. Transaction-cost sensitivity

One-way cost (bps per unit traded)0125102050
Annualised Sharpe0.580.580.570.570.550.520.44
Annualised mean return11.0%10.9%10.9%10.7%10.4%9.9%8.3%

Turnover 5.4 units per year; break-even cost 203.9 bps.

9. Parameter sensitivity

Annualised Sharpe across the parameter gridlookback ↓ / skip →0160.50lookback=6, skip=0: 0.5030.52lookback=6, skip=1: 0.51990.53lookback=9, skip=0: 0.5290.48lookback=9, skip=1: 0.483120.50lookback=12, skip=0: 0.5030.50lookback=12, skip=1: 0.503
Annualised Sharpe by parameter pair (averaged over top). The outlined cell holds the chosen variant.

Chosen variant lookback=12|skip=1|top=3 ranks 4 of 12. Grid median Sharpe 0.49, best 0.59; 100% of cells positive; neighbour ratio 0.89.

10. Drawdown distribution

Bootstrap distribution of maximum drawdown40%60%80%realised
Maximum drawdown in 1000 stationary-bootstrap resamples. Realised: 48.1%; median 41.5%; 95th percentile 60.8%. 26% of resamples were worse than the realised drawdown.

What is fragile, and how to test the fix

FAIL Selection-adjusted significance vs benchmark (Deflated Sharpe)

Excess Sharpe over buy-and-hold of the traded asset (total return): 0.07 a year over 26.6 years. After allowing for 12 variant(s) tried, the best-of-N noise hurdle is an annualised excess Sharpe of 0.08; the probability that the true excess Sharpe exceeds it is 0.491.

How to test a fix: Declare every variant you tried (including abandoned ones), then re-test on data you have not looked at: a sealed forward period or a later holdout. A DSR that only passes at N = 1 is not evidence.

FAIL Probability of backtest overfitting (CSCV)

Across 12870 combinatorial in-sample/out-of-sample splits of 12 variants, ranked by excess Sharpe, the in-sample winner ranked in the bottom half out of sample 98.8% of the time and trailed the benchmark out of sample 70.8% of the time.

How to test a fix: Shrink the search: fewer free parameters, coarser grids, rules fixed from theory before looking at results. Then re-run CSCV on the smaller family; PBO should fall.

FAIL Multiple-testing haircut vs benchmark (Harvey-Liu)

Excess-return t-statistic 0.37 (single-test p 0.3546). Adjusted for 12 tests (BHY), p = 1.0000; the Bonferroni-haircut excess Sharpe is 0.00.

How to test a fix: A longer sample raises the t-statistic without new selection: extend the test to earlier or later data the rule was not designed on, keeping the rule frozen.

FAIL Beats the benchmark after data snooping (Hansen SPA)

Benchmark: buy-and-hold of the traded asset (total return). Testing whether the best of 12 variant(s) beats it in mean return, with a stationary bootstrap (1000 draws, mean block 19 periods): Hansen SPA p = 0.696, White Reality Check p = 0.670.

How to test a fix: If the rule's value is lower risk rather than higher return, state that claim before testing it (e.g. drawdown or volatility against the benchmark) and test it on unseen data.

FAIL Holdout degradation (excess Sharpe)

In-sample (before 2018-01-02) annualised excess Sharpe 0.16; holdout -0.12 over 2194 periods. Consistency p-value 0.422 (low values mean the holdout is unlikely to share the in-sample excess Sharpe).

How to test a fix: Lock a fresh holdout before any further tuning, or run a sealed forward trial with preregistered pass criteria; only data you have not seen can confirm a fix.

CAUTION Period and regime stability (vs benchmark)

The strategy's compounded return beat the benchmark's in 14 of 27 calendar years. Annualised excess Sharpe by volatility tercile: low vol 0.08, mid vol 0.25, high vol -0.03.

How to test a fix: Check whether the edge over the benchmark is concentrated in a few years or one regime; if so, state that regime as part of the hypothesis and test it on a period the rule has not seen.

These are suggestions for further statistical testing, not suggestions to trade. Any change to the rules creates a new variant: count it in N and confirm it on data it was not designed on.

What held up

Methods appendix

Sharpe ratio and its standard error

Per-period mean over standard deviation of the audited return series (risk-free rate taken as zero), annualised by the square root of periods per year. The standard error uses the IID-normal formula of Lo (2002) and the non-normal correction of Mertens (2002), which widens the error for negative skew and fat tails. Lo's autocorrelation-adjusted annualisation is reported alongside. [1], [2]

Probabilistic Sharpe ratio (PSR) and minimum track record

The probability that the true Sharpe ratio exceeds a benchmark (here zero), given the sample length, skewness and kurtosis; and the minimum sample length for 95% confidence. [3]

Deflated Sharpe ratio (DSR)

The PSR measured against the Sharpe ratio one would expect from the best of N skill-less trials, where N is the number of variants tried and the dispersion of trial Sharpe ratios is estimated from the variant matrix (or, without one, set to the null sampling variance 1/T). It corrects for selection and non-normality at once. [4]

Probability of backtest overfitting (PBO) via CSCV

The variant matrix is cut into 16 time blocks; for each of the 12,870 ways of picking half of them as in-sample, the in-sample best variant is ranked out of sample. PBO is the share of splits where it falls to or below the median. It assumes blocks long enough to preserve serial dependence and a variant set that represents the real search. [5]

Multiple-testing haircut

The Sharpe ratio is turned into a t-statistic and p-value, the p-value is adjusted for the number of tests (Bonferroni always; Holm and BHY when every variant's returns are supplied), and the adjusted p-value is mapped back to a haircut Sharpe ratio. [6], [7]

White's Reality Check and Hansen's SPA test

Tests whether the best of all variants beats the benchmark in mean return once the search over variants is accounted for. Uses the stationary bootstrap with mean block length T^(1/3) (at least 5) to keep short-range dependence. SPA studentises and recentres, so poor variants do not dilute the power. [8], [9], [10]

Holdout degradation

In-sample versus holdout Sharpe ratio at a fixed split date (by default the last 30% of the sample). The consistency p-value asks how surprising the holdout Sharpe would be if the in-sample Sharpe were the truth. A holdout only counts if it was not used to design the rule. [11]

Period and regime stability

Returns and Sharpe ratio by calendar year and by tercile of the underlying market's trailing 21-day volatility (lagged one day). Descriptive: the tercile cut points use the full sample. [12]

Parameter-sensitivity surface

Sharpe ratio across the supplied parameter grid. The neighbour ratio compares the chosen cell with its one-step neighbours: a plateau (ratio near 1) is less fragile than an isolated peak. [13]

Transaction-cost sensitivity

Net return = gross return minus turnover times a one-way cost per unit traded, over a grid of costs; the break-even cost is where the mean net return reaches zero. Market impact beyond a flat cost is not modelled. [12]

Drawdown distribution by bootstrap

The maximum drawdown is recomputed on 1,000 stationary-bootstrap resamples of the return series, showing how much deeper (or shallower) the worst loss could plausibly have been with the same return distribution in a different order. [10]

References

  1. Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal 58(4), 36-52.
  2. Mertens, E. (2002). Comments on Variance of the IID Estimator in Lo (2002). Working paper, University of Basel.
  3. Bailey, D. H. & Lopez de Prado, M. (2012). The Sharpe Ratio Efficient Frontier. Journal of Risk 15(2), 3-44.
  4. Bailey, D. H. & Lopez de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management 40(5), 94-107.
  5. Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2016). The Probability of Backtest Overfitting. Journal of Computational Finance 20(4), 39-69.
  6. Harvey, C. R. & Liu, Y. (2015). Backtesting. Journal of Portfolio Management 42(1), 13-28.
  7. Harvey, C. R., Liu, Y. & Zhu, H. (2016). ... and the Cross-Section of Expected Returns. Review of Financial Studies 29(1), 5-68.
  8. White, H. (2000). A Reality Check for Data Snooping. Econometrica 68(5), 1097-1126.
  9. Hansen, P. R. (2005). A Test for Superior Predictive Ability. Journal of Business & Economic Statistics 23(4), 365-380.
  10. Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association 89(428), 1303-1313.
  11. Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the AMS 61(5), 458-471.
  12. Lopez de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.
  13. Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies, 2nd ed. Wiley.

Data and hash appendix

ItemSHA-256 / value
packageholdout-audit 0.3.0
yahoo_xlb.csv722cb8bb69953816b2323071f7759a6f914a686d1b9e7da34b9e24e51cb80d9a
yahoo_xle.csv6ccc223ccaf10ad125da1781752055f0470ff2705416865ec8e88b06e9ebcb12
yahoo_xlf.csveb9cb569ccb0ceda593be91e3ecb045cc0a008e654aa26281a778f13367f433d
yahoo_xli.csv3016cd9968910b3287f73924628fa9d566478a0363fb1d9ddb5a05d975d13e86
yahoo_xlk.csvdd9d8889f4b5c7b0072326a4f003a6018a1124c67398d317bf47d9a7badabfa5
yahoo_xlp.csvdbb4c43329c9c5826195d0e467123bbba723075db63ff09c3920f66611bb3412
yahoo_xlu.csv79fa525b66aa5292c217a98d242faa1a97801b0da6a261f2c5b76b731ea83ce6
yahoo_xlv.csv6debedf2e83df2875d445167466c77d21c3dec02787611a33cbd50a195b2229f
yahoo_xly.csvc81319403d30a30bb6c1fa9db9ffabb6861cf3d5246813a47c3908b7b479f83e
yahoo_spy.csv8f10432b608927a6a88b67e6e98d72153ba70eb0254c2b5dce1fe3548fb3ceec
audited_series_sha256610578f414c742c1f5483afe15b0e2c318c24050687c09c7f6b2c3215969a739
objective_declaration_sha256ff8cace63cad05977286a8a9a33dbbc88071c2f12a1d4a0d78f87f8d33f98250
preregistration_sha256ff8cace63cad05977286a8a9a33dbbc88071c2f12a1d4a0d78f87f8d33f98250
config_sha25693ebca5673ec5dbbb44e4917789349633b006093e0088e7159095bb2fc6e0792
bootstrap seed20260925
rubric version1.1 (docs/RUBRIC-v1.1.md)
declared objectivebeat_benchmark
preregistration seal35f52f363e7cb128df873803d8e10268977adb2cf06204cb6dfd1b2e57f18b11 (digicert, freetsa; earliest 2026-09-25T20:11:44Z)
objective declaration seal35f52f363e7cb128df873803d8e10268977adb2cf06204cb6dfd1b2e57f18b11 (digicert, freetsa; earliest 2026-09-25T20:11:44Z)
rubric document SHA-2562311b126336bab9f426a84f9c9cd9cb937e346b403a62f60e3c96ed8b5868e46
rubric seal record SHA-256ac882c422df5154ec7e6e928904c9099d6b286dc9e5369fece4b85df57044138
rubric sealed by digicert (RFC 3161)2026-09-25T14:20:10Z
rubric sealed by freetsa (RFC 3161)2026-09-25T14:20:11Z

Anyone with the same inputs and package version can re-run the audit and must obtain the same audited-series hash and the same statistics.

Important. This report is a statistical analysis of historical data supplied by or selected for the client. It describes how fragile past results are under standard tests; it says nothing reliable about future results. It is not investment advice, not a recommendation to buy, sell or hold any security or other instrument, and not an offer of any service regulated as investment or trading advice. Past performance, simulated or real, does not guarantee future results. No warranty is given that the data, code or conclusions are free of error. Backtested and simulated results are hypothetical and have inherent limitations.