HoldoutLabs
Scientific audits for trading strategies
Strategy Fragility Audit
Report 38d40a2be940 · 2026-09-25 20:42 UTC
holdout-audit 0.4.0 · rubric v1.2

Batch 3: Time-series momentum (Moskowitz, Ooi & Pedersen)

Time-series momentum (Moskowitz, Ooi & Pedersen) on SPY, EFA, EEM, IEF, TLT. Claim tested (declared before the run): A diversified time-series momentum portfolio earns a high Sharpe ratio and performs well in extreme markets; here judged on Sharpe improvement over holding the same assets. Benchmark: equal-weight buy-and-hold of the same five ETFs, rebalanced monthly. Part of Indicator Audit Batch 3, 'Strongest published evidence' (6 audits, sealed, Holm across the batch).

Executive summary

Fragility grade
F
3/16 points (19%)

Very fragile. The backtest shows no robust edge over its benchmark once selection, holdout and stability are accounted for. The grade measures how fragile the historical evidence is. It is not a forecast and not a recommendation.

Declared objective: improve Sharpe vs benchmark, declared by the client on 2026-09-25T20:39:29Z (declaration SHA-256 35fd8b3302c094fa…).

Benchmark: buy-and-hold of the traded asset. This is the rubric v1.2 default. The Sharpe-ratio improvement was not robust, so the grade is at most C.

Sharpe: strategy vs benchmark
0.53 vs 0.65
Sharpe improvement (5th pct)
-0.13 [-0.52]
Deflated dSR (N=8)
0.045
PBO (CSCV, dSR)
0.297
Holdout dSR (in → out)
-0.02 → -0.27
Excess Sharpe vs benchmark (95% CI)
0.22 [-0.20, 0.63]
Excess return a year
5.9%
Absolute Sharpe
0.53
Deflated Sharpe (N=8)
0.322
PBO (CSCV)
0.008
Holdout excess Sharpe (in → out)
0.31 → 0.06
Max drawdown
47.6%
Buy-and-hold Sharpe, same period
0.65

The second row is context on the beat-the-benchmark view; it is not graded under the declared objective.

Checks that failed: Selection-adjusted Sharpe improvement (deflated dSR); Multiple-testing adjusted significance of the Sharpe improvement; Primary: Sharpe-ratio improvement vs benchmark; Holdout: Sharpe improvement; Stability of the Sharpe improvement; Parameter plateau (dSR).

Scorecard

CheckResultValueRule
Selection-adjusted Sharpe improvement (deflated dSR)FAILD 0.045 (N = 8)PASS deflated dSR >= 0.95; CAUTION >= 0.80; else FAIL (DSR construction on the Sharpe-ratio improvement)
Probability of backtest overfitting (CSCV, dSR)CAUTIONPBO 0.297 (12870 splits)PASS PBO <= 0.20; CAUTION <= 0.50, or PBO > 0.50 with P(OOS dSR < 0) <= 0.10; else FAIL (variants ranked by dSR)
Multiple-testing adjusted significance of the Sharpe improvementFAILadj. p 1.0000 (BHY)PASS one-sided adjusted p <= 0.05; CAUTION <= 0.10; else FAIL (bootstrap z of dSR)
Primary: Sharpe-ratio improvement vs benchmarkFAILdSR -0.13 (5th pct -0.52)PASS 5th-percentile bootstrap dSR > 0; CAUTION point dSR > 0; else FAIL
Holdout: Sharpe improvementFAILdSR -0.02 -> -0.27PASS holdout and in-sample dSR > 0 and holdout >= 50% of in-sample; CAUTION holdout dSR > 0; else FAIL
Stability of the Sharpe improvementFAIL30% of years higherPASS >= 60% of years with a higher Sharpe than the benchmark and dSR > 0 in every volatility tercile; CAUTION >= 50%; else FAIL
Transaction-cost headroom (absolute)PASSbreak-even 144.5 bpsPASS break-even cost >= 20 bps per unit traded; CAUTION >= 5 bps; else FAIL (absolute)
Parameter plateau (dSR)FAILneighbours n/a% of chosenPASS neighbour dSR >= 70% of chosen dSR; CAUTION >= 40%; else FAIL, or FAIL if chosen dSR <= 0

How the grade is set (Holdout Labs Fragility Rubric v1.2, sealed before this report was produced). PASS = 2 points, CAUTION = 1, FAIL = 0; N/A checks are excluded. Share of available points: A ≥ 85%, B ≥ 70%, C ≥ 55%, D ≥ 40%, otherwise F. Hard caps: Objective cap (beat the benchmark): if the strategy does not beat its benchmark (SPA check FAIL, or annualised mean excess return <= 0), the grade is at most C. Objective cap (reduce drawdown): if the drawdown reduction is not robust (drawdown check FAIL) or the declared return-cost tolerance is breached (tolerance check FAIL), the grade is at most C. Objective cap (improve Sharpe): if the Sharpe-ratio improvement is not robust (Sharpe check FAIL), the grade is at most C. Selection cap: if the deflated check FAILS, the grade is at most C. Overfitting cap: if the PBO check FAILS, the grade is at most D. Disclosure cap: if the number of variants tried was not declared and no variant matrix was supplied, the grade is at most B.

What was audited

Source: Moskowitz, T. J., Ooi, Y. H. & Pedersen, L. H. (2012). Time Series Momentum. Journal of Financial Economics 104(2), 228-250. Rule: At each month end, long each of SPY, EFA, EEM, IEF, TLT with a positive trailing 12-month return and short each with a negative one; each position sized to 40% ex-ante volatility (EWMA, centre of mass 60 days) and averaged across the five (commodities and currencies from the paper are excluded). Audited variant: lookback=12|scaling=vol40. Grid: 8 variants.
Sample2004-05-03 to 2026-09-24 (5635 periods, 252 per year)
Variants tried (N)8 · variant matrix of 8 columns supplied
Base-case cost5 bps per unit traded (one way)
Holdout split2018-01-02
Compound annual return, strategy (net)11.0%
Buy-and-hold of the underlying, same periodcompound annual return 7.5%; annualised Sharpe 0.65; max drawdown 33.6%

Results in detail

Growth of 1 (log scale)holdout12510200420062008201020122014201620182020202220242026StrategyBuy and hold
Growth of 1 on a log scale, net of base-case costs, with buy-and-hold of the underlying in grey. The shaded area is the holdout period.

1. Sharpe ratio and its uncertainty

Rubric checks use the excess series (strategy minus buy-and-hold of the traded asset); the absolute series is shown for reference.

Excess over benchmarkAbsolute
Annualised Sharpe (sqrt-time scaling)0.2180.528
Annualised Sharpe (Lo 2002 autocorrelation-adjusted)0.2630.639
Standard error, annualised (non-normal)0.2120.213
Standard error per period: IID-normal / non-normal0.0133 / 0.01340.0133 / 0.0134
Skewness / kurtosis-0.34 / 20.14-0.25 / 8.21
Probabilistic Sharpe ratio vs 00.84770.9935
Minimum track record for 95% confidence57.4 years9.8 years
Annualised mean / volatility5.9% / 27.3%13.9% / 26.4%

2. The variant-count effect (Deflated Sharpe)

Deflated Sharpe ratio against number of variants tried00.20.40.60.811: 0.84812: 0.69025: 0.42458: 0.322810: 0.2801020: 0.1802050: 0.09750100: 0.060100200: 0.036200500: 0.0185001000: 0.0111000variants tried (log scale)
The same backtest, judged as the best of N tries. The dark dot is the declared N = 8; the dashed line is the 0.95 pass mark. Computed on the excess Sharpe over the benchmark. The more variants tried, the higher the bar.
N triedNoise hurdle (annual excess SR)DSR
10.0000.8477
20.1130.6900
50.2580.4241
80.3160.3216
100.3410.2805
200.4120.1803
500.4930.0972
1000.5480.0597
2000.5990.0361
5000.6610.0183
10000.7050.0108

Dispersion of trial Sharpe ratios used: 0.217 (annualised standard deviation), estimated from the variant matrix.

3. Probability of backtest overfitting (CSCV)

Logit of the in-sample winner's out-of-sample rank-2-1012median
PBO0.008
Splits (blocks)12870 (16)
Variants8
P(IS winner trails benchmark out of sample)0.213
Median OOS excess Sharpe of IS winner (annual)0.13
Degradation slope (OOS on IS)-0.85

Histogram of logit ranks: values left of the dashed line are splits where the in-sample winner finished at or below the out-of-sample median (logit 0).

4. Multiple-testing haircut

t-statistic (excess SR x sqrt(years))1.03
p-value, single test0.15163
Bonferroni p (8 tests)1.00000
Sidak p0.73166
Holm p1.00000
BHY p1.00000
Haircut excess Sharpe (Bonferroni)0.000 (100% haircut)
Haircut excess Sharpe (BHY)0.000

5. Reality Check and SPA

Benchmarkbuy-and-hold of the traded asset
Variants in the test8
Best mean excess return (annualised)5.95%
White Reality Check p0.329
Hansen SPA p (consistent / lower / upper)0.300 / 0.228 / 0.313
Bootstrapstationary, 1000 draws, mean block 18

6. Holdout degradation

PeriodsAnnualised excess Sharpe
In-sample (before 2018-01-02)34410.310
Holdout21940.059
Holdout / in-sample0.19
Consistency p-value0.461

7. Period and regime stability

Return minus benchmark return, by calendar year-25%0%25%50%2004: 27.9%2005: 11.0%2006: -11.9%2007: 34.2%2008: 67.6%2009: -37.4%2010: -12.3%2011: 2.7%2012: -5.4%2013: 22.8%2014: -2.9%2015: 8.0%2016: -22.8%2017: 41.5%2018: -20.9%2019: 3.4%2020: -23.1%2021: -8.5%2022: 48.3%2023: -13.9%2024: 3.4%2025: -0.4%2026: 4.0%200420062008201020122014201620182020202220242026
Strategy compounded return minus benchmark compounded return, by year: ahead in 12 of 23 years (absolute return positive in 17).
Volatility regime (trailing 21-day, lagged)PeriodsAnnualised mean excessAnnualised excess Sharpe
low vol187210.3%0.48
mid vol18716.2%0.30
high vol18711.0%0.03

8. Transaction-cost sensitivity

One-way cost (bps per unit traded)0125102050
Annualised Sharpe0.550.540.540.530.510.470.36
Annualised mean return14.4%14.3%14.2%13.9%13.4%12.4%9.4%

Turnover 10.0 units per year; break-even cost 144.5 bps.

9. Parameter sensitivity

Annualised Sharpe across the parameter gridlookback ↓ / scaling →equalvol4010.02lookback=1, scaling=equal: 0.0190.01lookback=1, scaling=vol40: 0.01230.36lookback=3, scaling=equal: 0.3550.42lookback=3, scaling=vol40: 0.42360.24lookback=6, scaling=equal: 0.2410.39lookback=6, scaling=vol40: 0.393120.35lookback=12, scaling=equal: 0.3470.53lookback=12, scaling=vol40: 0.528
Annualised Sharpe by parameter pair. The outlined cell holds the chosen variant.

Chosen variant lookback=12|scaling=vol40 ranks 1 of 8. Grid median Sharpe 0.35, best 0.53; 100% of cells positive; neighbour ratio 0.70.

10. Drawdown distribution

Bootstrap distribution of maximum drawdown30%40%50%60%70%80%realised
Maximum drawdown in 1000 stationary-bootstrap resamples. Realised: 47.6%; median 50.3%; 95th percentile 70.7%. 58% of resamples were worse than the realised drawdown.

What is fragile, and how to test the fix

FAIL Selection-adjusted Sharpe improvement (deflated dSR)

Annualised Sharpe 0.53 against 0.65 for the benchmark: an improvement of -0.13 (bootstrap SE 0.24). The best-of-8 noise hurdle is +0.27; the probability that the true improvement exceeds it is 0.045.

How to test a fix: Declare every variant tried and confirm the improvement on data the rule has not seen.

CAUTION Probability of backtest overfitting (CSCV, dSR)

Ranking 8 variants by Sharpe improvement, the in-sample winner ranked in the bottom half out of sample 29.7% of the time and had a lower Sharpe than the benchmark out of sample 88.2% of the time.

How to test a fix: Shrink the search and re-run CSCV on the smaller family.

FAIL Multiple-testing adjusted significance of the Sharpe improvement

One-sided p that the Sharpe improvement is zero or worse: 0.7040; adjusted for 8 tests (BHY): 1.0000.

How to test a fix: A longer sample, with the rule frozen, sharpens this test without new selection.

FAIL Primary: Sharpe-ratio improvement vs benchmark

Annualised Sharpe 0.53 vs 0.65; improvement -0.13, paired stationary-bootstrap 5th percentile -0.52 (1000 draws).

How to test a fix: Test the improvement on a period with different market conditions, or in a sealed forward trial.

FAIL Holdout: Sharpe improvement

Sharpe improvement -0.02 before 2018-01-02 and -0.27 after.

How to test a fix: Lock a fresh holdout or run a sealed forward trial.

FAIL Stability of the Sharpe improvement

The strategy's Sharpe was higher than the benchmark's in 7 of 23 years. dSR by volatility tercile: low vol -0.11, mid vol -0.17, high vol -0.26.

How to test a fix: State the regime in which the improvement is expected and test it where the rule was not designed.

FAIL Parameter plateau (dSR)

The chosen variant ranks 1 of 8 by Sharpe improvement; its one-step neighbours average n/a% of its improvement.

How to test a fix: Prefer the centre of a plateau.

These are suggestions for further statistical testing, not suggestions to trade. Any change to the rules creates a new variant: count it in N and confirm it on data it was not designed on.

What held up

Methods appendix

Sharpe ratio and its standard error

Per-period mean over standard deviation of the audited return series (risk-free rate taken as zero), annualised by the square root of periods per year. The standard error uses the IID-normal formula of Lo (2002) and the non-normal correction of Mertens (2002), which widens the error for negative skew and fat tails. Lo's autocorrelation-adjusted annualisation is reported alongside. [1], [2]

Probabilistic Sharpe ratio (PSR) and minimum track record

The probability that the true Sharpe ratio exceeds a benchmark (here zero), given the sample length, skewness and kurtosis; and the minimum sample length for 95% confidence. [3]

Deflated Sharpe ratio (DSR)

The PSR measured against the Sharpe ratio one would expect from the best of N skill-less trials, where N is the number of variants tried and the dispersion of trial Sharpe ratios is estimated from the variant matrix (or, without one, set to the null sampling variance 1/T). It corrects for selection and non-normality at once. [4]

Probability of backtest overfitting (PBO) via CSCV

The variant matrix is cut into 16 time blocks; for each of the 12,870 ways of picking half of them as in-sample, the in-sample best variant is ranked out of sample. PBO is the share of splits where it falls to or below the median. It assumes blocks long enough to preserve serial dependence and a variant set that represents the real search. [5]

Multiple-testing haircut

The Sharpe ratio is turned into a t-statistic and p-value, the p-value is adjusted for the number of tests (Bonferroni always; Holm and BHY when every variant's returns are supplied), and the adjusted p-value is mapped back to a haircut Sharpe ratio. [6], [7]

White's Reality Check and Hansen's SPA test

Tests whether the best of all variants beats the benchmark in mean return once the search over variants is accounted for. Uses the stationary bootstrap with mean block length T^(1/3) (at least 5) to keep short-range dependence. SPA studentises and recentres, so poor variants do not dilute the power. [8], [9], [10]

Holdout degradation

In-sample versus holdout Sharpe ratio at a fixed split date (by default the last 30% of the sample). The consistency p-value asks how surprising the holdout Sharpe would be if the in-sample Sharpe were the truth. A holdout only counts if it was not used to design the rule. [11]

Period and regime stability

Returns and Sharpe ratio by calendar year and by tercile of the underlying market's trailing 21-day volatility (lagged one day). Descriptive: the tercile cut points use the full sample. [12]

Parameter-sensitivity surface

Sharpe ratio across the supplied parameter grid. The neighbour ratio compares the chosen cell with its one-step neighbours: a plateau (ratio near 1) is less fragile than an isolated peak. [13]

Transaction-cost sensitivity

Net return = gross return minus turnover times a one-way cost per unit traded, over a grid of costs; the break-even cost is where the mean net return reaches zero. Market impact beyond a flat cost is not modelled. [12]

Drawdown distribution by bootstrap

The maximum drawdown is recomputed on 1,000 stationary-bootstrap resamples of the return series, showing how much deeper (or shallower) the worst loss could plausibly have been with the same return distribution in a different order. [10]

References

  1. Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal 58(4), 36-52.
  2. Mertens, E. (2002). Comments on Variance of the IID Estimator in Lo (2002). Working paper, University of Basel.
  3. Bailey, D. H. & Lopez de Prado, M. (2012). The Sharpe Ratio Efficient Frontier. Journal of Risk 15(2), 3-44.
  4. Bailey, D. H. & Lopez de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management 40(5), 94-107.
  5. Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2016). The Probability of Backtest Overfitting. Journal of Computational Finance 20(4), 39-69.
  6. Harvey, C. R. & Liu, Y. (2015). Backtesting. Journal of Portfolio Management 42(1), 13-28.
  7. Harvey, C. R., Liu, Y. & Zhu, H. (2016). ... and the Cross-Section of Expected Returns. Review of Financial Studies 29(1), 5-68.
  8. White, H. (2000). A Reality Check for Data Snooping. Econometrica 68(5), 1097-1126.
  9. Hansen, P. R. (2005). A Test for Superior Predictive Ability. Journal of Business & Economic Statistics 23(4), 365-380.
  10. Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association 89(428), 1303-1313.
  11. Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the AMS 61(5), 458-471.
  12. Lopez de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.
  13. Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies, 2nd ed. Wiley.

Data and hash appendix

ItemSHA-256 / value
packageholdout-audit 0.4.0
yahoo_spy.csv8f10432b608927a6a88b67e6e98d72153ba70eb0254c2b5dce1fe3548fb3ceec
yahoo_efa.csv259e2b5afa686d8fd0605dc472b7e84f8c94a5f9b7cc95e54c8433e9f001e75f
yahoo_eem.csvef60cb9acb5fcb2801a3e1576cf3d4d74eeae1d2fc6e21efeb08b44e5c1118b8
yahoo_ief.csvc80eb1f9e013d37ab945b6c150cfbfee0dc3c6eec9a1c6463587fc58d771ecce
yahoo_tlt.csv9e3ba9c9ac9caa8aa53096e89de089bafb91734efe725d1ef22502764afced84
audited_series_sha2564c73589d927d5f4feaa03ce0d109c924b293caf0b9761d31282673cdc8827840
objective_declaration_sha25635fd8b3302c094fad51dda9235b8ea45af2796687e9861016cf10b05d2f2a144
preregistration_sha2562142d17c29db1f2b0080e225cfbcd15ba67582b4122adcc6c78b0aad1325d690
config_sha25638d40a2be9405cc868a0582dead307e10b869f7300505106d8d3133fd3654df7
bootstrap seed20260925
rubric version1.2 (docs/RUBRIC-v1.2.md)
declared objectiveimprove_sharpe
preregistration seal90ff232e75d552219386a15d1bc0990a86b40a925d2a7dc1405e7be67995e009 (digicert, freetsa; earliest 2026-09-25T20:39:31Z)
objective declaration seala899ba2115046c75752bac2d09a1fe594d1618f62ef4fe9536b2a0d79c88f555 (digicert, freetsa; earliest 2026-09-25T20:39:46Z)
rubric document SHA-25681bc60b5bf4520fdfb91c43b37883c30a9a2061903442d1e8a8b05f3d37a71fc
rubric seal record SHA-256142cbfb45314faab8a578b62a17532ac323b0f619f74ab008244163d682db4dd
rubric sealed by digicert (RFC 3161)2026-09-25T20:33:45Z
rubric sealed by freetsa (RFC 3161)2026-09-25T20:33:46Z

Anyone with the same inputs and package version can re-run the audit and must obtain the same audited-series hash and the same statistics.

Important. This report is a statistical analysis of historical data supplied by or selected for the client. It describes how fragile past results are under standard tests; it says nothing reliable about future results. It is not investment advice, not a recommendation to buy, sell or hold any security or other instrument, and not an offer of any service regulated as investment or trading advice. Past performance, simulated or real, does not guarantee future results. No warranty is given that the data, code or conclusions are free of error. Backtested and simulated results are hypothetical and have inherent limitations.