Batch 3: Faber GTAA timing on unseen markets (EEM, VNQ, TLT)
Faber GTAA timing on unseen markets (EEM, VNQ, TLT) on EEM, VNQ, TLT. Claim tested (declared before the run): Timing each asset class with a 10-month SMA gives equity-like returns with bond-like volatility and much smaller drawdowns than buy-and-hold. Benchmark: equal-weight buy-and-hold of the same three ETFs, rebalanced monthly. Part of Indicator Audit Batch 3, 'Strongest published evidence' (6 audits, sealed, Holm across the batch).
Executive summary
Mixed. Part of the historical edge looks explainable by selection, period or cost effects. The grade measures how fragile the historical evidence is. It is not a forecast and not a recommendation.
Declared objective: reduce drawdown at acceptable cost, declared by the client on 2026-09-25T20:39:29Z (declaration SHA-256 e4fb88120aad0e08…). Declared return-cost tolerance: 1.0% a year.
Benchmark: buy-and-hold of the traded asset. This is the rubric v1.2 default. The drawdown reduction was not robust or exceeded the declared cost, so the grade is at most C.
The second row is context on the beat-the-benchmark view; it is not graded under the declared objective.
Checks that failed: Primary: return cost within the declared tolerance.
Scorecard
| Check | Result | Value | Rule |
|---|---|---|---|
| Selection-adjusted tail-loss reduction (deflated dCVaR) | PASS | D 1.000 (N = 8) | PASS deflated dCVaR >= 0.95; CAUTION >= 0.80; else FAIL (DSR construction on the CVaR reduction) |
| Probability of backtest overfitting (CSCV, dCVaR) | CAUTION | PBO 0.534 (12870 splits) | PASS PBO <= 0.20; CAUTION <= 0.50, or PBO > 0.50 with P(OOS dCVaR < 0) <= 0.10; else FAIL (variants ranked by dCVaR) |
| Multiple-testing adjusted significance of the tail-loss reduction | PASS | adj. p 0.0001 (BHY) | PASS one-sided adjusted p <= 0.05; CAUTION <= 0.10; else FAIL (bootstrap z of dCVaR) |
| Primary: drawdown and tail-loss reduction vs benchmark | PASS | dMDD 29.4% (5th pct 3.5%) | PASS 5th-percentile bootstrap dMDD > 0 and dCVaR > 0; CAUTION both point estimates > 0; else FAIL |
| Primary: return cost within the declared tolerance | FAIL | shortfall 3.2%/yr vs 1.0% | PASS 95th-percentile bootstrap return shortfall <= declared tolerance; CAUTION point shortfall <= tolerance; else FAIL |
| Holdout: tail-loss and drawdown reduction | PASS | dCVaR 129.1 bps -> 81.7 bps | PASS holdout and in-sample dCVaR > 0, holdout >= 50% of in-sample, and holdout dMDD > 0; CAUTION holdout dCVaR > 0; else FAIL |
| Stability of the drawdown reduction | PASS | 68% of years shallower | PASS >= 60% of years with a shallower drawdown than the benchmark and dCVaR > 0 in every volatility tercile; CAUTION >= 50%; else FAIL |
| Transaction-cost headroom (absolute) | PASS | break-even 224.3 bps | PASS break-even cost >= 20 bps per unit traded; CAUTION >= 5 bps; else FAIL (absolute) |
| Parameter plateau (dCVaR) | PASS | neighbours 95% of chosen | PASS neighbour dCVaR >= 70% of chosen dCVaR; CAUTION >= 40%; else FAIL, or FAIL if chosen dCVaR <= 0 |
How the grade is set (Holdout Labs Fragility Rubric v1.2, sealed before this report was produced). PASS = 2 points, CAUTION = 1, FAIL = 0; N/A checks are excluded. Share of available points: A ≥ 85%, B ≥ 70%, C ≥ 55%, D ≥ 40%, otherwise F. Hard caps: Objective cap (beat the benchmark): if the strategy does not beat its benchmark (SPA check FAIL, or annualised mean excess return <= 0), the grade is at most C. Objective cap (reduce drawdown): if the drawdown reduction is not robust (drawdown check FAIL) or the declared return-cost tolerance is breached (tolerance check FAIL), the grade is at most C. Objective cap (improve Sharpe): if the Sharpe-ratio improvement is not robust (Sharpe check FAIL), the grade is at most C. Selection cap: if the deflated check FAILS, the grade is at most C. Overfitting cap: if the PBO check FAILS, the grade is at most D. Disclosure cap: if the number of variants tried was not declared and no variant matrix was supplied, the grade is at most B.
What was audited
| Sample | 2005-10-03 to 2026-09-24 (5277 periods, 252 per year) |
| Variants tried (N) | 8 · variant matrix of 8 columns supplied |
| Base-case cost | 5 bps per unit traded (one way) |
| Holdout split | 2018-01-02 |
| Compound annual return, strategy (net) | 4.5% |
| Buy-and-hold of the underlying, same period | compound annual return 6.8%; annualised Sharpe 0.48; max drawdown 43.9% |
Results in detail
1. Sharpe ratio and its uncertainty
Rubric checks use the excess series (strategy minus buy-and-hold of the traded asset); the absolute series is shown for reference.
| Excess over benchmark | Absolute | |
|---|---|---|
| Annualised Sharpe (sqrt-time scaling) | -0.223 | 0.536 |
| Annualised Sharpe (Lo 2002 autocorrelation-adjusted) | -0.269 | 0.674 |
| Standard error, annualised (non-normal) | 0.218 | 0.219 |
| Standard error per period: IID-normal / non-normal | 0.0138 / 0.0137 | 0.0138 / 0.0138 |
| Skewness / kurtosis | -0.32 / 35.10 | -0.08 / 7.77 |
| Probabilistic Sharpe ratio vs 0 | 0.1532 | 0.9928 |
| Minimum track record for 95% confidence | ∞ years | 9.5 years |
| Annualised mean / volatility | -3.2% / 14.5% | 4.8% / 8.9% |
2. The variant-count effect (Deflated Sharpe)
| N tried | Noise hurdle (annual excess SR) | DSR |
|---|---|---|
| 1 | 0.000 | 0.1532 |
| 2 | 0.010 | 0.1427 |
| 5 | 0.023 | 0.1299 |
| 8 | 0.028 | 0.1250 |
| 10 | 0.030 | 0.1229 |
| 20 | 0.036 | 0.1172 |
| 50 | 0.043 | 0.1109 |
| 100 | 0.048 | 0.1067 |
| 200 | 0.053 | 0.1030 |
| 500 | 0.058 | 0.0986 |
| 1000 | 0.062 | 0.0956 |
Dispersion of trial Sharpe ratios used: 0.019 (annualised standard deviation), estimated from the variant matrix.
3. Probability of backtest overfitting (CSCV)
| PBO | 0.917 |
| Splits (blocks) | 12870 (16) |
| Variants | 8 |
| P(IS winner trails benchmark out of sample) | 0.987 |
| Median OOS excess Sharpe of IS winner (annual) | -0.27 |
| Degradation slope (OOS on IS) | -0.90 |
Histogram of logit ranks: values left of the dashed line are splits where the in-sample winner finished at or below the out-of-sample median (logit 0).
4. Multiple-testing haircut
| t-statistic (excess SR x sqrt(years)) | -1.02 |
| p-value, single test | 0.84647 |
| Bonferroni p (8 tests) | 1.00000 |
| Sidak p | 1.00000 |
| Holm p | 1.00000 |
| BHY p | 1.00000 |
| Haircut excess Sharpe (Bonferroni) | -0.000 (100% haircut) |
| Haircut excess Sharpe (BHY) | -0.000 |
5. Reality Check and SPA
| Benchmark | buy-and-hold of the traded asset |
| Variants in the test | 8 |
| Best mean excess return (annualised) | -2.75% |
| White Reality Check p | 0.947 |
| Hansen SPA p (consistent / lower / upper) | 1.000 / 1.000 / 1.000 |
| Bootstrap | stationary, 1000 draws, mean block 17 |
6. Holdout degradation
| Periods | Annualised excess Sharpe | |
|---|---|---|
| In-sample (before 2018-01-02) | 3083 | -0.235 |
| Holdout | 2194 | -0.210 |
| Holdout / in-sample | n/a | |
| Consistency p-value | 0.941 |
7. Period and regime stability
| Volatility regime (trailing 21-day, lagged) | Periods | Annualised mean excess | Annualised excess Sharpe |
|---|---|---|---|
| low vol | 1752 | -1.8% | -0.39 |
| mid vol | 1752 | -3.6% | -0.55 |
| high vol | 1752 | -4.4% | -0.18 |
8. Transaction-cost sensitivity
| One-way cost (bps per unit traded) | 0 | 1 | 2 | 5 | 10 | 20 | 50 |
|---|---|---|---|---|---|---|---|
| Annualised Sharpe | 0.55 | 0.55 | 0.54 | 0.54 | 0.52 | 0.50 | 0.43 |
| Annualised mean return | 4.9% | 4.9% | 4.8% | 4.8% | 4.7% | 4.5% | 3.8% |
Turnover 2.2 units per year; break-even cost 224.3 bps.
9. Parameter sensitivity
Chosen variant months=10|basis=price ranks 5 of 8. Grid median Sharpe 0.54, best 0.56; 100% of cells positive; neighbour ratio 1.01.
10. Drawdown distribution
11. Declared objective: drawdown reduction at acceptable cost
| Strategy | Benchmark | Reduction | Bootstrap bound | |
|---|---|---|---|---|
| Maximum drawdown | 14.5% | 43.9% | 29.4% | 5th pct 3.5% |
| Daily CVaR95 (bps) | 135.4 | 245.8 | 110.4 | 5th pct 69.7 |
| Return shortfall a year (tolerance 1.0%) | 3.2% | 95th pct 7.1% |
Deflated dCVaR 1.000 against a best-of-8 hurdle of 7.0 bps; adjusted one-sided p <0.0001. Holdout: dCVaR 129.1 bps before 2018-01-02, 81.7 bps after; holdout drawdown reduction 17.5%. Years with a shallower drawdown than the benchmark: 15 of 22.
What is fragile, and how to test the fix
CAUTION Probability of backtest overfitting (CSCV, dCVaR)
Ranking 8 variants by tail-loss reduction, the in-sample winner ranked in the bottom half out of sample 53.4% of the time and had worse tails than the benchmark out of sample 0.0% of the time.
How to test a fix: Shrink the search and re-run CSCV on the smaller family.
FAIL Primary: return cost within the declared tolerance
The strategy gave up 3.2% a year of arithmetic return against the benchmark (bootstrap 95th percentile 7.1%). The client declared a tolerance of 1.0% a year on 2026-09-25T20:39:29Z.
How to test a fix: The tolerance was fixed in advance and cannot be revised for this audit; a changed rule is a new variant and needs a new declaration.
These are suggestions for further statistical testing, not suggestions to trade. Any change to the rules creates a new variant: count it in N and confirm it on data it was not designed on.
What held up
- Selection-adjusted tail-loss reduction (deflated dCVaR): Daily CVaR95 falls from 245.8 bps (benchmark) to 135.4 bps (strategy): a reduction of 110.4 bps (bootstrap SE 26.1 bps). The best-of-8 noise hurdle is 7.0 bps; the probability that the true reduction exceeds it is 1.000.
- Multiple-testing adjusted significance of the tail-loss reduction: One-sided p that the tail-loss reduction is zero or worse: 0.0000 for the audited strategy; adjusted for 8 tests (BHY): 0.0001.
- Primary: drawdown and tail-loss reduction vs benchmark: Maximum drawdown 14.5% against 43.9% for the benchmark (reduction 29.4%; bootstrap 5th percentile 3.5%). Daily CVaR95 reduction 110.4 bps (5th percentile 69.7 bps). Paired stationary bootstrap, 1000 draws.
- Holdout: tail-loss and drawdown reduction: Tail-loss reduction 129.1 bps before 2018-01-02 and 81.7 bps after; holdout drawdown reduction 17.5%.
- Stability of the drawdown reduction: The strategy's within-year drawdown was shallower than the benchmark's in 15 of 22 years. dCVaR by volatility tercile: low vol 16.6 bps, mid vol 34.8 bps, high vol 214.9 bps.
- Transaction-cost headroom (absolute): The gross mean return falls to zero at 224.3 bps per unit traded.
- Parameter plateau (dCVaR): The chosen variant ranks 1 of 8 by tail-loss reduction; its one-step neighbours average 95% of its reduction.
Methods appendix
Sharpe ratio and its standard error
Per-period mean over standard deviation of the audited return series (risk-free rate taken as zero), annualised by the square root of periods per year. The standard error uses the IID-normal formula of Lo (2002) and the non-normal correction of Mertens (2002), which widens the error for negative skew and fat tails. Lo's autocorrelation-adjusted annualisation is reported alongside. [1], [2]
Probabilistic Sharpe ratio (PSR) and minimum track record
The probability that the true Sharpe ratio exceeds a benchmark (here zero), given the sample length, skewness and kurtosis; and the minimum sample length for 95% confidence. [3]
Deflated Sharpe ratio (DSR)
The PSR measured against the Sharpe ratio one would expect from the best of N skill-less trials, where N is the number of variants tried and the dispersion of trial Sharpe ratios is estimated from the variant matrix (or, without one, set to the null sampling variance 1/T). It corrects for selection and non-normality at once. [4]
Probability of backtest overfitting (PBO) via CSCV
The variant matrix is cut into 16 time blocks; for each of the 12,870 ways of picking half of them as in-sample, the in-sample best variant is ranked out of sample. PBO is the share of splits where it falls to or below the median. It assumes blocks long enough to preserve serial dependence and a variant set that represents the real search. [5]
Multiple-testing haircut
The Sharpe ratio is turned into a t-statistic and p-value, the p-value is adjusted for the number of tests (Bonferroni always; Holm and BHY when every variant's returns are supplied), and the adjusted p-value is mapped back to a haircut Sharpe ratio. [6], [7]
White's Reality Check and Hansen's SPA test
Tests whether the best of all variants beats the benchmark in mean return once the search over variants is accounted for. Uses the stationary bootstrap with mean block length T^(1/3) (at least 5) to keep short-range dependence. SPA studentises and recentres, so poor variants do not dilute the power. [8], [9], [10]
Holdout degradation
In-sample versus holdout Sharpe ratio at a fixed split date (by default the last 30% of the sample). The consistency p-value asks how surprising the holdout Sharpe would be if the in-sample Sharpe were the truth. A holdout only counts if it was not used to design the rule. [11]
Period and regime stability
Returns and Sharpe ratio by calendar year and by tercile of the underlying market's trailing 21-day volatility (lagged one day). Descriptive: the tercile cut points use the full sample. [12]
Parameter-sensitivity surface
Sharpe ratio across the supplied parameter grid. The neighbour ratio compares the chosen cell with its one-step neighbours: a plateau (ratio near 1) is less fragile than an isolated peak. [13]
Transaction-cost sensitivity
Net return = gross return minus turnover times a one-way cost per unit traded, over a grid of costs; the break-even cost is where the mean net return reaches zero. Market impact beyond a flat cost is not modelled. [12]
Drawdown distribution by bootstrap
The maximum drawdown is recomputed on 1,000 stationary-bootstrap resamples of the return series, showing how much deeper (or shallower) the worst loss could plausibly have been with the same return distribution in a different order. [10]
References
- Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal 58(4), 36-52.
- Mertens, E. (2002). Comments on Variance of the IID Estimator in Lo (2002). Working paper, University of Basel.
- Bailey, D. H. & Lopez de Prado, M. (2012). The Sharpe Ratio Efficient Frontier. Journal of Risk 15(2), 3-44.
- Bailey, D. H. & Lopez de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management 40(5), 94-107.
- Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2016). The Probability of Backtest Overfitting. Journal of Computational Finance 20(4), 39-69.
- Harvey, C. R. & Liu, Y. (2015). Backtesting. Journal of Portfolio Management 42(1), 13-28.
- Harvey, C. R., Liu, Y. & Zhu, H. (2016). ... and the Cross-Section of Expected Returns. Review of Financial Studies 29(1), 5-68.
- White, H. (2000). A Reality Check for Data Snooping. Econometrica 68(5), 1097-1126.
- Hansen, P. R. (2005). A Test for Superior Predictive Ability. Journal of Business & Economic Statistics 23(4), 365-380.
- Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association 89(428), 1303-1313.
- Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the AMS 61(5), 458-471.
- Lopez de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.
- Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies, 2nd ed. Wiley.
Data and hash appendix
- Prices from Yahoo Finance's public chart endpoint; we publish derived statistics only, never raw prices; no redistribution. Signals use the split-adjusted close; total returns use the split- and dividend-adjusted close.
- yahoo_eem.csv: 5900 rows ending 2026-09-24, SHA-256 ef60cb9acb5fcb2801a3e1576cf3d4d74eeae1d2fc6e21efeb08b44e5c1118b8 (fetched after the seal)
- yahoo_vnq.csv: 5532 rows ending 2026-09-24, SHA-256 8bc567ef38c35aa59aacd9d40c713590455970c2ac58faff9da458728d8dd3ac (fetched after the seal)
- yahoo_tlt.csv: 6078 rows ending 2026-09-24, SHA-256 9e3ba9c9ac9caa8aa53096e89de089bafb91734efe725d1ef22502764afced84 (fetched after the seal)
- Audit period 2005-10-03 to 2026-09-24; holdout from 2018-01-02; benchmark: equal-weight buy-and-hold of the same three ETFs, rebalanced monthly (all fixed in the sealed batch preregistration).
- Cash earns 0% and borrowing above 100% exposure costs 0%; shorting has no borrow cost (stated in the preregistration).
| Item | SHA-256 / value |
|---|---|
| package | holdout-audit 0.4.0 |
| yahoo_eem.csv | ef60cb9acb5fcb2801a3e1576cf3d4d74eeae1d2fc6e21efeb08b44e5c1118b8 |
| yahoo_vnq.csv | 8bc567ef38c35aa59aacd9d40c713590455970c2ac58faff9da458728d8dd3ac |
| yahoo_tlt.csv | 9e3ba9c9ac9caa8aa53096e89de089bafb91734efe725d1ef22502764afced84 |
| audited_series_sha256 | 27b4db7e08dc651d624d99f214a1eb13aa980f36da7e8f017239cd106b61bdb4 |
| objective_declaration_sha256 | e4fb88120aad0e081f71500717f4f73490f446fb92462f7adc9772e8b062a97f |
| preregistration_sha256 | 2142d17c29db1f2b0080e225cfbcd15ba67582b4122adcc6c78b0aad1325d690 |
| config_sha256 | 36ee2e7779ea3926c514054ab234911b6f840cc5820cf06fbd247342396b3a92 |
| bootstrap seed | 20260925 |
| rubric version | 1.2 (docs/RUBRIC-v1.2.md) |
| declared objective | reduce_drawdown |
| preregistration seal | 90ff232e75d552219386a15d1bc0990a86b40a925d2a7dc1405e7be67995e009 (digicert, freetsa; earliest 2026-09-25T20:39:31Z) |
| objective declaration seal | 5520ca3004e61782747c0696bec84ee5d2d2e876a86534e31735a433b7905127 (digicert, freetsa; earliest 2026-09-25T20:39:50Z) |
| rubric document SHA-256 | 81bc60b5bf4520fdfb91c43b37883c30a9a2061903442d1e8a8b05f3d37a71fc |
| rubric seal record SHA-256 | 142cbfb45314faab8a578b62a17532ac323b0f619f74ab008244163d682db4dd |
| rubric sealed by digicert (RFC 3161) | 2026-09-25T20:33:45Z |
| rubric sealed by freetsa (RFC 3161) | 2026-09-25T20:33:46Z |
Anyone with the same inputs and package version can re-run the audit and must obtain the same audited-series hash and the same statistics.