ONCHAIN-1 A6: Stock-to-flow (PlanB 2019), implied rule, post-publication
Implied trading rule of a published Bitcoin indicator, frozen as first published and tested only after its publication (ONCHAIN-1, preregistered and sealed). Benchmark: buy-and-hold BTC. Objective (a).
Executive summary
Very fragile. The backtest shows no robust edge over its benchmark once selection, holdout and stability are accounted for. The grade measures how fragile the historical evidence is. It is not a forecast and not a recommendation.
Declared objective: beat the benchmark, the default; no objective declaration was supplied.
Benchmark: buy-and-hold of the traded asset. This is the rubric v1.1 default. The strategy did not beat its benchmark under the rubric definition, so the grade is at most C.
Checks that failed: Selection-adjusted significance vs benchmark (Deflated Sharpe); Multiple-testing haircut vs benchmark (Harvey-Liu); Beats the benchmark after data snooping (Hansen SPA); Holdout degradation (excess Sharpe); Period and regime stability (vs benchmark).
Scorecard
| Check | Result | Value | Rule |
|---|---|---|---|
| Selection-adjusted significance vs benchmark (Deflated Sharpe) | FAIL | DSR 0.082 (N = 1) | PASS DSR >= 0.95; CAUTION >= 0.80; else FAIL (on excess returns) |
| Probability of backtest overfitting (CSCV) | N/A | not computed | PASS PBO <= 0.20; CAUTION <= 0.50, or PBO > 0.50 with P(OOS excess SR < 0) <= 0.10; else FAIL |
| Multiple-testing haircut vs benchmark (Harvey-Liu) | FAIL | adj. p 0.9253 (Bonferroni) | PASS one-sided adjusted p <= 0.05; CAUTION <= 0.10; else FAIL (excess-return t-statistic; evidence for the strategy only) |
| Beats the benchmark after data snooping (Hansen SPA) | FAIL | SPA p 1.000; RC p 0.920 | PASS p <= 0.05; CAUTION <= 0.10; else FAIL |
| Holdout degradation (excess Sharpe) | FAIL | excess SR -0.63 -> 0.00 | PASS holdout and in-sample excess SR > 0 and holdout >= 50% of in-sample; CAUTION holdout excess SR > 0; else FAIL |
| Period and regime stability (vs benchmark) | FAIL | 0% of years ahead | PASS >= 60% of years ahead of benchmark and excess SR > 0 in every volatility tercile; CAUTION >= 50%; else FAIL |
| Transaction-cost headroom (absolute) | PASS | break-even 1414.1 bps | PASS break-even cost >= 20 bps per unit traded; CAUTION >= 5 bps; else FAIL (absolute) |
| Parameter plateau | N/A | no grid | PASS neighbour SR >= 70% of chosen SR; CAUTION >= 40%; else FAIL, or FAIL if chosen SR <= 0 (absolute) |
How the grade is set (Holdout Labs Fragility Rubric v1.1, sealed before this report was produced). PASS = 2 points, CAUTION = 1, FAIL = 0; N/A checks are excluded. Share of available points: A ≥ 85%, B ≥ 70%, C ≥ 55%, D ≥ 40%, otherwise F. Hard caps: Objective cap (beat the benchmark): if the strategy does not beat its benchmark (SPA check FAIL, or annualised mean excess return <= 0), the grade is at most C. Objective cap (reduce drawdown): if the drawdown reduction is not robust (drawdown check FAIL) or the declared return-cost tolerance is breached (tolerance check FAIL), the grade is at most C. Selection cap: if the deflated check FAILS, the grade is at most C. Overfitting cap: if the PBO check FAILS, the grade is at most D. Disclosure cap: if the number of variants tried was not declared and no variant matrix was supplied, the grade is at most B.
What was audited
| Sample | 2019-03-23 to 2026-09-24 (2743 periods, 365 per year) |
| Variants tried (N) | 1 |
| Base-case cost | 10 bps per unit traded (one way) |
| Holdout split | 2024-06-24 |
| Compound annual return, strategy (net) | 32.8% |
| Buy-and-hold of the underlying, same period | compound annual return 50.1%; annualised Sharpe 0.97; max drawdown 76.7% |
Results in detail
1. Sharpe ratio and its uncertainty
Rubric checks use the excess series (strategy minus buy-and-hold of the traded asset); the absolute series is shown for reference.
| Excess over benchmark | Absolute | |
|---|---|---|
| Annualised Sharpe (sqrt-time scaling) | -0.526 | 0.814 |
| Annualised Sharpe (Lo 2002 autocorrelation-adjusted) | -0.496 | 0.808 |
| Standard error, annualised (non-normal) | 0.377 | 0.363 |
| Standard error per period: IID-normal / non-normal | 0.0191 / 0.0197 | 0.0191 / 0.0190 |
| Skewness / kurtosis | 1.90 / 82.88 | 0.31 / 9.28 |
| Probabilistic Sharpe ratio vs 0 | 0.0816 | 0.9875 |
| Minimum track record for 95% confidence | ∞ years | 4.1 years |
| Annualised mean / volatility | -19.2% / 36.5% | 41.1% / 50.5% |
2. The variant-count effect (Deflated Sharpe)
| N tried | Noise hurdle (annual excess SR) | DSR |
|---|---|---|
| 1 | 0.000 | 0.0816 |
| 2 | 0.190 | 0.0289 |
| 5 | 0.435 | 0.0054 |
| 10 | 0.574 | 0.0018 |
| 20 | 0.693 | 0.0006 |
| 50 | 0.830 | 0.0002 |
| 100 | 0.923 | <0.0001 |
| 200 | 1.009 | <0.0001 |
| 500 | 1.114 | <0.0001 |
| 1000 | 1.187 | <0.0001 |
Dispersion of trial Sharpe ratios used: 0.365 (annualised standard deviation), set to the null sampling error 1/sqrt(T).
3. Probability of backtest overfitting (CSCV)
Not computed: no variant matrix supplied.
4. Multiple-testing haircut
| t-statistic (excess SR x sqrt(years)) | -1.44 |
| p-value, single test | 0.92527 |
| Bonferroni p (1 tests) | 0.92527 |
| Sidak p | 0.92527 |
| Holm p | n/a |
| BHY p | n/a |
| Haircut excess Sharpe (Bonferroni) | -0.000 (100% haircut) |
| Haircut excess Sharpe (BHY) | n/a |
5. Reality Check and SPA
| Benchmark | buy-and-hold of the traded asset |
| Variants in the test | 1 |
| Best mean excess return (annualised) | -19.20% |
| White Reality Check p | 0.920 |
| Hansen SPA p (consistent / lower / upper) | 1.000 / 1.000 / 1.000 |
| Bootstrap | stationary, 1000 draws, mean block 14 |
6. Holdout degradation
| Periods | Annualised excess Sharpe | |
|---|---|---|
| In-sample (before 2024-06-24) | 1920 | -0.629 |
| Holdout | 823 | 0.000 |
| Holdout / in-sample | n/a | |
| Consistency p-value | 0.346 |
7. Period and regime stability
| Volatility regime (trailing 21-day, lagged) | Periods | Annualised mean excess | Annualised excess Sharpe |
|---|---|---|---|
| low vol | 908 | 9.7% | 0.56 |
| mid vol | 907 | -1.3% | -0.03 |
| high vol | 907 | -66.4% | -1.38 |
8. Transaction-cost sensitivity
| One-way cost (bps per unit traded) | 0 | 1 | 2 | 5 | 10 | 20 | 50 |
|---|---|---|---|---|---|---|---|
| Annualised Sharpe | 0.82 | 0.82 | 0.82 | 0.82 | 0.81 | 0.81 | 0.79 |
| Annualised mean return | 41.4% | 41.4% | 41.3% | 41.2% | 41.1% | 40.8% | 39.9% |
Turnover 2.9 units per year; break-even cost 1414.1 bps.
9. Parameter sensitivity
Not computed: no parameter grid supplied.
10. Drawdown distribution
What is fragile, and how to test the fix
FAIL Selection-adjusted significance vs benchmark (Deflated Sharpe)
Excess Sharpe over buy-and-hold of the traded asset (total return): -0.53 a year over 7.5 years. After allowing for 1 variant(s) tried, the best-of-N noise hurdle is an annualised excess Sharpe of 0.00; the probability that the true excess Sharpe exceeds it is 0.082.
How to test a fix: Declare every variant you tried (including abandoned ones), then re-test on data you have not looked at: a sealed forward period or a later holdout. A DSR that only passes at N = 1 is not evidence.
N/A Probability of backtest overfitting (CSCV)
No variant return matrix was supplied, so PBO could not be computed.
How to test a fix: Send the daily returns of every variant you tried; PBO measures how often the in-sample winner disappoints out of sample.
FAIL Multiple-testing haircut vs benchmark (Harvey-Liu)
Excess-return t-statistic -1.44 (single-test p 0.9253). Adjusted for 1 tests (Bonferroni), p = 0.9253; the Bonferroni-haircut excess Sharpe is -0.00.
How to test a fix: A longer sample raises the t-statistic without new selection: extend the test to earlier or later data the rule was not designed on, keeping the rule frozen.
FAIL Beats the benchmark after data snooping (Hansen SPA)
Benchmark: buy-and-hold of the traded asset (total return). Testing whether the best of 1 variant(s) beats it in mean return, with a stationary bootstrap (1000 draws, mean block 14 periods): Hansen SPA p = 1.000, White Reality Check p = 0.920.
How to test a fix: If the rule's value is lower risk rather than higher return, state that claim before testing it (e.g. drawdown or volatility against the benchmark) and test it on unseen data.
FAIL Holdout degradation (excess Sharpe)
In-sample (before 2024-06-24) annualised excess Sharpe -0.63; holdout 0.00 over 823 periods. Consistency p-value 0.346 (low values mean the holdout is unlikely to share the in-sample excess Sharpe).
How to test a fix: Lock a fresh holdout before any further tuning, or run a sealed forward trial with preregistered pass criteria; only data you have not seen can confirm a fix.
FAIL Period and regime stability (vs benchmark)
The strategy's compounded return beat the benchmark's in 0 of 8 calendar years. Annualised excess Sharpe by volatility tercile: low vol 0.56, mid vol -0.03, high vol -1.38.
How to test a fix: Check whether the edge over the benchmark is concentrated in a few years or one regime; if so, state that regime as part of the hypothesis and test it on a period the rule has not seen.
N/A Parameter plateau
No parameter grid was supplied.
How to test a fix: Name variant columns like 'fast=50|slow=200' so the parameter surface can be mapped.
These are suggestions for further statistical testing, not suggestions to trade. Any change to the rules creates a new variant: count it in N and confirm it on data it was not designed on.
What held up
- Transaction-cost headroom (absolute): Turnover is 2.9 units a year. The gross mean return falls to zero at a one-way cost of 1414.1 bps per unit traded. Base case in this report: 10 bps.
Methods appendix
Sharpe ratio and its standard error
Per-period mean over standard deviation of the audited return series (risk-free rate taken as zero), annualised by the square root of periods per year. The standard error uses the IID-normal formula of Lo (2002) and the non-normal correction of Mertens (2002), which widens the error for negative skew and fat tails. Lo's autocorrelation-adjusted annualisation is reported alongside. [1], [2]
Probabilistic Sharpe ratio (PSR) and minimum track record
The probability that the true Sharpe ratio exceeds a benchmark (here zero), given the sample length, skewness and kurtosis; and the minimum sample length for 95% confidence. [3]
Deflated Sharpe ratio (DSR)
The PSR measured against the Sharpe ratio one would expect from the best of N skill-less trials, where N is the number of variants tried and the dispersion of trial Sharpe ratios is estimated from the variant matrix (or, without one, set to the null sampling variance 1/T). It corrects for selection and non-normality at once. [4]
Probability of backtest overfitting (PBO) via CSCV
The variant matrix is cut into 16 time blocks; for each of the 12,870 ways of picking half of them as in-sample, the in-sample best variant is ranked out of sample. PBO is the share of splits where it falls to or below the median. It assumes blocks long enough to preserve serial dependence and a variant set that represents the real search. [5]
Multiple-testing haircut
The Sharpe ratio is turned into a t-statistic and p-value, the p-value is adjusted for the number of tests (Bonferroni always; Holm and BHY when every variant's returns are supplied), and the adjusted p-value is mapped back to a haircut Sharpe ratio. [6], [7]
White's Reality Check and Hansen's SPA test
Tests whether the best of all variants beats the benchmark in mean return once the search over variants is accounted for. Uses the stationary bootstrap with mean block length T^(1/3) (at least 5) to keep short-range dependence. SPA studentises and recentres, so poor variants do not dilute the power. [8], [9], [10]
Holdout degradation
In-sample versus holdout Sharpe ratio at a fixed split date (by default the last 30% of the sample). The consistency p-value asks how surprising the holdout Sharpe would be if the in-sample Sharpe were the truth. A holdout only counts if it was not used to design the rule. [11]
Period and regime stability
Returns and Sharpe ratio by calendar year and by tercile of the underlying market's trailing 21-day volatility (lagged one day). Descriptive: the tercile cut points use the full sample. [12]
Parameter-sensitivity surface
Sharpe ratio across the supplied parameter grid. The neighbour ratio compares the chosen cell with its one-step neighbours: a plateau (ratio near 1) is less fragile than an isolated peak. [13]
Transaction-cost sensitivity
Net return = gross return minus turnover times a one-way cost per unit traded, over a grid of costs; the break-even cost is where the mean net return reaches zero. Market impact beyond a flat cost is not modelled. [12]
Drawdown distribution by bootstrap
The maximum drawdown is recomputed on 1,000 stationary-bootstrap resamples of the return series, showing how much deeper (or shallower) the worst loss could plausibly have been with the same return distribution in a different order. [10]
References
- Lo, A. W. (2002). The Statistics of Sharpe Ratios. Financial Analysts Journal 58(4), 36-52.
- Mertens, E. (2002). Comments on Variance of the IID Estimator in Lo (2002). Working paper, University of Basel.
- Bailey, D. H. & Lopez de Prado, M. (2012). The Sharpe Ratio Efficient Frontier. Journal of Risk 15(2), 3-44.
- Bailey, D. H. & Lopez de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management 40(5), 94-107.
- Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2016). The Probability of Backtest Overfitting. Journal of Computational Finance 20(4), 39-69.
- Harvey, C. R. & Liu, Y. (2015). Backtesting. Journal of Portfolio Management 42(1), 13-28.
- Harvey, C. R., Liu, Y. & Zhu, H. (2016). ... and the Cross-Section of Expected Returns. Review of Financial Studies 29(1), 5-68.
- White, H. (2000). A Reality Check for Data Snooping. Econometrica 68(5), 1097-1126.
- Hansen, P. R. (2005). A Test for Superior Predictive Ability. Journal of Business & Economic Statistics 23(4), 365-380.
- Politis, D. N. & Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association 89(428), 1303-1313.
- Bailey, D. H., Borwein, J. M., Lopez de Prado, M. & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the AMS 61(5), 458-471.
- Lopez de Prado, M. (2018). Advances in Financial Machine Learning. Wiley.
- Pardo, R. (2008). The Evaluation and Optimization of Trading Strategies, 2nd ed. Wiley.
Data and hash appendix
- Prices from the Coinbase Exchange public market-data API; we publish derived statistics only, never raw prices; no redistribution. BTC-USD daily closes (UTC).
- Audit period 2019-03-23 to 2026-09-24: the indicator's post-publication window, fixed in the sealed ONCHAIN-1 preregistration. N = 1 declared: the rule is the creator's, frozen as published.
- Block headers fetched by us from the Bitcoin P2P network (every header prev-hash-linked and proof-of-work checked); issuance = consensus subsidy by height.
| Item | SHA-256 / value |
|---|---|
| package | holdout-audit 0.3.0 |
| btc_daily.csv | 064d2d183581eebe94c24ab3e69e318c369068c03adfe61bac39528f6a2d5c17 |
| headers.csv | a69bc2d867edf4e5c46ea5b87aeb0e1fbc24ff795d7776b72407abe4f9025b15 |
| audited_series_sha256 | ca122abd7a4d1df0196a31c7bf58712a727b5627440142796d87642b3e9b4cb2 |
| preregistration_sha256 | aafe3c37f5bbf83bc6e4b9a986ef1c936c3b74a927c253b833a5a7626482de40 |
| config_sha256 | 570f9a67ce043a1ff0e55c6aaef7f06d6317a64f996a1c4e37c775ce9d2ef8c7 |
| bootstrap seed | 20260925 |
| rubric version | 1.1 (docs/RUBRIC-v1.1.md) |
| declared objective | beat_benchmark |
| preregistration seal | f8d7d587466a4bb61b5c179aba44bb766825a6a495cc3ecb68d16e39c1810275 (digicert, freetsa; earliest 2026-09-25T18:22:38Z) |
| rubric document SHA-256 | 2311b126336bab9f426a84f9c9cd9cb937e346b403a62f60e3c96ed8b5868e46 |
| rubric seal record SHA-256 | ac882c422df5154ec7e6e928904c9099d6b286dc9e5369fece4b85df57044138 |
| rubric sealed by digicert (RFC 3161) | 2026-09-25T14:20:10Z |
| rubric sealed by freetsa (RFC 3161) | 2026-09-25T14:20:11Z |
Anyone with the same inputs and package version can re-run the audit and must obtain the same audited-series hash and the same statistics.