Mean excess return
Mean daily net log return of the strategy minus the identically costed buy-and-hold benchmark. A member could pass only if this quantity was positive and its dependence-aware lower confidence bound was also positive.
Published strategy out-of-sample validation
We froze Grobys, Ahmed and Sapkota's published log-price VMA horizons before evaluation and tested them on Binance BTCUSDT daily data from 2019 through 2025. Some longer horizons looked economically attractive, but none established positive excess return under the frozen HAC criterion.
Here, FALSIFIED means this experiment failed its predeclared acceptance rule—not that every version of the strategy is disproved or that the original source is wrong.
Reading the result
The headline metrics answer different questions. These short definitions are the interpretation used on this page, not additional evaluation criteria.
Mean daily net log return of the strategy minus the identically costed buy-and-hold benchmark. A member could pass only if this quantity was positive and its dependence-aware lower confidence bound was also positive.
A one-sided 95% lower confidence bound for mean daily excess log return using Bartlett-HAC with a frozen 20-day lag. Positive-looking point estimates did not qualify when this lower bound remained below zero.
The 20-day VMA was frozen as the named primary because the source emphasizes that horizon. Under the primary-plus-robustness family rule, robustness members could describe sensitivity but could not rescue a failed primary.
The primary numerical cost was 15 bps per executed leg. Ten- and 20-bps cases were frozen in advance as descriptive diagnostics and could not replace the primary classification.
Five published log-price moving-average horizons were frozen before evaluation and tested on Binance BTCUSDT daily data from 2019 through 2025. None met the predeclared excess-return criterion.
That headline hides the more interesting result. The 50-day and 100-day variants looked good by ordinary backtest standards. The 100-day rule compounded to +3152.03%, annualized at 64.38%, posted a log-return Sharpe of 1.078, and cut maximum drawdown to 44.35% versus 76.63% for buy-and-hold. Yet its dependence-aware lower confidence bound for mean excess return remained below zero, so it did not survive the frozen test.
The source-emphasized 20-day rule was less ambiguous: it annualized at 30.53% versus 57.04% for the identically costed benchmark and had a negative mean daily excess log return.
This is exactly the kind of case where a visually impressive backtest and a predeclared evidentiary standard give different answers.
Klaus Grobys, Shaker Ahmed and Niranjan Sapkota published Technical trading rules in the cryptocurrency market in Finance Research Letters (Volume 32, 2020, article 101396).
The paper studies daily prices for 11 cryptocurrencies over 2016–2018. Its variable moving-average rule uses a one-day short average, log(P_t), against a long moving average of the most recent n log prices. The published long horizons are 20, 50, 100, 150 and 200 days. The analysis focuses on buy-side positions rather than short positions.
The paper’s main market-wide result is not a Bitcoin-only backtest. It uses a multivariate Seemingly Unrelated Regression (SUR) framework to test average strategy payoffs across cryptocurrencies while accounting for contemporaneous dependence. Its headline result excludes Bitcoin and reports an annualized excess return of 8.76% for the 20-day VMA after controlling for the average market return.
Bitcoin appears in the paper’s 11-coin robustness analysis. At the individual-coin level, the paper reports statistically significant positive Bitcoin payoffs for all five VMA horizons, with the strongest significance at 20 days and declining economic/statistical strength at longer horizons.
Those details matter because they set the boundary of what this new experiment can claim.
The strategy source is essential here, but the word replication would overstate what was done.
A strict reproduction of the paper would require, at minimum, the original 2016–2018 cryptocurrency panel, the authors’ data source and return construction, the 10-coin and 11-coin cross-sectional tests, and the paper’s SUR inference. Today’s Strategy Lab production capability is deliberately narrower: Binance Public Data, Spot, BTCUSDT, daily bars.
So this experiment asks a different but still source-anchored question:
If the paper’s published Bitcoin VMA horizons are frozen before seeing a later sample, do they establish excess return over an identically costed buy-and-hold benchmark on Binance BTCUSDT under one explicit execution and inference contract?
The answer is no for this frozen 2019–2025 evaluation.
The important distinction is:
That makes the paper useful without pretending the experiment is something it is not.
The external source PDF was frozen by SHA-256 before the research run:
8a1783f51898b59a3293707f760f7ebcf859ed232738e07809c94b814f11e83a
The five family members were:
| Member | Role | Long lookback |
|---|---|---|
vma_20 | primary | 20 days |
vma_50 | robustness | 50 days |
vma_100 | robustness | 100 days |
vma_150 | robustness | 150 days |
vma_200 | robustness | 200 days |
The 20-day rule was named primary before evaluation because the paper emphasizes that horizon. The family rule was primary_plus_robustness, with k=0: the primary had to meet its criterion, while the four longer horizons were preserved as source-defined robustness members and could not rescue it.
The evaluation interval was fixed as 1 January 2019 through 31 December 2025, immediately after the source paper’s 2016–2018 historical sample. The shared capture started on 15 June 2018 only to supply the longest 200-day warm-up. Every member started independently FLAT at the common first evaluation open.
Signals used completed UTC daily closes. A changed target position executed at the next daily open. The strategy was LONG/FLAT, fractional all-in/all-out, with terminal liquidation at the 1 January 2026 open.
The primary transaction-cost assumption was 15 basis points per executed leg, with 10 and 20 bps frozen in advance as diagnostics. Buy-and-hold was charged under the same per-leg cost convention.
None of those execution, cost or inference choices are claims about the original paper’s exact implementation. They are the explicit Strategy Lab contract used for this validation.
The deterministic research plan required 2,758 daily rows, including warm-up and the terminal-support candle.
The production run retrieved 92 official Binance Public Data archives and their published checksum companions. Zero archives were reused from earlier retained evidence. The normalized daily dataset contained exactly the required 2,758 rows.
Key immutable identities were:
e0c123864b33c5d45e228af42f7430a37a04837d1f4fbd8e76e74174f99961ba3744822a0e3cb1b34384c40dbe45b626f0d41a7d4af762d5fbeff0da471ac98a7920dbed13fa862d5beade58b981d647318c859d778c756e62358039cedcc49c21f579e0113ee397f1946083ef0e0692a1978845e9760dfa97036a7091ca7d54715e03437132a099035865351d4ca5c2ffae5bfdd0ca27aae6b1d1a1baa1db801d17ab812433b05ce503570427365b6ad1089b4afd82af09f13cea2d5a81d150The final canonical run state was REPORTED / COMPLETE. The evidence bundle binds the dataset, captured source lineage, evaluator/build identity, result, report and deterministic article scaffold.
At the primary 15-bps-per-leg cost, buy-and-hold compounded to +2260.99%, corresponding to 57.04% annualized, with a log-return Sharpe of 0.709 and maximum drawdown of 76.63%.
The strategy family produced a much less uniform picture:
| Variant | Annualized return | Log Sharpe | Max drawdown | Mean excess bps/day | HAC 95% lower bound bps/day | Frozen outcome |
|---|---|---|---|---|---|---|
vma_20 | 30.53% | 0.619 | −69.83% | −5.07 | −11.62 | FAIL |
vma_50 | 60.44% | 1.052 | −60.04% | +0.59 | −5.84 | FAIL |
vma_100 | 64.38% | 1.078 | −44.35% | +1.25 | −5.14 | FAIL |
vma_150 | 56.63% | 0.966 | −45.97% | −0.07 | −6.31 | FAIL |
vma_200 | 40.34% | 0.713 | −61.85% | −3.08 | −9.27 | FAIL |
The family result was DID_NOT_MEET_CRITERIA / COMPLETE, with 0 of 5 members meeting excess_return_v2.
If the test had been an ordinary retrospective backtest, the 100-day row would be easy to sell.
It beat buy-and-hold on cumulative and annualized return. Its log-return Sharpe was materially higher. Its maximum drawdown was more than 30 percentage points shallower. At the lower 10-bps diagnostic it annualized at 65.39%; at the higher 20-bps diagnostic it still annualized at 63.38%.
But the test was not “find an attractive row.”
For vma_100, mean daily excess log return was positive at about +1.25 basis points per day. Its HAC standard error was large enough that the one-sided 95% lower bound was still approximately −5.14 bps/day.
So this was not evidence of a reliably positive excess-return mean under the frozen criterion.
Promoting vma_100 after seeing those descriptive metrics would also violate the family design. It was frozen as a robustness member. The source-emphasized vma_20 was the primary, and robustness members were never allowed to replace it after outcome access.
The attractive 100-day result is therefore worth reporting precisely because it was not allowed to become the strategy after the fact.
The named primary did not suffer from a subtle confidence-interval problem around an otherwise attractive point estimate.
At 15 bps per leg it returned +546.44% cumulatively and 30.53% annualized. Those numbers sound large in isolation. Bitcoin buy-and-hold over the same period returned +2260.99% cumulatively and 57.04% annualized.
The primary strategy’s mean daily excess log return was about −5.07 bps/day, and its one-sided 95% HAC lower bound was about −11.62 bps/day.
Because the primary failed, the predeclared family rule failed. In this particular run the robustness members did not create a difficult rescue question anyway: none of them met excess_return_v2 either.
The primary cost model was intentionally simple: 15 bps per executed leg, not a reconstruction of every historical Binance fee tier, spread or account discount from 2019–2025.
Two predeclared diagnostics tested 10 and 20 bps per leg. The results moved in the expected direction, but the broad interpretation did not become knife-edge.
For example, the 100-day variant’s annualized return was:
The diagnostics are descriptive only. They cannot change the classification bound to the primary 15-bps case.
It does not show that Grobys, Ahmed and Sapkota’s original 2016–2018 results were wrong.
The paper used a different historical period, a broader cryptocurrency universe, CoinMarketCap data and a cross-sectional SUR research design. This experiment used one modern BTCUSDT spot series, a next-open execution convention, explicit costs, an identically costed buy-and-hold benchmark and a HAC test of mean daily excess log return.
The supported conclusion is narrower:
the five published Bitcoin VMA horizons did not establish positive excess return in this later 2019–2025 Binance BTCUSDT validation under the frozen Strategy Lab contract.
That conclusion also does not establish that moving-average rules are universally ineffective. Different assets, periods, execution assumptions, objectives or statistical hypotheses are outside this experiment.
The result does show that the paper’s source-emphasized 20-day Bitcoin rule did not carry its historical evidence into this particular later test, and that even the descriptively attractive longer horizons did not clear the predeclared uncertainty threshold.
Because the source changes what can be chosen after looking at the data.
Without an external source, it would be easy to inspect 2019–2025 BTC, discover that a 100-day log-price filter looks attractive, and present that parameter as if it had been the hypothesis all along.
The paper gives the experiment an externally defined mechanism and parameter family before the later sample is evaluated. That makes the subsequent validation meaningful even when the original study cannot be reconstructed exactly with the current capture capability.
But the distinction must remain visible. A useful publication program should separate at least three questions:
This case combines the second and third. It should not be relabeled as the first.
A future true replication of Grobys et al. would require extending the lab to the paper’s multi-cryptocurrency 2016–2018 panel and its joint SUR inference. If that were done, the strongest design would be a two-stage publication: first reproduce the original result, then compare it with this already-frozen later-period validation.
This publication was written from the already CLOSED canonical run. The strategy was not rerun, retuned or reclassified for the article.
The evaluator was built from clean Git revision:
738d671f8ff7d3fc01f9a8b75eb584c6e7653f71
with numerical implementation identity:
e380304a47088ec103a7cca35b00f5de448dd11fb38fa28046bfdc00f1b5dd99
and executable SHA-256:
f23943f6fbf43c9f068cb4abb6e8a5bf6dc63d03372b59c2ab6b7ebf0a023a2b
The deterministic scaffold itself explicitly leaves source critique, interpretation and conclusion operator-authored. That is intentional: the software can prove what was captured, evaluated and persisted; it cannot decide whether a source substitution deserves the word replication.
This later-sample test did not preserve the source paper’s strongest Bitcoin VMA signal under Strategy Lab’s frozen validation contract.
The source-emphasized 20-day rule substantially lagged buy-and-hold. Two longer variants — especially 100 days — produced attractive descriptive performance and materially smaller drawdowns, but none established a positive mean excess return with a positive one-sided HAC lower bound. The family therefore closed DID_NOT_MEET_CRITERIA, with no missing, inconclusive or failed evaluations.
The broader lesson is methodological rather than anti-moving-average: a backtest can look excellent and still fail the question that was actually frozen before the result.
That is why the source, the substitutions, the acceptance rule and the evidence lineage all need to remain visible at the same time.
Frozen decision rule
The named 20-day primary did not meet excess_return_v2, so the family failed under the predeclared primary-plus-robustness rule. Independently, all four longer-horizon robustness members also failed because their one-sided HAC lower bounds remained below zero.
Inspect the evidence
These are exact artifacts copied from the CLOSED canonical Strategy Lab run. The local canonical bundle additionally preserves the full artifact index, build provenance, result.json, report.md, normalized bars, research plan, and all 92 official Binance archives with checksum companions; those larger/raw objects are not all republished here.
These excerpts support inspection of the reported work; they are not a complete package for independently rerunning the experiment. Read the evidence policy.
b0bcc6a9552644b29e120ee87ac095854575e2b5580bc021c06f0c99092bfa74
Terminal run state
Exact canonical state showing REPORTED / COMPLETE and binding the final manifest and artifact-index hashes.
SHA-256
4ca0276ace8d0feae12df8fbd92ed7aaf76a865b739e77be841f923bbcac26b2