A reproduction does not have to reach the backtest to teach you something.
We recently evaluated Wu and Pinsky (2026), On the Performance of Lagged Momentum and Reversal Strategies Across Daytime and Overnight Sessions in Bitcoin and Ethereum Cryptocurrencies as a possible external numerical-conformance case for Strategy Lab.
On paper, it looked unusually attractive. The study used simple hourly BTC/USD and ETH/USD prices, named Kraken as the source, published executable Python code, and committed machine-readable result tables in a public replication repository.
The intended exercise was straightforward: independently acquire the same market data, freeze it, reproduce one published result, and only then decide whether Strategy Lab should implement the corresponding deterministic mechanism.
That is not what happened.
The source data were authentic. The input was still not reproducible.
Authenticate before interpreting
The first rule was simple: do not modify Strategy Lab until the external reference path works independently.
Kraken publishes its historical OHLCVT archive as a set of large downloadable parts together with SHA-256 checksums. We downloaded the official Kraken_OHLCVT_Full_2026Q2 release, verified every part independently, reconstructed the full ZIP, and verified the reconstructed archive against Kraken’s published hash.
All five parts matched. The final archive matched exactly.
The authoritative BTC/USD hourly member was XBTUSD_60.csv, with SHA-256 287a0598715c0892636e6035e6958414000685163ef8d7f015c4077c3721a601.
So this was not a stale mirror, a third-party CSV, a corrupted download, or an accidental exchange substitution. The bytes came from the source named by the paper and passed the source’s own integrity chain.
That distinction mattered because the next failure could no longer be dismissed as bad acquisition.
The authentic data did not satisfy the required input contract
The reproduction target covers every UTC hour from 2016-01-01 00:00:00 through 2025-12-31 23:00:00.
A complete hourly grid over that interval contains 87,672 timestamps.
The checksum-authenticated Kraken file contained only 87,408 observed rows in that interval.
The differences were concrete:
| Check | Required target | Official Kraken data |
|---|---|---|
| Hourly rows | 87,672 | 87,408 |
| First timestamp | 2016-01-01 00:00 UTC | 2016-01-01 04:00 UTC |
| Last timestamp | 2025-12-31 23:00 UTC | 2025-12-31 23:00 UTC |
| Missing hours | 0 | 264 |
| Internal discontinuities | 0 | 136 |
The first four sample hours were absent. The final traded hourly row before the sample was 2015-12-31 21:00 UTC, with a close of 431.31.
This behavior is consistent with an important property of Kraken’s historical OHLCVT archive: intervals with no trades are not necessarily represented as synthetic candles.
That does not make the Kraken archive defective. It means an additional data-construction rule is required if a downstream analysis expects a completely populated clock-time grid.
And that rule is part of the experiment.
Why we did not simply fill the gaps
It would have been easy to create 264 missing candles and continue.
The temptation was to start fixing a hole in the hourly grid. But the moment we chose how to fill those 264 missing hours, we would be making a methodological decision the published materials did not specify.
It would also have changed the question.
If the original analysis used a particular convention for no-trade hours — carry the previous close, insert zero returns, use another historical feed, merge an older snapshot, or something else — then independently choosing a convenient rule would create a new input dataset.
The authors’ replication repository makes this boundary unusually visible. It includes the analysis code and frozen output tables, but not the BTC/USD and ETH/USD input files. Its README explicitly notes that the raw crypto data were not committed and that provenance and the treatment of hours with no trades require independent verification.
So we did not search over filling methods.
We allowed exactly one predeclared forensic hypothesis.
One hypothesis, chosen before seeing the answer
The hypothesis was:
For an hourly interval with no trade, carry forward the most recent observed close.
This is plausible enough to test because it preserves a zero return across an empty interval without looking into the future.
The reconstruction was performed outside Strategy Lab. It used the authentic Kraken member as its only market-data source, seeded the beginning of 2016 from the last prior traded close, inserted only missing hourly timestamps, and then sliced the exact 2016–2025 sample.
The resulting candidate had:
- exactly 87,672 hourly rows;
- exactly 264 inserted hours;
- no remaining timestamp gaps;
- 137 gap runs including the initial four-hour run;
- a maximum consecutive fill of 23 hours.
Its deterministic CSV SHA-256 was 2ee16601652b7fee38652ec055cff98f59e2f048d0d1527612f8d9e2f59e9d11.
The lineage remained explicit:
- official Kraken
XBTUSD_60.csv; - carry forward the preceding close only across absent hours;
- produce the forensic complete-grid candidate.
We did not call this reconstructed file “raw Kraken data.” It was a hypothesis about an undocumented transformation.
Then we ran the authors’ published formulas against it.
The shape matched. The results did not.
The most visible BTC result in the paper is the 08:00 UTC Reversal/Reversal rule. The mismatch was not a rounding issue.
| Cost | Candidate final wealth | Published final wealth | Candidate Sharpe | Published Sharpe | Candidate trades | Published trades |
|---|---|---|---|---|---|---|
| 0 bps | 39,364.9263 | 95,015.4594 | 1.21439 | 1.35287 | 3,750 | 3,743 |
| 1 bp | 18,610.0238 | 44,946.1341 | 1.10496 | 1.24202 | 3,750 | 3,743 |
| 2 bps | 8,796.6915 | 21,258.1457 | 0.99544 | 1.13109 | 3,750 | 3,743 |
We also compared broadly instead of focusing only on one headline row.
Across the authors’ two committed BTC result tables, 1,425 rows were compared. Only 12 rows matched completely. In the full 12-cutoff table, zero rows were complete exact matches.
The largest discrepancies were economically material, not floating-point noise. For the seven-cutoff result table, the maximum absolute final-wealth difference exceeded 55,650.
At that point the predeclared hypothesis had failed.
Why we stopped instead of trying another repair
Once you know the answer you are trying to reproduce, preprocessing becomes a dangerous search space.
We could have tried:
- another fill convention;
- another historical Kraken snapshot;
- a different interpretation of timestamps;
- forward/backward combinations;
- an alternative exchange with similar prices;
- special handling for long gaps.
One of those choices might have moved the output closer to the published table.
But selecting a data transformation because it improves agreement with the target is no longer independent reproduction. It is reverse-engineering the answer.
Our stopping rule was therefore simple:
one plausible reconstruction rule, declared in advance; if it materially fails, stop.
The conclusion was correspondingly narrow:
The Wu–Pinsky BTC input was not reconstructible from the published materials under the predeclared hypothesis.
That statement does not mean the published calculations are necessarily wrong. The authors may have used a different preprocessing step or a historical input file that is no longer available in the public package.
It means we could not independently recover the numerical object needed to treat their committed outputs as an external oracle.
Three different questions that are easy to confuse
This attempt clarified three separate properties of research data.
1. Is the source authentic?
Yes.
The archive matched Kraken’s published checksums, and the selected hourly member came from that authenticated archive.
2. Is the authentic source suitable for the claimed experiment as-is?
No.
The official hourly file did not contain the complete clock-time grid required by the reproduction target.
3. Can the published result be reproduced from the available materials?
Not under the one independently declared reconstruction hypothesis we tested.
A cryptographic checksum can answer the first question very strongly while saying nothing about the other two.
That is why provenance is necessary but not sufficient.
The software consequence mattered too
During the investigation, Strategy Lab briefly gained a narrow local Kraken archive importer. It authenticated the source, selected XBTUSD_60.csv, validated the requested interval, and was designed to publish only fully verified evidence.
In practice, that strict boundary did exactly what it was supposed to do: the real archive failed admission rather than being silently repaired.
But once Wu and Pinsky was rejected as the reproduction candidate, the Kraken product path no longer had an active justification.
So it was removed.
This is another useful form of discipline. Experimental support should not become permanent architecture merely because engineering effort has already been spent on it. In a maintenance-first system, a capability that exists only for a rejected experiment is a liability unless another current requirement justifies it.
The evidence from the attempt was kept. The unused product surface was not.
A stricter gate for the next reproduction
The failed attempt changed the selection rule for future external conformance cases.
Before Strategy Lab gains any paper-specific capability, the external candidate should first demonstrate:
- an exact frozen input dataset, or a public acquisition path that reproduces an exact frozen hash;
- a pinned executable reference implementation;
- machine-readable expected outputs;
- successful execution of that implementation from the independently recovered input;
- a clear deterministic calculation worth implementing independently.
Only after those conditions pass should product architecture enter the discussion.
That moves the cheapest failure mode to the beginning of the process.
Failure before the backtest is still evidence
No Strategy Lab strategy verdict exists for this candidate. We never reached the stage where a deterministic trading evaluation would have been scientifically meaningful.
That is not a missing result. It is the result of the reproducibility gate.
The official data passed authenticity checks. The required input could not be recovered without an undocumented transformation. The one transformation we agreed to test did not reproduce the published numerical outputs. We stopped rather than search for a preprocessing rule that would make the answer fit.
For a system built around frozen specifications and deterministic evidence, that is preferable to a successful-looking backtest built on an input whose lineage we cannot defend.
No research was replayed for this note. It documents a completed reproduction preflight and the engineering decisions made from its preserved diagnostics.