The final result in Published Statistic Reproduction 001 is almost comically small:

  • published oracle: -0.05187475902840215
  • Strategy Lab: -0.05187475902840218
  • difference: 5 binary64 ULP
  • predeclared external tolerance: 0x1p-53
  • conformance: PASS

If the task were only “write the same equation in Go,” that would have been a short piece of work.

It was not.

The harder question was whether the number could be tied to an independently reconstructed dataset, produced by a deterministic implementation, closed into immutable evidence, verified later without rerunning the evaluator, and exported for publication without silently recalculating anything.

This note is not a development diary. It pulls out the engineering lessons that became visible only because a real external reproduction pushed the research system outside its previous comfort zone.

The formula was the easy part

The reproduced statistic is model-free and compact. Each 15-minute BTCUSDT bar becomes a sign from its open and close. The current sign is correlated with a strictly lagged field built from the previous 12 signs, weighted by inverse lag.

There is no optimizer, portfolio simulator, transaction-cost model or machine-learning fit in that calculation.

But “same formula” is not the same as “same experiment.” A numerical reproduction also depends on at least four other objects:

  1. the exact input sample;
  2. the normalization rules that turn provider files into canonical observations;
  3. the floating-point contract used by the implementation;
  4. the evidence that binds the result back to those choices.

The Kitron reproduction made all four visible.

A second research shape exposed the first hidden assumption

Strategy Lab had been built around trading-family evaluations: one admitted family, multiple strategy members, simulation, costs, benchmark comparison and a family-level acceptance rule.

The Kitron target was not a trading family at all. It was one deterministic scalar statistic.

The quickest implementation would have been to add a parallel scalar pipeline: separate captured payloads, separate result files, separate terminal verification and separate replay logic. That would have made this one reproduction easier and every future research type harder to maintain.

Instead, the persistence boundary was generalized once.

The admitted research plan became the single research authority. Captured evidence binds to that plan and its approved research identity. The durable research_result_v1 envelope contains exactly one finite payload: the existing trading-family result or the kernel-sign-correlation result. Terminal verification, retained evidence and publication then stay on one lifecycle with a narrow plan-selected branch.

That is a deliberately modest abstraction. There is no generic indicator DSL, evaluator plugin framework or arbitrary research engine. The system knows about the research shapes it actually supports and no more.

The practical lesson is that low maintenance does not always mean the smallest immediate diff. Sometimes it means refusing the tempting second pipeline when a real second use case finally proves which boundary should have been research-level all along.

Evidence size is part of the data contract

Daily strategy studies had also hidden another assumption: some evidence artifacts were small enough that a generic metadata limit looked reasonable.

A 38,977-row intraday sample changed that.

The accepted Kitron run produced:

ArtifactFinal size
Canonical normalized bars7,631,760 bytes
Row-level source attribution4,780,723 bytes

The attribution file is not ordinary fixed metadata. It records row-level lineage: which source archive and source row produced each canonical observation. Its size therefore grows with the dataset, just like the normalized bars.

The important fix was not “remove the limits.” It was the opposite: classify the evidence correctly and keep explicit bounded capacities for the artifacts that scale with rows, while leaving ordinary metadata under smaller limits.

That distinction matters because a fail-closed research system should reject unexpectedly large artifacts. The bound simply has to describe the artifact it is actually protecting.

The same episode exposed a second engineering rule: write and read limits must agree. Accepting a large canonical artifact during capture but rejecting the same bytes during staging verification is not conservatism; it is an inconsistent contract. Publication, loading, terminal verification and replay now use the same artifact-specific capacity rules.

External tolerance does not imply approximate internal evidence

The most interesting numerical problem appeared before Strategy Lab ran the study.

The frozen external oracle was:

-0.05187475902840215

Re-running the authors’ current implementation on the reconstructed public data produced:

-0.0518747590284022

An independent implementation produced the same current value. The data geometry and all inspected supporting statistics matched.

A predeclared diagnostic then evaluated mathematically equivalent binary64 reduction paths on the same operands. They differed by a handful of ULP. The result depended on floating-point accumulation order and numerical backend, not on a different research definition.

That forced an important distinction:

External numerical conformance can require a narrowly justified tolerance. Internal Strategy Lab integrity should not.

Strategy Lab therefore uses its own deterministic accumulation contract. Its persisted result is exact for that implementation. Result hashes, terminal verification and internal replay comparisons remain exact. Only the comparison against the external oracle uses the predeclared 0x1p-53 tolerance justified by the portability diagnostic.

That separation prevents a dangerous shortcut: once “close enough” becomes a generic internal concept, evidence verification can quietly stop proving what people think it proves.

Completed should mean verification-only

The accepted run reached a durable terminal state with the result, data identities and provenance already fixed.

A later terminal re-entry did not ask the provider for data and did not invoke the evaluator. It verified the retained evidence and completed artifacts only. During the acceptance check, that re-entry was performed in a network-disabled environment; the run remained unchanged.

That is a stronger property than “the program gives me the same answer when I run it again.”

Rerunning an evaluator can hide changes in external data, dependencies or runtime behavior. Verification-only closure instead asks a narrower question:

Do the immutable artifacts still prove the completed run we already recorded?

For a research publication, that is often the more useful question.

Replay should test a change, not perform a ritual

Strategy Lab also has a changed-evaluator replay path. It reconstructs the retained evidence and asks whether a new numerical implementation produces exactly the same scientific payload.

For the final Kitron acceptance, replay was reported as NOT_APPLICABLE_REDUNDANT.

That was intentional. The current numerical implementation identity was identical to the source run’s identity. Running the same evaluator again would not test a changed implementation, so the replay command refused to create a meaningless child run.

This is a small example of a broader principle: controls should have a purpose. A green checkbox is not stronger evidence if the operation behind it tests nothing new.

Publication is downstream of verification

The public Kitron article was not generated by asking an editorial process to reopen the research and calculate the headline number again.

The completed run first passed terminal verification. The publication-handoff command then projected stored, verified facts into a structured external artifact: research identity, dataset identity, sample geometry, statistic, oracle, tolerance, error, conformance and provenance.

The publishing repository consumes that handoff. It does not need access to the private runner, raw exchange archives or evaluator implementation in order to quote the result.

This boundary reduces a subtle publication risk. A charting script, template or AI-assisted editorial step should not become an accidental second implementation of the research.

The direction stays one-way:

captured evidence → deterministic evaluation → verified completed run → publication handoff → article

not:

article needs a number → recalculate something until a number appears.

What became reusable

Most of the work exposed by this case is not specific to the Kitron statistic.

The reusable pieces now include:

  • a research-level persistence envelope that can hold a trading family or a finite scalar study;
  • one captured-evidence lifecycle tied to the admitted research plan;
  • Binance intraday archive geometry with explicit gap and duplicate rejection;
  • artifact-specific bounded handling for row-scaled canonical evidence;
  • deterministic scalar result persistence;
  • exact terminal verification and changed-evaluator comparison semantics;
  • a verified publication handoff that does not recalculate research;
  • an explicit separation between exact internal integrity and justified external numerical tolerance.

What remains specific to this reproduction is much smaller: BTCUSDT, 15-minute bars, the fixed sample period, the 12-lag alpha-1 kernel statistic, and its frozen external oracle.

That is the payoff from doing the integration work at the research boundary rather than hiding it behind a one-off script.

What this engineering does not prove

A disciplined pipeline can make a result more inspectable. It cannot make the underlying scientific claim broader than the experiment.

The Kitron case still establishes only that one BTCUSDT statistic on one fixed sample was independently reconstructed and reproduced within its predeclared numerical contract.

It does not establish profitable mean reversion, general cross-market validity, live tradability, or correctness of every result in the source paper.

Engineering discipline helps answer what produced this number? It does not turn one reproduced number into a stronger economic conclusion.

Inspect the worked example

The full Published Statistic Reproduction 001 case study gives the scientific result, exact sample, five-ULP comparison, limitations and public publication handoff.

For the broader research rules behind the distinction between reproduction, extension and evidence closure, see the methodology.

No research was replayed or recalculated for this engineering note. It is an explanation of the already-completed and verified reproduction.