Attention-shock cycle
Real-data deterministic evaluation
FALSIFIED
Reached evaluation on 1,000 rows and failed. The negative result was preserved rather than used as a prompt to retune the tested rule.
Backtest Discipline
A selective record of empirical falsifications, pre-evaluation stops, and the reproduction and prospective milestones that closed the project.
The closed research program was not a single strategy test. Its lineage contains many candidate mechanisms and several different kinds of terminal outcome. A preserved checkpoint recorded 16 completed empirical studies and 16 empirical falsifications. Other candidates stopped earlier because a frozen source requirement, data-admission rule, or capability boundary could not be satisfied.
This page is deliberately selective. It does not turn every internal iteration into a public article. The entries below are included because they show materially different ways controlled research can end.
Representative completed evaluations that reached a frozen empirical verdict.
Real-data deterministic evaluation
FALSIFIED
Reached evaluation on 1,000 rows and failed. The negative result was preserved rather than used as a prompt to retune the tested rule.
Deterministic evaluation
FALSIFIED
Generated 344 completed round trips, then failed its required aggregate cost-adjusted return and chronological-block criteria.
Negative research outcomes where a frozen prerequisite failed before a strategy verdict could be produced.
Prospective experiment freeze
REJECTED BEFORE CAPTURE
Required 100,000 historical rows per symbol series, but the frozen request budget could supply only 4,000 per series. No public-data request or evaluator run followed.
Data admission
CLOSED BEFORE EVALUATION
Required at least 5,000 normalized rows per symbol; the real capture produced 4,995. The five-row shortfall was not filled or worked around after inspection.
The two final scientific milestones closed the project with different kinds of result.
Reproduction benchmark
FULL_REPLICATION_MATCH
Independent research machinery reproduced the externally specified executable behavior and independently recaptured the frozen official-data panel exactly. This established reproducibility, not profitability.
Prospective evaluation
FALSIFIED
The frozen rule failed because net Sharpe was about 0.368 versus 0.50 required and the one-sided 95% HAC lower bound was negative. The 3-of-4 positive-fold criterion and the −50% drawdown boundary both passed.
Benchmark 004: reproduce the result before trusting it →
Independent reimplementation and official-data recapture established that the external channel-breakout result was reproducible, while preserving the audit findings that reproduction did not resolve.
Research 005: why a positive return still failed →
The final prospective experiment shows the next step in the lineage: the reproduced executable strategy was frozen into a new 365-day holdout and failed two of its seven predeclared survival criteria.
FALSIFIED is reserved here for an empirical proposal that reached its frozen evaluation and failed the predeclared survival criteria. A pre-capture feasibility rejection or a data-admission failure is reported separately because no empirical strategy verdict was produced. FULL_REPLICATION_MATCH belongs to a different category again: it says the independent reproduction machinery matched the specified behavior, not that a trading strategy was profitable or deployable.
Keeping those categories separate is part of the research method. It prevents a source failure from being presented as market evidence, a reproduction benchmark from being presented as alpha, or a positive-looking metric from replacing the acceptance rule that was fixed in advance.
The internal research history contains more proposals, capability checks, source constraints, rejected discoveries, and implementation iterations than belong on a useful public page. Publishing all of them would blur the distinction between engineering history and research evidence.
The public record will grow only when an additional case teaches something distinct and can be tied cleanly to preserved evidence. The aim is not to maximize the number of strategy names on the site. It is to make the research process inspectable without overstating what any individual experiment demonstrated.