How a claim becomes a verdict.
Abstract
The pipeline behind every Praxis verdict: a crawler reduces the research firehose to finance-relevant papers (§1), extraction distils each into a structured spec (§2), the consequential ones are rebuilt end-to-end under documented deviations, known-value tests and negative controls (§3), then pushed through the C0–C7 gate ladder (§4) to a fixed-function verdict (§5). Nothing here forecasts. The engine describes what a published rule did: out of sample, net of costs.
1The corpus
Everything downstream depends on knowing what exists. The crawler runs daily across the open research firehose: 415,961 documents scanned, 42,984 strategy papers mapped, ≈480 new every month.
2Extraction
A paper is only comparable to another paper once both are the same shape.
A paper reduced to fixed pillars: asset · data · predictor · method · universe · validation. Each field is either quoted from the text or marked undisclosed. 143 papers carry a full spec to date.
The pillars are the engine's coordinate system: they make every spec searchable over MCP, rank which claims deserve a week of reproduction, and expose the white space no single paper can see.
3Reconstruction
The rule, rebuilt from the paper's own description and nothing else.
Each reproduction is built from the published text alone. Where the paper is silent, the gap is filled with a documented estimate and flagged as such; where its data cannot be sourced, the closest survivorship-clean proxy is substituted and recorded. The deviation ledger ships with every report: the reader never has to trust us on what changed.
- extract the spec · quote every disclosed parameter
- source survivorship-clean reference data · record the vintage
- rebuild the strategy · fill undisclosed gaps, flag each
- test every block against known values · run negative controls
- validate through C0–C7 · publish claimed vs. measured
The controls are adversarial by construction. A look-ahead trap feeds tomorrow's data through the pipeline: if the honest run were peeking it would light up (on the network-momentum paper the trap explodes the Sharpe to 22.3; the honest run reads 0.62). A signal shuffle scrambles the predictor and the edge must collapse to zero (it does: 0.01). Closed-form cases pin the numerics.
4The gate suite
Eight gates, fixed before any paper is run. A claim that survives all eight has nowhere left to hide.
| Gate | What it tests | |
|---|---|---|
| C0 | Replication | Faithful re-implementation on the paper's own period, universe (or closest free proxy), and cost assumptions. Does the mechanism exist at all? |
| C1 | Trading costs | Realistic transaction costs, fees and implementation frictions applied: the numbers the paper should have printed. |
| C2 | Significance | Statistical reality: multiple-testing corrections (Bonferroni / Holm / BHY across the paper's own configuration grid) and the Deflated Sharpe Ratio [1]. |
| C3 | Crisis regimes | The worst regimes: 2008, Volmageddon 2018, COVID 2020, the 2022 rate shock. Advertised protection is tested, not assumed. |
| C4 | Generalisation | Other universes, countries and asset classes, where the claim implies them. |
| C5 | Capacity | Capacity, borrow and tradability stress, with cost sweeps to breakeven. |
| C6 | Persistence | True post-publication out-of-sample: the window the authors could not have fit. |
| C7 | Originality | Spanning regressions on standard factor models (HAC / Newey–West errors): alpha, or relabelled beta [3]? Probability of Backtest Overfitting via CSCV assessed alongside [2]. |
4.1The deflation arithmetic
Gate C2 is where most claims die. The engine treats a reported Sharpe as the maximum of a search the authors ran, disclosed or not, and asks what survives once that search is priced in.
5The verdict
A fixed function, applied after the gates. Never an opinion.
| Verdict | Meaning |
|---|---|
| validated | The claim replicates cleanly on honest data. |
| overstated | The effect is real but materially smaller than claimed. |
| fails | The claim does not survive honest reproduction. |
Across the audits to date, the median Sharpe haircut on Sharpe-denominated claims is ≈ 60% of the published number, before anyone traded a dollar against it.
6Access
Every verdict is public. Readers open ten exemplar reproductions in full (gates, deviation ledgers, notebooks) and query them over MCP where they already work. The full corpus rides with the desk. The engine never emits position-level output.