MGPraxis42,984 papers mapped
Validation · #363490 · Equity indices (US · Germany · Japan)

Statistical jump-model regime timing

Downside Risk Reduction Using Regime-Switching Signals: A Statistical Jump Model Approach, 2024

Yizhan Shu, Chenyu Yu & John M. Mulvey · Princeton ORFE · arXiv
Download notebookNotebook
The rule

Let a two-state jump model read the index's own recent returns and say whether the market is in a bull or a bear regime; hold the whole index while it says bull, hold 3-month bills while it says bear, switching one day after the label flips. The jump penalty that makes the label sticky is re-picked every month from an 8-year cross-validation. Claimed on the S&P 500: Sharpe 0.68 against 0.48 for buy-and-hold, with the same lift on the DAX and the Nikkei 225.

Fails0.680.50claimed → measured Sharpe
Praxis Pro members only

The full audit — gate ladder, deviation ledger, inside-the-model statistics and the reproduction notebook — opens with Praxis Pro.

Become a member — £29/mo ↗
0.50
Measured S&P 500 Sharpeclaimed 0.68 · buy-and-hold measures 0.50 — the strategy earns exactly the index's Sharpe while holding 72% of it
0.52
The 200-day moving averagesame delay, same costs, smaller drawdown (−23% vs −34%) — the oldest timing rule beats the jump model; a constant 72% mix ties it at 0.49
p = 0.52
Sharpe gain vs buy-and-holdpaired block bootstrap on the S&P 500; the DAX is worse than the index (p = 0.82), the Nikkei keeps the direction but not significance (p = 0.23)
AssetS&P 500, DAX, Nikkei 225 — the first non-US markets in the library
StrategyRegime timing, 100% index or 100% bills — the first regime-switching paper here
Period1990–2023 out-of-sample (DAX 2008–2023 on public data), then 2024–2026
Costs10 bp one-way, one-day delay; 5- and 10-day delays tested
BenchmarkBuy-and-hold — plus the two the paper omits: a constant mix at the strategy's own exposure and a 200-day moving average
Instruments^SP500TR^GDAXI^N225DTB3Each index is traded alone in local currency against its own 3-month bill. S&P 500 total return: the real total-return index from 1988, Shiller's dividend series accrued onto the price index before that (0.9996 correlation on the overlap). The DAX is already a total-return index; the Nikkei 225 price index carries a documented dividend-yield proxy.

The exact rules

Each day, from the index's own excess returnsCompute three smoothed features: downside deviation (10-day half-life) and two Sortino ratios (20- and 60-day half-lives)
Every six monthsRefit a two-state jump model on the last 3,000 days (ten random starts, keep the best); the state with the higher cumulative excess return in training is bull, the other bear
Every day, after the closeRun the dynamic programme over the trailing 3,000 days with the centroids fixed; today's regime is the last state of the best path — a switch costs the jump penalty λ, so the label is sticky
At each month-endRe-pick λ from the candidate grid: the value whose delayed, cost-adjusted 0/1 strategy had the highest Sharpe over the past 8 years; it applies from two days later
Model says bullHold 100% of the index
Model says bearHold 100% 3-month bills
The label flips at the close of day tTrade at the close of day t+1; the new position is live from t+2; every switch costs 10 bp

The backtest, re-run

$1$2$4$8$16$3219901995200020052010201520202025Jump-model 0/1 timing (S&P 500, net of 10bp)S&P 500 buy-and-hold (total return)
Growth of $1 · log scale
0%-17%-34%1990200020102020
Drawdown · deepest -34.4%
021991200020092018claimed 0.68
252-day rolling Sharpe (excess of bills)

Inside the model

Position mix
S&P 500: 72% index / 28% bills on average — 28% of days in bills across 23 bear spells (median 47 days). DAX 69 / 31%. Nikkei 37 / 63%: whole years in cash in 1990, 1992, 1998, 2001, 2008, 2016 and 2022
Switches
0.44 round trips a year on the S&P (66% turnover; paper 44%), 0.47 on the DAX, 0.49 on the Nikkei
Penalty in force
S&P: λ = 150 in 39% of months, ≥ 100 in 58%; changed in 7.8% of months (32 of 408). DAX median 50, changed 9.9%. Nikkei median 35, changed 5.7%
Win rate
S&P 67.2% of days, 75.2% of months — and beats buy-and-hold in only 13.0% of months. DAX 53.2% / 56.5% / 15.7%. Nikkei 75.4% / 74.6% / 32.2%
Skew / kurtosis
S&P −0.87 / 22.4 daily; DAX −0.52 / 8.1; Nikkei −0.84 / 18.9
Best / worst day
S&P +9.32% / −11.98% · month +11.44% / −24.23% (March 2020, fully invested into the crash)
Annual returns, S&P 500 (strategy vs index, %)
1990 −0.7/3.9 · 91 9.0/30.5 · 92 7.6/7.6 · 93 10.1/10.1 · 94 1.3/1.3 · 95 37.6/37.6 · 96 8.7/23.0 · 97 33.4/33.4 · 98 28.6/28.6 · 99 4.6/21.0 · 2000 −4.6/−9.1 · 01 5.4/−11.9 · 02 1.6/−22.1 · 03 7.1/28.7 · 04 10.9/10.9 · 05 4.9/4.9 · 06 7.3/15.8 · 07 4.9/5.5 · 08 1.4/−37.0 · 09 17.3/26.5 · 10 5.9/15.1 · 11 −3.0/2.1 · 12 6.7/16.0 · 13 32.4/32.4 · 14 6.0/13.7 · 15 −4.7/1.4 · 16 4.3/12.0 · 17 21.8/21.8 · 18 −4.4/−4.4 · 19 31.5/31.5 · 20 −19.7/18.4 · 21 28.7/28.7 · 22 −5.1/−18.1 · 23 10.4/26.3
Where the return goes
In the crash years the strategy is in bills and wins big; in the recovery years it is still in bills, or re-enters late, and loses it back. Same shape on all three indices
Sharpe by fixed penalty (S&P, 1990–2023)
λ 5 0.40 · 10 0.42 · 15 0.49 · 25 0.54 · 35 0.58 · 50 0.55 · 70 0.51 · 100 0.59 · 150 0.53 — the paper's monthly re-selection lands at 0.50, buy-and-hold at 0.50

The validation ladder

C0ReplicateThe mechanics reproduce — the authors' own package and a from-scratch coordinate-descent fitter return identical centroids and paths on all 90 S&P refits, the buy-and-hold rows land on the paper's (10.47% vs 10.2%, vol 18.2%, drawdown −55.3% vs −55.2%), and even the paper's HMM benchmark lands where they put it (0.58 vs 0.54). The headline does not: S&P 500 Sharpe 0.495 against 0.68 claimed, CAGR 8.3% against 11.2%, drawdown −34% against −27%. DAX 0.15 against 0.44 on the honest 2008–2023 window. Nikkei 0.29 against 0.31 — the one index that lands, on a buy-and-hold that is itself higher than tabled.
C1HonestyCosts are not the story — the S&P Sharpe is 0.487 at zero cost and 0.503 at 50 bp, because the cross-validation simply picks a smoother penalty. The honesty test is the two baselines the paper never shows. A constant mix holding the strategy's own average exposure scores 0.493 on the S&P (the strategy: 0.495) and 0.326 on the DAX (the strategy: 0.151). A 200-day moving average with the same delay and costs scores 0.523 on the S&P with a smaller drawdown. Regime identification adds nothing over holding less.
C2DeflateAcross the nine-penalty grid, no fixed λ has a positive active return on the S&P or the DAX (best t = −0.61 and −0.78); Deflated Sharpe 0.02 and 0.02 against a 0.95 bar; PBO 0.49 and 0.72. The Nikkei: best t 0.11, DSR 0.22, PBO 0.50. The best fixed penalty on the S&P (λ = 100, Sharpe 0.59) is an in-sample pick and still below the claim.
C3CrisisThe dodges are real: the S&P strategy sits in bills through 2001–02 (+5.4% and +1.6% against −11.9% and −22.1%), 2008 (+1.4% against −37.0%) and 2022 (−5.1% against −18.1%); the DAX and Nikkei versions sit out 2008 too. But it is still in bills for the rebounds — 2003 +7% against +29%, 2009 +17% against +27%, 2023 +10% against +26% — and COVID-2020 is a whipsaw: −24.2% against −1.1% for the index over February–April, fully invested into the crash and out for the recovery. Four cash years earn about +67 points against the index; six missed rebounds cost about −93.
C5Frictions & delaysCost sweep flat. The delay robustness of Table 5 replicates — S&P 0.495 / 0.473 / 0.467 at 1, 5 and 10 days against 0.68 / 0.71 / 0.70 claimed — because there is little edge to lose to a delay.
C6DecaySince publication (2024-01 → 2026-08) the strategy trails buy-and-hold on all three indices: S&P 0.71 against 1.05 (CAGR 12.4% against 21.3%), DAX 0.29 against 0.97, Nikkei 0.74 against 1.18 — and the 200-day moving average beats it on all three as well. Over the full 1990–2026 span the S&P reads 0.51 against 0.53.
C7OriginalityA Henriksson-Merton timing regression on daily returns finds no timing ability on any index: the timing coefficient is negative on all three (S&P −0.10, t = −1.97; DAX −0.11; Nikkei −0.06), the signature of a strategy that damps the upside more than the downside. CAPM alpha: S&P +1.9%/yr (t = 1.2), DAX −1.0%, Nikkei +2.4% (t = 1.4).
P.S.FragilityEvery friendlier assumption was re-run in full on the S&P. Extending the penalty grid to 500 (the 150 cap bound in 39% of months): 0.525. Adding the 3σ clipping from the authors' own example code, which the paper text says it does not use: 0.568. Fixing the best single penalty with no monthly re-selection: 0.591. The reduced HMM benchmark: 0.580. All of them sit between the 200-day moving average and the paper's own HMM row; none reaches 0.68, none clears the bootstrap. The paper's advertised contribution — re-picking λ monthly to maximise the strategy's Sharpe — is itself worse than nearly every fixed λ ≥ 25 (0.495 against 0.51–0.59): after a long calm it moves to the smoothest penalty, which is then slowest to leave the next crash.
QAIs the wiring honest?Shifting the regime label forward by 5, 10 or 20 days explodes the S&P Sharpe to 1.66, 1.77 and 1.73 (a one-day peek buys nothing on features smoothed over 10–60 days); the same bear spells placed at random offsets give −0.10 ± 0.10 against buy-and-hold, and the real strategy's −0.001 sits inside that distribution. Truncating the data at 2010 reproduces every earlier state exactly. One fit on the whole window with its in-sample labels traded at the same delay reaches Sharpe 1.06 — the model finds clean regimes after the fact and keeps none of it in real time.
Reproduces mechanicsReproduces magnitudeBeats buy-and-hold on SharpeBeats a constant mix at its own exposureBeats the 200-day moving averageReduces drawdown and volatilityStatistically realHolds post-publication

How we rebuilt it

Data
Public substitutes for the paper's Bloomberg total-return indices and GFD bills, each checked. S&P 500: the price index with Shiller's monthly dividends accrued daily to 1987, spliced onto the real total-return index from 1988 — 0.9996 daily correlation and a 6 bp/yr gap on the 36-year overlap, and the buy-and-hold row lands on the paper's. DAX: the performance index itself. Nikkei 225: the price index plus an annual dividend-yield proxy; with zero dividends the direction and the gap are unchanged. Bills from FRED (US daily; OECD/IMF 3-month rates for Germany and Japan).
Method
The authors' own jumpmodels package for the fit, cross-checked against an independent coordinate-descent implementation (identical on every refit); a vectorised form of their online dynamic programme (identical on every sampled window); the monthly Sharpe-maximising cross-validation of λ over a nine-value grid, 5 to 150, the paper's stated range.
Windows
S&P 500 and Nikkei 225 out-of-sample from 1990 as in the paper (the Nikkei from August, its first 3,000 feature days). The DAX's public history starts at the end of 1987, so twelve years of training plus eight of validation put its honest out-of-sample start at 2008 — the window is 2008–2023, stated and not stretched. All three extended to 2026-08 for the post-publication read.
Baselines
The claim is about regime identification, so it has to beat holding less and the oldest timing rule: a constant mix at the strategy's own average exposure, rebalanced daily, and a 200-day moving average with the same delay and costs. Neither is in the paper.
Deviations from the paper
  • S&P 500 total return: Bloomberg's index replaced by ^GSPC plus Shiller dividends before 1988 and the real ^SP500TR after; 0.9996 correlation, −6 bp/yr on the overlap, buy-and-hold reproduces the paper's row. Effect on the verdict: nil.
  • DAX window: public history begins 1987-12-30, so the out-of-sample window is 2008-02 → 2023-12 rather than 1990–2023. The paper's own Figure 5 shows the DAX strategy flat-to-up over 2008–2023; on public data it is worse than buy-and-hold there.
  • Nikkei dividends: a hand-tabulated annual yield proxy (±0.3 pp), accrued daily. The zero-dividend run keeps the direction and the gap (0.24 vs 0.10). The Nikkei window opens 1990-08, six months after the paper's.
  • Bills: German 3-month interbank and a Japanese T-bill-to-interbank splice replace GFD; a few basis points of level on a strategy that holds bills a third of the time.
  • λ grid {5, 10, 15, 25, 35, 50, 70, 100, 150} is ours (paper: 'typical values 5 to 150'); refit calendar (June and December month-ends), per-window standardisation and no outlier clipping follow the paper text. The grid-to-500 and 3σ-clipping variants are reported as fragility rungs.
  • HMM benchmark reproduced in reduced form only (refit every 21 days instead of daily, S&P 500 only): 0.58 against the paper's 0.54. Reported as reduced, never as the paper's benchmark.
  • Sample end pinned at 2026-08-31; Yahoo, FRED and Shiller vintage 2026-09-07. The reproduction notebook ships the exact frozen series.

ProvenanceBeta

Engine
v1
Blocks
5 new (public total-return index and bill adapter with the Shiller splice; jump-model features, package fit, from-scratch fit, vectorised online DP and monthly cross-validation; 0/1 timing backtest; paired block bootstrap; Henriksson-Merton timing regression), 3 reused (multiple testing, deflated Sharpe, PBO)
Data
Public: Yahoo ^GSPC / ^SP500TR / ^GDAXI / ^N225, Shiller monthly dividends, FRED DTB3 and OECD/IMF 3-month rates; vintage 2026-09-07; S&P and Nikkei 1970-01 → 2026-08, DAX 1988-01 → 2026-08
Source
arXiv 2402.05272 · Authors' jumpmodels package · Shiller S&P data · FRED DTB3
Tests
7 known-value tests (our DP ≡ the authors' package DP; online recursion ≡ predict_online; from-scratch fit ≡ package fit; feature and accounting values by hand; no look-ahead by truncation; one-λ grid composes to the fixed-λ strategy) + the peek ladder, circular-shift control, truncation identity and in-sample ceiling
Reproduce
download the notebook ↓
Validation Report · #363490

Downside Risk Reduction Using Regime-Switching Signals: A Statistical Jump Model Approach, 2024

One email opens this reproduction in full, and the rest of the library: ten systematic strategies, re-run and audited. Free.

Descriptive research, not investment advice. No card, no trial clock.
© MMXXVI MarketGraphs · descriptive research, not investment advice