Praxis · #358063
Validation · #358063 · Cross-asset futures

Dynamic momentum learning

Trend-Following Strategies via Dynamic Momentum Learning, 2021

Bruno P. C. Levy & Hedibert F. Lopes · Insper / arXiv q-fin
The rule

Instead of trading the sign of past returns, let a model learn which momentum speed to trust: each month, for every futures contract, 127 small regressions — one per combination of look-back windows (1–12 months) — each predict the odds the contract rises next month; the strategy follows whichever regression has recently predicted best. Long if those odds ≥ 50%, else short, sized to equal risk. The pitch: it senses turning points that a fixed 12-month rule rides into the ground.

Overstated 1.230.49 claimed → measured Sharpe
1.23
Claimed SharpeDMS-TVP net of costs, 56 futures, 1980–2020 — a +52% lift over the naive 12-month rule's 0.81
0.49
Measured Sharpesame rule on the 39-contract public panel, 2003–2020; the naive baseline scores 0.21 — the lift is +274bp/yr at t = 0.85
−179bp
The lift since publication2020-10 → 2026-06: dynamic momentum 0.42, the naive rule it was meant to beat 0.60
Asset39 futures — all 25 paper commodities + US equity, Treasury, FX
StrategyDynamic-logistic classifier over momentum speeds
Period2003–2026, monthly
CostsPer-class bps + roll legs (paper's values cited, never printed)
BenchmarkNaive 12-month TSMOM, identical sizing
Assets tradedCL GC NG HG ZC KC LE ES ZN 6E…and 29 more front-month continuous chains. The paper's full 25-commodity sleeve is covered; its 9 non-US equity-index and 7 non-US bond futures have no free source and are absent — the panel skews toward the class the paper itself crowns.

The exact rules

Each month, per contractCompute momentum over 1, 2, 4, 6, 8, 10 and 12-month look-backs (through last month)
127 model combinationsEvery subset of those look-backs feeds a dynamic logistic regression — coefficients drift as random walks, discounted by λ ∈ {0.98, 0.99, 1} picked each step by realised predictive likelihood
Model averaging (DMA) or selection (DMS)Weight each model by its recent forecasting record (forgetting factor α ∈ {0.99, 1}); DMS trades the single best model's forecast — the headline config
Forecast P(next month up) ≥ 50%Go long one unit of risk: weight = (1/N) × 40% ÷ trailing EWMA volatility; below 50% → short the same size
Every month-endRebalance all positions; pay per-class transaction plus roll costs; portfolio presented at 10% annualised volatility

The backtest, re-run

Inside the model

Position mix
69% long / 31% short across 7,776 asset-months; average gross leverage 2.5× (p95 3.9×) from the 40%-vol sizing
Win rate
62% of months (in-sample window)
Skew / kurtosis
−0.38 / 5.7 monthly — mild left tail, nothing exotic
Best / worst month
+11.8% (May 2009, the crash-dodge paying) / −11.8% (Oct 2008, the crisis it didn't dodge)
Annual returns
2007 +29 · 2008 −15 · 2009 +30 · 2013–15 −2/−3/−7 · 2017 +12 · 2020 +8 · 2022 −0.2 · 2024 +8 — the naive rule is its mirror: +17 in 2008, −14 in 2009, +15 in 2022
Turnover
102%/mo one-way (paper: 101.3) — the signal flips on 25% of asset-months; the naive 12-month rule trades 61%/mo
The model's own vote
1-month momentum is the most-included predictor (inclusion probability 0.47) and 12-month the least (0.28) — the paper's Figure-2 pattern, reproduced
Accuracy
52.6% directional vs 51.1% naive — the paper's '~1% accuracy gain' claim is real; it just doesn't buy the advertised Sharpe

The validation ladder

C0ReplicateThe paper's structure replicates on independent data: every classifier config beats every naive config, selection beats averaging, commodities are the best class, and the 2009 momentum-crash escape lands almost to the number (naive −23.9% measured vs −25.3% claimed; DMS-TVP +25.2% vs +22.1%). The levels don't: 0.49 vs the claimed 1.23, with the naive baseline suffering the same haircut (0.21 vs 0.81) — era, panel and roll-bias, not implementation.
C1HonestyCosts are real but small at monthly speed: 83bp/yr of drag at our estimated per-class rates on 102%/mo turnover. The +274bp/yr lift over naive is a gross phenomenon that survives netting.
C2DeflateThe lift never clears significance: best t-statistic across all 18 classifier configs is 1.02, no multiple-testing survivor, and the paper itself never significance-tests its headline spread. Deflated Sharpe against the paper's own 61-row disclosed search grid: 0.58 — fail. PBO 0.25 passes: the config family is stable, just weak.
C3CrisisThe advertised skill is real but narrow: +19.8%/yr (Sharpe 1.59) through the 2009 momentum crash the paper showcases — and −5.7%/yr through the 2008 crisis itself, while the naive rule's shorts earned +17%. The model senses rebounds, not bear markets.
C5FrictionsCost-robust: the lift stays positive to ~5× our cost estimates (+164bp/yr at 5×). Frictions are not what separates 0.49 from 1.23.
C6DecayTrue post-publication window (2020-10 → 2026-06): the strategy itself holds (0.42 vs 0.485 in-sample) but the naive benchmark does better (0.60) — the claimed edge is −179bp/yr out of sample. The one survivor: the simpler constant-parameter DMS on commodities, 0.80 vs 0.56.
C7OriginalityGenuinely orthogonal: β 0.09, R² 0.01 against the naive strategy it replaces — a different return stream, not an overlay — with α 4.7%/yr at t = 1.79, just shy of significance. Against a long-only futures index: β 0.61, R² 0.37 — a substantial long tilt through a rising decade.
P.S.Roll-true re-runRe-run on our roll-adjusted reference panel — exchange daily settlements (CME Globex, via Databento) assembled into continuous series and ratio-adjusted at every roll, so returns are what a rolled position actually earns. Its history begins 2010, hence the 2013 trade start after the paper's 36-month training. Commodity sleeve, identical code: the naive benchmark improves to 0.34 — front-month chains had been destroying its roll-carry capture — while the classifier flips to −0.13; the lift becomes −472bp/yr (t = −1.48). A same-data control shows most of the flip is warmup fragility: under the paper's own 36-month training rule (without the 13-year warmup our 2000-start data silently provided), the classifier goes negative even on unchanged data. Both findings harden the verdict.
Mechanics replicateMagnitude replicatesStatistically realOriginal vs benchmarkCost-robustCrisis-robustHolds post-publication

How we rebuilt it

Data
Yahoo front-month continuous futures chains (39 contracts, 2000→today), monthly closes. The paper's Refinitiv/Datastream panel (56 roll-adjusted series from 1980) has no free source; front-month chains carry roll bias symmetrically across all strategies compared.
Method
McCormick-style dynamic logistic regression — one-Newton-step Laplace filter per model, per-step discount selection, Raftery model-forgetting — over all 127 look-back subsets; DMA weights by recent predictive record, DMS picks the winner; positions and sizing per the paper's Eq 18/19.
Universe
Portfolios trade only months with ≥10 eligible contracts (each contract enters 36 months after its data starts, per the paper), so the book starts 2003-09 — the paper's 56-contract panel never runs as thin as our 2000-02 tail.
Regime stated
The paper's fair sub-sample (its Table 6, 2010–2020) claims 1.10 vs 0.50; we measure 0.39 vs 0.10 on the matching window — direction confirmed, half the size, t = 0.62.
Deviations from the paper
  • Panel: 39/56 contracts (no free source for non-US equity-index/bond futures or SEK), 2003–2026 vs 1980–2020 — the naive baseline's own haircut (0.81 claimed → 0.21 measured) calibrates how much is era/panel rather than method.
  • Roll treatment: Yahoo chains are not roll-adjusted; absolute Sharpe levels carry the bias (measured per contract on our adjusted panel: natgas chains overstate +29.8%/yr, hogs +11.8, wheat +10.8). The roll-true re-run (P.S. rung) shows what changes on investable data.
  • Transaction costs: the paper adopts Baltas-Kosowski (2020) per-class values without printing them → estimated {commodity 3bp, financial 1bp} one-way + roll legs, swept 0–10× (conclusion cost-invariant).
  • Undisclosed state prior → C0 = 100·I; the lift moves +130 → +274bp/yr across C0 ∈ {1, 10, 100} — direction robust, size prior-sensitive.
  • Commodity-sleeve claimed values read off the paper's Figure 1 bars (±0.02) — the paper prints no per-class table.

ProvenanceBeta

Engine
v1
Blocks
3 new (Yahoo futures monthly panel, dynamic-logistic DMA/DMS classifier, vol-scaled TSMOM portfolio), 6 reused (metrics, G1/G2/G3 gates, regime, spanning)
Data
Yahoo front-month futures chains (run 10, primary) + roll-adjusted reference panel re-run (run 11, vintage databento_glbx_2026-07-17)
Source
arXiv 2106.08420 ↗ · Yahoo Finance futures ↗ · Baltas & Kosowski (2020) cost model ↗ · Databento CME Globex daily settlements (re-run panel) ↗
Tests
8 known-value tests (vectorised filter ≡ reference loop at 1e-8, look-ahead traps, sizing arithmetic) + 4 negative controls (signal shuffle, white-noise classifier, peek trap, prior sweep)
Reproduce
view code ↗
Validation Report · #358063

Trend-Following Strategies via Dynamic Momentum Learning, 2021

One email opens this reproduction in full, and the rest of the library: ten systematic strategies, re-run and audited. Free.

Descriptive research, not investment advice. No card, no trial clock.
© MMXXVI MarketGraphs · descriptive research, not investment advice