Historical replay and custom shocks

What history's worst market declines would have done to the portfolio, and a simple linear shock model.

What it does

For each of seven crises, from the dot-com bust to the 2025 tariff shock, the portfolio is bought at the S&P 500's peak and held to its trough, using each holding's actual daily returns over that window. The custom shock applies a move in the stock market and in interest rates through each holding's estimated sensitivities.

Why it is used

Volatility and VaR describe ordinary bad days. Crises are different: correlations rise and losses compound. Replaying real episodes shows the path, not just the endpoint, and needs no distributional assumption.

Inputs

  • Full price history of each holding and SPY.
  • For holdings that did not trade during a crisis: their beta to SPY from their first three years of trading.
  • For custom shocks: each holding's regression on SPY and IEF (7–10 year Treasuries) over the last three years.

Formulas

replay: R_p = Σ_i w_i · (P_i,trough / P_i,peak − 1) (buy and hold, starting weights) beta-scaled: r_i,t = β_i · r_SPY,t for holdings not yet trading custom shock: move_i = β_M,i · market_move + β_T,i · (−7.5 · Δyield)

Assumptions

  • Windows run from the S&P 500's closing high to its closing low; a holding could have fallen further on other dates.
  • No rebalancing inside the window.
  • Custom shocks are instantaneous and linear, and IEF's duration is taken as 7.5 years.

How to read the results

The comparison against the S&P 500 answers "compared to what?". Each holding is labeled with the method used: actual returns, or beta-scaled when it did not exist yet. The share of weight replayed with actual data is shown for every scenario; below 100%, treat the result as an estimate.

Limitations

  • Most ETFs did not exist in 2000, so portfolios of ETFs are largely beta-scaled in the dot-com scenario. Beta-scaling assumes the holding behaved like a leveraged S&P 500, which misses anything specific to it: bonds and gold, for example, rose in several of these crises.
  • Seven episodes are not a distribution. The next crisis will differ.
  • Custom shocks ignore changes in correlation during stress, which is when diversification tends to fail.

Where it can fail

  • For a holding with a short history, its estimated beta is noisy and the beta-scaled result inherits that noise.

Changes from the original version

DeanOS began as a personal tool. Rebuilding it for the public meant rechecking each model; these are the changes that came out of that.

  • The original version scaled SPY's move by each holding's beta and a hand-set sector multiplier for every holding. It now replays actual returns wherever they exist.
  • The 'recovery days' estimate (loss × 200) is removed; it had no basis.
  • Hardcoded custom scenarios are replaced by sliders with the assumptions shown.

Validation on current data

Share of each example portfolio's weight replayed with actual data, by scenario.

References

  • Basel Committee on Banking Supervision (2018). Stress testing principles.
  • Kupiec, P. (1998). Stress testing in a value at risk framework. Journal of Derivatives 6(1).

See it on a portfolio