Hidden Markov model regimes
A four-state model of the S&P 500 that labels each period calm, normal, volatile or crisis, using only data available at the time.
What it does
A hidden Markov model assumes the market moves between a few unobserved states, each with its own typical returns and volatility, and switches between them with fixed probabilities. The model is fitted to three features of the S&P 500: its 20-day average return, 20-day volatility and 60-day average return.
Why it is used
Risk figures averaged over all market conditions hide how differently a portfolio behaves in calm and turbulent periods. Regimes give a data-driven way to split history and ask how this portfolio did in each.
Inputs
- Daily SPY returns since 2000.
- The portfolio's daily returns, only for the per-regime performance table.
Formulas
Assumptions
- Four states, each a multivariate normal distribution of the three features.
- Constant transition probabilities.
- Model parameters are estimated once on the full sample. The label at each date uses only features up to that date, but the parameters that turn features into labels have seen the whole history.
- States are named by their average 20-day volatility, lowest to highest.
How to read the results
The stacked chart shows, for each week, how probable each state looked at the time. The current reading never shows more than 95% confidence: the model's raw probabilities are often near 100%, which overstates how sure anyone can be about an unobservable state. Expected durations come from the transition probabilities.
Limitations
- Regimes are descriptive. A regime label says what recent volatility and trend look like, not what comes next.
- Features are backward-looking 20- and 60-day windows, so the model recognizes a change weeks after it starts.
- Four states is a modeling choice. More states fit better and mean less.
Where it can fail
- The fit depends on its random starting point. On the current data, the original model's fixed seed landed on a clearly worse fit (lower likelihood) whose labels agreed with the best fit on only about a third of days. The engine now fits from eight starting points, keeps the most likely, and reports how well the runners-up agree.
- The original model named states bull, bear, recovery and crisis by their average same-day return. On this data every non-crisis state has a positive average return, so "bear" meant the lowest of three positive returns and flipped between fits that were otherwise identical. The states differ mainly in volatility, and are now named that way.
- If the HMM fails to fit, a Gaussian mixture without time dynamics is used instead and labeled as such.
Changes from the original version
DeanOS began as a personal tool. Rebuilding it for the public meant rechecking each model; these are the changes that came out of that.
- Labels use forward-filtered probabilities instead of the most likely path and smoothed probabilities, which used future data.
- Best of eight random starts; stability across starts is reported.
- States renamed by volatility rank.
- A feature that was an exact multiple of another (annualized and daily 20-day volatility) was removed; it made the covariance matrices singular.
Validation on current data
Agreement between the best fit and the other random starts, from the nightly fit on the current snapshot.
References
- Hamilton, J. (1989). A new approach to the economic analysis of nonstationary time series and the business cycle. Econometrica 57(2).
- Rabiner, L. (1989). A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE 77(2).
- Ang, A. and Timmermann, A. (2012). Regime changes and financial markets. Annual Review of Financial Economics 4.