CAGR stands for Compound Annual Growth Rate. It answers a simple question: if a strategy grew at a perfectly steady rate every year, what would that annual rate be? It is the standard way to compare returns across strategies with different track record lengths, because raw cumulative return is misleading - a strategy with 20 years of history and a strategy with 3 years of history need a common yardstick.
The formula compounds the total return over the full period and then scales it to one year:
CAGR = (Ending Value / Starting Value) ^ (1 / Years) - 1
A strategy that turned $10,000 into $18,000 over 5 years has a CAGR of about 12.5%, regardless of how bumpy or smooth the ride was along the way.
How the current month is handled
The most recent month shown on the site is usually still in progress. Its allocation is published daily, but its return is not final until the month closes, and you will see it labeled accordingly.
When a date range ends in an in-progress month, DMS counts that month as a fraction of a month rather than a whole one, based on how many days have elapsed. This matters more than it might sound. On the third day of a month, only a sliver of that month's return exists. Annualizing it as though it represented a full month would badly distort the result - a small partial gain would be scaled up as if it had taken a month to earn, inflating CAGR, and a small partial loss would do the reverse.
The practical effect:
A consequence worth knowing: a CAGR quoted for a range ending in the current month is a moving number. It will shift day to day as the month accumulates, and it will settle once the month closes. That is the calculation working correctly, not a data problem.
The same fractional logic applies at the other end of the track record. If a strategy went live partway through a month, its first month is scaled to the portion of the month it was actually invested, so it is never credited with a full month of return it did not earn.
Why the number on your screen may differ from one you noted earlier
Three global toggles change the CAGR displayed:
If a figure does not match what you recorded previously, check these first - the underlying data has not changed, only the lens.
What CAGR doesn't tell you
CAGR is a two-point measurement. It sees only where the strategy started and where it ended, and it is blind to everything in between. Two strategies can post an identical CAGR while one lost half its value and clawed back while the other barely dipped. The ride is invisible to this number.
For a fuller picture, pair it with Max Drawdown, Ulcer Index, and the Sortino ratio.
Maximum Drawdown (Max DD) is the largest peak-to-trough decline a strategy has experienced, measured from a peak (high-water mark) down to the lowest point that followed it. It is expressed as a negative percentage. A Max DD of -15% means the strategy fell 15% from its peak before recovering.
It answers a pointed question: what is the worst loss an investor in this strategy would have had to endure, and how bad did it get before things turned around?
How it is calculated
DMS computes Max DD on a month-end basis. Walking forward through the equity curve, it tracks the highest value reached so far and measures how far below that high-water mark each subsequent month falls. The deepest of those declines is the Max DD.
Because DMS strategies are evaluated monthly, the figure reflects month-end to month-end moves. Intra-month, a strategy will have dipped lower than any month-end close shows, so a daily-measured drawdown would be deeper. Every strategy and benchmark on the site is measured the same way, so comparisons between them remain fair even though all of them understate the intra-month extreme.
The global toggles feed into it. Turning on Inflation Adjusted Returns, Include Trading Friction, or Taxable Account changes the return stream the curve is built from, and the Max DD moves with it.
What it is useful for
Max DD is one of the most honest stress-tests available for a strategy. It tells you: this is the real-world pain that an investor would have experienced at the worst moment in the strategy's recorded history. Paired with the Max DD Recovery metric, which shows how long it took to get back to the prior peak, it gives a sense of both depth and duration of the worst episode.
For comparison, the S&P 500 has experienced drawdowns exceeding 50%, and a traditional 60/40 portfolio has seen drawdowns around 32%. DMS strategies are designed with low drawdowns as a central goal, not an afterthought.
What it does not tell you
This is the most important thing to understand about Max DD: it is a historical figure, not a guarantee of maximum future drawdowns.
The Max DD shown is the worst drawdown the strategy has experienced to date, based on the specific market environments in the backtest and live history. It is not a promise, a guarantee, or a prediction of the worst that could ever happen. As Meb Faber so aptly has said: your largest drawdown is still to come.
Max DD and the date range selector
The Max DD shown reflects the currently selected date range, not the full history. Narrow the range to the last five years and you will see the worst drawdown within those five years, which may be far smaller than the all-time figure if the worst episode fell outside the window. For the all-time number, use the maximum available range for the strategy.
There is a subtlety here worth knowing. The high-water mark resets at the start of your selected range. The calculation has no memory of anything before it, so the opening month is treated as the first peak.
If your range begins partway into a decline, the real peak that preceded it is invisible, and the drawdown gets measured from an already-depressed starting value. A range beginning at a market bottom will make almost any strategy look serene. When you want to understand risk rather than study a specific episode, start from the full range.
Pairing Max DD with other risk metrics
Max DD captures the single worst episode but says nothing about how often or how persistently a strategy draws down. Two strategies can share the same Max DD while feeling very different to hold - one might recover quickly, another might grind sideways for years. For a fuller picture of drawdown behavior, pair Max DD with:
Two of the detailed metrics answer a retirement question rather than an investment one: given this strategy's actual history, how much could you have pulled out every year without running out of money?
Both follow the method William Bengen introduced in 1994, which is stricter than it first sounds. The number you see is not what one retirement would have supported. It is the worst outcome across every 30-year retirement the strategy's history contains.
How the cohorts work
Take a strategy with 46 years of monthly returns. A retirement beginning in January 1985 and running 30 years is one cohort. February 1985 is another. March 1985 is another. Roll that window forward one month at a time and a 46-year record yields roughly 200 distinct 30-year retirements, each with its own sequence of good and bad years.
Every one of them is tested. The published figure is the worst of them.
That distinction matters more than it might appear. A single run starting at the beginning of a strategy's record tests exactly one entry point, and an early-1980s start happens to be close to the most favourable moment in the modern record. Reporting that number would describe a lucky retirement rather than a safe rate. The minimum across all cohorts is what the word "safe" is doing.
Because a 30-year cohort needs 30 years of data, and because a minimum is only meaningful if there are enough cohorts to take a minimum of, both metrics require at least 35 years of history. Strategies with less show N/A rather than a number built from too few retirements.
Safe Withdrawal Rate (SWR)
SWR is the largest annual withdrawal, as a percentage of your starting balance, that would have carried a full 30-year retirement through to the end without the account reaching zero, beginning in any month on record.
The withdrawal is held constant in real terms. A withdrawal fixed in dollars shrinks every year in what it actually buys, so holding it constant in purchasing power is what makes the figure describe a standard of living rather than a dollar amount.
Perpetual Withdrawal Rate (PWR)
PWR asks a stricter question over the same cohorts: what could you have withdrawn while leaving the account, in real terms, at least as large at the end of 30 years as it was at the start? SWR permits you to spend the balance down toward zero. PWR does not touch the principal.
PWR is therefore always at or below SWR. The gap is usually small, often a few tenths of a percentage point, and it is smaller for strategies that compound faster. That is not a rounding artefact: when a portfolio grows a great deal over 30 years, the extra draw that spending down the principal would buy you is small next to what the growth itself already supports.
They ignore the date range
This is the one behaviour that surprises people, and it is deliberate.
Every other figure in Detailed Metrics answers "over the period you selected." These two do not. They always use the strategy's full history, whatever range is on screen. A sustainable withdrawal rate is a property of the strategy, not of the window you happen to be looking at, and tying it to the range would mean the number vanished on any view shorter than 35 years. The row labels carry "(full history)" so the exception is visible rather than silent.
What the toggles do
Where they appear
How to read them
What these numbers are not
They are not a retirement plan and not a recommendation. They describe a strategy's historical capacity to support withdrawals across the retirements its own record contains.
A real retirement introduces everything the model leaves out: a specific horizon rather than a fixed 30 years, taxes on the withdrawals themselves, Social Security and other income, spending that is lumpy rather than smooth, and the near certainty that you would change your behaviour after a bad year rather than mechanically withdrawing the same amount into a falling account.
An equity curve shows one path: the particular sequence of months that happened. It cannot tell you how much of the result came from the strategy and how much came from the order those months arrived in.
Range of Outcomes answers that. Switch the equity chart to it using the toggle above the chart, and instead of one line you get a spread of paths the same strategy could plausibly have produced, with the actual result drawn over the top.
How it is built
The strategy's own monthly returns are resampled 2,000 times to produce 2,000 alternative 46-year histories, each using the same pool of months in a different arrangement. The percentile bands show where those 2,000 paths sit at each point in time.
The resampling is done in blocks of consecutive months, not one month at a time, and that detail is the difference between a useful chart and a misleading one.
Real market declines are made of bad months arriving in a row. Draw months independently and those runs get scattered apart, so simulated portfolios recover between shocks and never experience a proper crash. The effect is not subtle. On a test series containing one sustained decline, independent-month resampling reported a typical worst drawdown of 34% where block resampling on the identical data reported 88%. Independent draws would have understated the risk by a factor of more than two.
Block resampling keeps those runs intact. Block lengths are random, averaging up to 24 months on a full history and scaling down on shorter ranges so there are always enough distinct blocks for genuine variety. The applied block length is shown beneath the chart.
Reading the chart
The bands start narrow and fan out. That is the point of the picture. Early on, sequence has had little chance to matter; over decades it compounds into an enormous spread.
The four figures below the chart
Median Outcome is the ending value of a $10,000 starting balance at the 50th percentile, with the 5th and 95th beneath it. The gap between those two is usually startling, and it is worth sitting with. Every path used the same returns.
Median CAGR is the same idea in annualised terms.
Median Max Drawdown is the deepest peak-to-trough fall at the 50th percentile, with the 95th shown as the unlucky case. Expect this to be worse than the strategy's actual historical drawdown. The realised figure is one draw; the simulation asks what the same months could have done in a crueller order, and the answer is usually "quite a bit worse."
Actual vs Range is where the real backtest landed among the 2,000. This one is routinely misread, so it is worth being explicit: a middling number here is the correct and expected result. The simulation is built from the strategy's own returns, so it is centred on them by construction. A figure near the 50th percentile means the machinery is working. It is not a measure of skill, and a high number would not be good news, it would be a sign something was wrong upstream.
What responds to what
Range of Outcomes uses the selected date range, so narrowing the range changes both the bands and the figures. In this it differs from the Safe and Perpetual Withdrawal Rates, which always use full history.
The Inflation Adjusted, Trading Friction and Taxable Account toggles all feed the return series being resampled, so the bands respond to them as the ordinary equity chart does.
The view needs at least 60 months in the selected range.
The bands do not move between visits. The simulation uses a fixed starting point for its random number generator, so the same strategy over the same range always produces the same picture. Without that, the bands would shift slightly every time you touched a toggle and the chart would look untrustworthy for no reason.
What this tells you
That the strategy's historical result was, or was not, heavily dependent on the order in which its months arrived. If the actual path sits comfortably inside the bands and the median lands near it, sequence luck is not what produced the backtest.
It also gives you a realistic sense of dispersion. A strategy with a 15% historical CAGR whose 5th-to-95th band spans 8% to 22% is telling you something a single number cannot.
What it does not tell you
It is not a forecast. The distribution is centred on returns the strategy has already earned. It assumes those returns keep coming from the same process. If markets change, nothing in this simulation would know.
Two thousand paths are not two thousand pieces of evidence. They are one dataset rearranged 2,000 times. The apparent precision is real in the sense that the arithmetic is exact, and misleading in the sense that it all rests on one historical record. If that record is optimistic, every percentile shown is optimistic by the same amount, and nothing inside the method can detect it.
It cannot validate the strategy. Resampling takes the edge as given and only reshuffles it. It answers "was this sequence luck," which is a different and easier question than "does this strategy work." Establishing the second requires a test where the strategy is allowed to fail, which resampling is not.
It does not correct for having chosen this strategy. DMS publishes many strategies. Looking at a strong one in isolation, however rigorously, does not account for the fact that it stands out partly because it performed well.
Why it is not called "Monte Carlo"
It is a Monte Carlo simulation, and the tooltip says so. But the label invites a particular misreading, that thousands of simulations amount to thousands of independent observations, when they are one dataset restated many times. "Range of Outcomes" describes what is actually on screen: the spread of results consistent with this strategy's historical behaviour, including the unfavourable ones.
Maximum Drawdown tells you how bad the single worst moment was. It says nothing about whether a strategy spent one month underwater or eleven years. The Ulcer Index closes that gap.
Ulcer Index
The Ulcer Index measures the depth and the duration of every drawdown across a period, not just the deepest one.
It is computed by walking the equity curve month by month, recording how far below the previous high-water mark each month sits, squaring that figure, and taking the square root of the average across all months. Months at a new high contribute zero. Months deep underwater contribute a great deal, because squaring them punishes depth disproportionately.
The name is literal. It is meant to approximate how much stress holding the strategy would have caused. Lower is better, and unlike most statistics on this site there is no theoretical maximum, only comparisons between strategies over the same window.
Two things follow from the construction:
Ulcer Performance Index (UPI)
UPI turns the Ulcer Index into a risk-adjusted return measure:
UPI = (CAGR - risk-free rate) / Ulcer Index
The numerator is what the strategy earned above cash. The denominator is how much discomfort it caused getting there. Higher is better.
The structure is the same idea as the Sharpe ratio, with one substitution that matters. Sharpe divides excess return by standard deviation, which treats upside volatility as risk. UPI divides by drawdown pain, which counts only the downside and counts prolonged recoveries as worse than quick ones. For a strategy designed around limiting drawdowns rather than limiting volatility, UPI is usually the more informative of the two.
Why UPI is our headline risk-adjusted number
Standard deviation punishes a strategy for a strong upside month exactly as hard as for a weak one. Nobody experiences those two months the same way. Drawdown-based measures line up much more closely with what actually causes an investor to abandon a strategy, which is the failure mode that destroys more returns than any market decline does.
An important caveat: UPI is not leverage-invariant
It is sometimes assumed that a risk-adjusted ratio like UPI stays roughly constant when you scale a position up or down, so that a 3x version of an asset would show similar UPI to the unlevered version. Our own measurement says otherwise, decisively.
Measured against daily data for an unlevered S&P 500 fund and its 2x and 3x counterparts, buy-and-hold UPI at 3x falls to roughly 27% of the unlevered figure. The Ulcer Index does not scale linearly with leverage; across seven leveraged asset-class families it scales at approximately the leverage factor raised to the power 1.3, and that exponent was stable across sub-periods.
The practical consequence: do not compare the UPI of a leveraged strategy against an unleveraged one and conclude the leverage was free. Leverage degrades UPI structurally, before any question of skill or timing enters. A leveraged strategy that holds its UPI near its unleveraged parent's has done something genuinely difficult, and the comparison to make is against the same strategy at the same leverage, not across leverage levels.
How to read these numbers
What they do not tell you
Neither statistic knows anything about why a strategy drew down, whether the conditions that caused it are likely to repeat, or what a drawdown outside the historical record might look like. A strategy with an excellent Ulcer Index has been comfortable to hold across the period tested. That is a real and useful thing to know, and it is not a forecast.
Both answer the same question in slightly different ways: how much return did this strategy earn for the risk it took? Raw return alone cannot distinguish a strategy that earned 12% smoothly from one that earned 12% through violent swings.
Sharpe ratio
Sharpe = (average monthly return above cash / standard deviation of those excess returns) x the square root of 12
The numerator strips out what you could have earned sitting in cash. The denominator measures how much the monthly returns scattered around their own average. The square root of 12 annualizes a figure computed from monthly data.
Higher is better. As a rough guide, above 1.0 is good and above 2.0 is unusual over a long period, though these rules of thumb depend heavily on the era and the asset class.
DMS uses the arithmetic mean of monthly returns here, not the compound annual growth rate. This is the standard construction and it is what makes the figure comparable to Sharpe ratios published elsewhere. A version built on CAGR runs systematically lower, by roughly half the annualized variance, purely as an artifact of the formula rather than anything about the strategy.
Sortino ratio
Sortino keeps the same shape and changes what counts as risk. Instead of the standard deviation of all returns, it uses the deviation of only the losing months.
Sortino = (average monthly return / deviation of months below zero) x the square root of 12
The reasoning is straightforward. Standard deviation treats a surprise gain as risk, identical to a surprise loss of the same size. Nobody experiences it that way. Sortino counts only the outcomes that actually hurt.
The threshold here is zero, not the risk-free rate, and it is zero on both sides of the calculation. The minimum acceptable return is "do not lose money," and the numerator is measured against that same standard. Using one threshold in the numerator and a different one in the denominator produces a figure that is not comparable to anything, which is worth knowing if you are checking our numbers against another source.
Sortino is essentially always higher than Sharpe for the same strategy, since the denominator is built from a subset of the same months and the numerator is not reduced by the cash rate. The two are not comparable to each other in absolute terms. Compare Sharpe against Sharpe and Sortino against Sortino.
The gap between them is informative in itself. A strategy whose Sortino greatly exceeds its Sharpe has volatility concentrated on the upside, which is exactly what you want. A strategy where the two sit close together has volatility distributed evenly in both directions.
The risk-free rate
Sharpe needs a risk-free rate. DMS uses the actual return on cash over the period you have selected, taken from the cash series in the return data rather than from a fixed assumption. When cash data is unavailable it falls back to 4% per year.
This is the honest approach, and it has a consequence worth understanding: the bar moves across eras. In the early 1980s cash yielded close to 10%, so a strategy needed to earn well into double digits before its excess return was even positive. Through the 2010s cash yielded nearly nothing and almost any positive return counted as excess. A Sharpe ratio from the 1980s and one from the 2010s are not measuring against the same standard.
Sortino, using a zero threshold, is unaffected by this. That is one of its advantages when comparing across long periods with very different interest rate regimes.
When the Inflation Adjusted toggle is on, the risk-free rate is deflated along with the strategy returns, so the Sharpe comparison stays internally consistent rather than measuring a real return against a nominal benchmark.
Comparing our figures against other sites
If a ratio here differs from one you have seen elsewhere for a similar strategy, the cause is usually one of these rather than a disagreement about the underlying returns:
Which one to use
For DMS strategies specifically, we lead with the Ulcer Performance Index rather than either of these. UPI divides excess return by drawdown pain instead of by volatility, and for strategies designed to limit drawdowns rather than to limit volatility, it captures the design goal more directly.
Sharpe and Sortino are here because they are the standard vocabulary of the field, and because a strategy that looks good on one measure and poor on another is telling you something worth investigating.
Limitations shared by both
Almost every statistic on this site is computed over the date range you have selected, not over the strategy's full history. Move the range and the numbers move with it. This is intended, but a few of the behaviors surprise people, and one of them can be genuinely misleading if you do not know about it.
The general rule
CAGR, Max Drawdown, Ulcer Index, UPI, Sharpe, Sortino, standard deviation, alpha, beta, and the withdrawal rates are all window-scoped. Each is recomputed from the months inside your selection and nothing outside it.
So a strategy showing a 14% CAGR over its full history and 9% over the last five years is not contradicting itself. Those are two different measurements of two different periods.
The drawdown clock restarts
This one deserves particular attention.
Max Drawdown measures the decline from a running high-water mark. That high-water mark resets to the first month of your selected range. The calculation has no memory of anything before your window opens.
If your range begins partway into a decline, the true peak that preceded it is invisible, and the drawdown is measured from an already-depressed value. A range that starts at a market bottom will make nearly any strategy look serene, because the calculation never sees the fall that created the bottom.
The practical consequence: when you are trying to understand risk rather than study a particular episode, use the full available range. Narrow windows are for examining a specific period, not for judging how much pain a strategy can inflict.
The risk-free rate changes with the window
Sharpe, Sortino, and UPI all measure return above cash, and DMS uses the actual cash return over your selected period rather than a fixed assumption.
Cash yielded close to 10% in the early 1980s and nearly nothing through the 2010s. A window covering the first is holding the strategy to a far higher standard than a window covering the second. Risk-adjusted figures from very different eras are not directly comparable, even for the same strategy.
Some things do not change with the window
A few figures are computed over the strategy's entire history regardless of your selection:
Short windows are noisy
A statistic computed over 24 months is a much weaker claim than the same statistic over 300 months, and the site does not visually distinguish between them.
Some measures refuse to compute below a minimum. The withdrawal rates require ten years. The Tax Profile requires twelve months of allocation data. Most others will happily return a number from a very short window, and that number deserves proportionally less trust.
Comparing strategies fairly
When comparing two strategies, make sure the window covers a period both actually lived through. A strategy that launched in 2015 compared against one going back to 1979 over the maximum range is not a comparison of strategies. It is a comparison of eras.
The comparison views handle the common cases by aligning the period, but it is worth checking the start dates yourself when a result looks surprising.
Quick checklist when a number looks wrong
Most surprises resolve at one of those five.