Shuffle Test on Triad & Global Navigator

I wanted an unbiased write up of this test which was performed on Triad & Global Navigator - and also I wanted it to be able to explain the details clearly, so I asked Claude Opus 5.0 to give it's thoughts on this test which was performed, a Tier 2 Monte Carlo Simulation.

2026-08-06-monte-carlo-simulation-tier-2

Do Triad & Global Navigator actually work, or did it just get lucky?

Results of the shuffle test, August 2026


The question

Every backtest has the same problem. You look at what a strategy did over 46 years, you see a good number, and you cannot tell how much of it was the strategy working and how much was the particular sequence of months that happened to occur.

This test was built to answer exactly that question, and it does it in a way that is easier to grasp than most statistics.

How the test works

Imagine every month from 1980 to 2026 written on its own index card. Each card records what every ETF did that month: US stocks up 3.1%, gold down 0.8%, treasuries up 0.4%, and so on. Whole months stay together on one card, so everything that happened at the same time stays happening at the same time.

Now shuffle the deck.

The cards are the same. The returns are the same. Every relationship between assets within a month survives untouched, so when stocks crashed and gold rallied in the same month, they still do. The only thing destroyed is the order.

Then re-run the entire strategy on the shuffled deck, from scratch, exactly as it was run on the real history. All the momentum signals get recalculated, all the buy and sell decisions get remade, and a new 46-year track record comes out the other end.

Do that 1,000 times, and you have 1,000 alternate histories in which momentum cannot possibly work, because there is no longer any connection between one month and the next.

This matters because momentum is the entire premise. Triad and Global Navigator both pick investments based on what has been going up recently. That only works if recent performance tells you something about near-future performance. Shuffling the deck breaks that link completely. The strategies should fail.

The real question is: how much better did the real history do than the 1,000 fake ones?

Reading the results table

Five columns, in plain terms.

Real is what the strategy actually did over the true, unshuffled history. This is the number you already know from the site.

Null p5, p50, p95 describe the 1,000 shuffled runs. "Null" is just the statistician's word for "the world where the strategy has no edge." The three numbers mark out that world:

So p5 to p95 is the range of ordinary outcomes when there is no skill involved, and p50 is the middle of it.

Percentile is where the real result landed among the fakes. 100% means the real history beat all 1,000 shuffled versions.

p-value is the one that gets quoted in research, and it means something specific and useful: if the strategy truly had no edge, how often would blind luck have produced a result this good? A p-value of 0.05 means once in 20. A p-value of 0.001 means once in 1,000.

Lower is stronger. Below 0.05 is the conventional bar for "this is unlikely to be luck." Below 0.01 is a strong result.


Results: Triad

MeasureRealNull p5p50p95Percentilep-value
CAGR (annual return)15.53%8.60%10.15%11.60%100%0.001
Max drawdown-13.75%-31.37%-22.05%-16.81%99.7%0.004
UPI (return per unit of pain)4.420.540.961.46100%0.001
Volatility10.15%10.03%10.56%11.07%9.6%0.097

In words. With the calendar shuffled, Triad typically returned about 10.2% a year. Even a lucky shuffle only reached 11.6%. The real Triad returned 15.53%, which beat all 1,000 shuffled versions.

The drawdown result is just as striking. Shuffled Triad typically fell 22% from peak to trough at its worst, and in unlucky orderings fell over 31%. Real Triad's worst decline was 13.75%, shallower than 997 of the 1,000.

Results: Global Navigator

MeasureRealNull p5p50p95Percentilep-value
CAGR (annual return)14.53%7.13%9.52%11.87%99.9%0.002
Max drawdown-22.33%-50.71%-34.13%-24.84%98.7%0.014
UPI (return per unit of pain)1.900.190.470.86100%0.001
Volatility12.29%13.35%14.06%14.68%0%0.001

In words. Shuffled Global Navigator typically returned 9.5% a year with a 34% worst decline. The real one returned 14.53% with a 22% worst decline.

Note the volatility row, which is the most striking single number in either table. Real GN's volatility of 12.29% is lower than all 1,000 shuffled versions, every one of which came in above 13.3%. Nothing about the returns themselves explains this. It exists only because GN moved to safety at the right moments, and shuffling the calendar destroys the ability to do that.


What this means

Both strategies pass, and they pass decisively. On return and on return-adjusted-for-pain, both landed outside the entire range of 1,000 no-edge simulations. The odds of that happening by chance are below 1 in 1,000. The results are not an artifact of a lucky sequence of months.

But they earn their results differently, and that is genuinely interesting.

Global Navigator's volatility sits below every single shuffled run. It is a defensive strategy in a measurable sense: its value comes substantially from being out of the market at the right times, and that ability vanishes the moment the calendar is scrambled.

Triad's volatility is statistically indistinguishable from the shuffled versions. Its ride is no smoother than chance would produce. Everything Triad gains comes from picking better assets and avoiding the deepest holes, not from running a calmer portfolio.

Two different routes to a strong result, and neither would have been visible from the headline numbers alone.

One result that is weaker than the others

Drawdown protection holds up less well than return under stress-testing.

A second version of the test was run in which the cards were shuffled in blocks of six months rather than one at a time, which leaves some short-term momentum intact rather than destroying all of it. That is a harder test to beat, because the fake histories are no longer completely skill-free.

Under that harder test the return advantage survives comfortably for both strategies, but the drawdown advantage no longer clears the conventional bar (p-values rise to 0.082 for Triad and 0.062 for Global Navigator).

The honest reading: the evidence that these strategies produce superior returns is very strong. The evidence that they protect against deep declines is good but less bulletproof, and depends more heavily on markets continuing to trend the way they have.


What this test does not prove

Worth being clear about, because this kind of result is easy to over-read.

It does not predict the future. It establishes that the historical record is unlikely to be a fluke of sequencing. It says nothing about whether the market conditions that produced it will continue.

It does not account for strategy selection. DMS runs roughly 19 strategies. Testing the two that performed well, in isolation, ignores the fact that they were partly chosen because they performed well. Correcting for that requires a different test.

It uses raw returns. Trading costs, taxes, and slippage are not included on either side. This affects the real and shuffled results roughly equally, so the comparison stays fair, but the absolute numbers are optimistic.

Early history relies on reconstructed data. Several of the funds these strategies hold did not exist in the 1980s and 1990s. Those years are built from index and proxy data, which is well documented but is not the same as observed fund returns.


The takeaways

  1. The historical results are not sequence luck. Both strategies beat essentially all 1,000 alternate histories on return and on risk-adjusted return.
  1. The momentum signals carry real information. When the link between one month and the next is severed, both strategies collapse toward ordinary. That collapse is the proof that the signals were doing work.
  1. Global Navigator is genuinely defensive. Its below-every-simulation volatility is the clearest evidence in either table that market timing added value rather than merely appearing to.
  1. Triad's edge is selection, not smoothness. It does not deliver a calmer ride than chance. It delivers a better destination.
  1. Trust the return evidence more than the drawdown evidence. Both are positive; only the return result survives the harder version of the test intact.
  1. This validates the past, not the future. It is one important question answered well, and it is not the same question as "will this keep working."

Method: 1,000 full row-wise permutations of the ETF return matrix, with the complete strategy re-run on each. The Python implementation reproduces every published monthly return for both strategies across all 560 months to within 6 parts in a billion, so the simulated results come from the same logic that produces the live figures.