Invest
Portfolio ManagementSmart Beta strategyAlpha strategyIndividual Investment AccountCompany CashHow it works
Company
AboutFAQContactInsightsThe app
Book a consultation

ENSL

All insights

A good backtest starts the questions, it does not end them

Someone shows you an investment strategy. Its rules are clear, and the line on the chart rises steadily across five decades while convincingly beating the market. No forecasts and no story about a miraculous company, just rules and a long history. Where is the catch?

Most readers look for it in the rules. Perhaps they are too simple to work, or too complicated to trust. Yet the problem with a persuasive backtest usually lies in what the chart leaves out: which companies disappeared from the data, how many similar rules were tested and quietly discarded, why the period begins and ends where it does, and how much trading costs would have consumed. A good backtest starts the questions. It does not end them.

The book that put rules before stories

In What Works on Wall Street, James O'Shaughnessy asked a different question. Instead of asking who had the best story, he asked which stock-selection rules had actually worked in the past. He tested them systematically over a long history. The fourth edition went further, combining several measures of cheapness and adding checks on financial strength and earnings quality. Most honestly, the author states plainly that every tested strategy also has bad periods.

Later independent research supported his broad lessons, not his exact recipes. Peer-reviewed studies support the idea that a stock which looks cheap on several economically related measures is a better-founded candidate than one selected through a single ratio, and that quality deserves attention alongside price. For the foundation, see our explanation of value, momentum and quality. None of those studies independently confirmed the book's precise formulas, exact weights and exact number of portfolio holdings. That distinction matters. A broad principle found in many forms and markets is not the same as a recipe that works in one implementation. The book also shaped how we think about evidence at JonatanMars Invest: respectful of data and suspicious of recipes.

Five traps that flatter the past

Before looking at the result of any backtest, ask five questions about the data.

First, where are the losers? If a database contains only companies and funds that still exist today, failed ones have silently disappeared from history and the past looks safer than it was. This is survivorship bias. Brown, Goetzmann, Ibbotson and Ross showed in 1992 how excluding failed funds can create the appearance of persistently successful management.

Second, could an investor have known this at the time? A test is broken if a 2005 decision uses information published in 2006 or an accounting correction released later still. Ask when each data point became public and when the test used it.

Third, how many rules were tested before this winner appeared? Test thousands of combinations of measures, periods and thresholds, and some results will look exceptional through luck alone. Researchers such as Harvey, Liu and Zhu therefore propose much stricter statistical standards for new findings. More searching requires stronger evidence.

Fourth, why does the chart start and end where it does? A period beginning after a major fall and ending near a peak flatters almost any strategy. Demand the full available history and several different starting points.

Fifth, who could really have traded it? A result built on the smallest, hardest-to-access stocks, or on constant portfolio changes, behaves differently on paper than it does in a world with trading costs, taxes and limited liquidity.

Original, reproduced, robust and investable are different things

A simple ladder helps to separate the questions. Every backtest result sits on one of four levels, each with a different meaning.

Original, reproduced, robust and investable are four stages of scrutiny, each narrower than the last.
Original, reproduced, robust and investable are four stages of scrutiny, each narrower than the last.

An original result is what a researcher measured in one data sample. By itself it is a hypothesis, not proof.

A reproduced result means that an independent researcher using the same data and definitions obtains the same numbers. When Andrew Chen and Tom Zimmermann closely followed the original methods of published studies, they reproduced a large majority of the results reported as statistically significant in the original papers. Reproducibility is necessary, but it is a low bar.

A robust result survives changed conditions: different definitions, weighting methods, periods and markets. The picture becomes less friendly here. Hou, Xue and Zhang placed hundreds of published effects into a consistent, more conservative implementation that reduced the influence of the smallest stocks. Most failed even the conventional statistical threshold. An international study by Jensen, Kelly and Pedersen was more forgiving. It found that most effects also appear around the world and cluster into a smaller number of common themes. The studies disagree less than it first seems because they answer different questions. Reproducing an original result is not the same as surviving changed conditions. McLean and Pontiff, and Linnainmaa and Roberts, likewise showed that published strategies tend to weaken outside their original periods. Part of the weakness comes from testing many rules before publication; part comes after other investors begin exploiting the published opportunity.

An investable result is the last and strictest level. It asks what remains after trading costs, taxes, liquidity constraints and the bad periods an investor must endure. Novy-Marx and Velikov showed that costs reduce the results of every strategy. Strategies that changed holdings infrequently generally held up better, while few high-turnover strategies survived.

Each level reduces uncertainty. None removes it.

What every investor can demand

This leads to a standard that does not require specialist training. The rule should be written before the result, together with an economic reason for why it might work. The data should reflect what was known at the time and include companies that failed. The result should survive nearby versions of the rule and different periods. Evidence should also come from periods and markets the researcher did not use. Finally, show the net result after costs and include the bad periods, not only the most attractive slice. The next step is to build your own investment process around those questions.

One candid qualification is essential. Even a strategy that passes every test is not a promise. Markets adapt, definitions change and a genuine effect can lag for years. Investment values fluctuate and can fall despite a disciplined process. Evidence filters out weak ideas; it does not insure against loss.

At JonatanMars Invest, every assessment therefore starts with questions, not the curve. We publish new articles on the Insights page.

Frequently asked questions

What is a backtest?

A calculation of how an investment rule would have performed if used in the past. It shows how the rule fitted history, not how it will perform in the future.

What is survivorship bias?

An error in which a test includes only companies or funds that still exist and omits those that failed. Because the worst outcomes disappear, the test makes history look safer and more profitable than it really was.

Does a good historical result imply a good future return?

No. Research shows that published strategies tend to weaken outside their original data periods. A good historical result reduces uncertainty but does not remove it. Investments can fall, and an investor can receive less than they invested.

Why do studies of strategy replication reach such different conclusions?

Because they measure different things. One type asks whether the original calculation can be reproduced with the original method. Another asks whether the result survives stricter and more consistent conditions. The first gives encouraging answers; the second is much harsher. Both questions matter.

How does JonatanMars Invest use these findings?

As a filter. We write rules before seeing the result, require evidence from several periods and markets, account for costs and expect bad periods in advance. This is process discipline, not a promise of returns. Risk remains part of every investment.

What should I ask first when I see a persuasive chart?

Ask where the losers are. Does the dataset include companies and funds that failed during the period? Without an answer, the rest of the analysis has little value.

Sources


This is a marketing communication and general educational material, not personal investment advice.

Investing involves risk. The value of investments may fall as well as rise, and you may receive less than you invested. Past and any simulated or tested returns are not a reliable indicator of future returns. Tax treatment depends on personal circumstances and applicable law, both of which may change.

More insights

Stay in touch

A couple of useful ideas a month, or a proper conversation.