SledgeKey Backtests
We ran a Piotroski-style financial-strength screen honestly. It lost to the index.
Joseph Piotroski's 2000 paper asked a narrow question about cheap stocks. If you sort the market by price to book and buy the cheapest slice, you end up holding a lot of companies that are cheap for good reason, so he proposed a nine-point scorecard covering profitability, leverage, liquidity, and operating efficiency, and showed that the high-scoring end of the cheap bucket did far better than the low-scoring end. The scorecard has been part of the value investor's toolkit ever since, and it is one of the most requested things people want to test.
We tested a version of that idea over nearly ten years, January 2017 to September 2026. The screen buys companies trading at or below 1.5 times book value that also clear four financial-strength tests at the same time: positive return on assets, positive operating cash flow, debt at or below half of equity, and a current ratio of at least 1.5. It holds the twenty largest qualifying companies in equal weight and rebalances once a year. What makes this run different from most published Piotroski backtests is the universe. Every company listed on each rebalance date was eligible, including the ones that were later acquired, sanctioned off the exchange, or delisted for any other reason, and any holding that left the market was booked out at its last traded price instead of being erased from the record.
Over this window the screen turned $100,000 into $241,447, a return of 9.52% a year. An S&P 500 index fund did far better over the same stretch, at 15.28% a year, and it got there with a much shallower worst-case drawdown.
Growth of $100,000 · 2017–2026
The Piotroski-style screen against the S&P 500, both starting from $100,000 in January 2017, with delisted companies kept in the strategy the whole way through.
View the data as a table (calendar year-end values)
| Year | Piotroski-style screen | S&P 500 |
|---|---|---|
| 2017 | $113,939 | $119,995 |
| 2018 | $109,011 | $127,361 |
| 2019 | $116,694 | $148,109 |
| 2020 | $104,203 | $175,813 |
| 2021 | $125,386 | $219,434 |
| 2022 | $126,042 | $201,434 |
| 2023 | $138,785 | $230,582 |
| 2024 | $169,678 | $306,717 |
| 2025 | $188,811 | $350,488 |
| 2026 | $241,447 | $397,034 |
The exact screen
The full configuration is below exactly as it ran, so anyone who wants to reproduce the result or argue with it can start from the same table we did.
| Value filter | Price to book of 1.5 or lower |
| Profitability | Return on assets of 0% or higher (trailing twelve months) |
| Cash generation | Operating cash flow of zero or higher (trailing twelve months) |
| Leverage | Debt to equity of 0.5 or lower |
| Liquidity | Current ratio of 1.5 or higher |
| Market cap | $1 billion and up |
| Selection | The 20 qualifying companies with the largest market cap at each rebalance |
| Weighting | Equal weight, 10% maximum position size |
| Rebalance | Once a year (10 rebalances over the window) |
| Transaction cost | 0.10% per trade ($1,595 in modeled costs over the run) |
| Initial capital | $100,000 |
| Benchmark | SPY, the S&P 500, over the same window |
| Universe | NYSE + NASDAQ operating companies (no SPACs, REITs, ETFs, or funds). Point-in-time: eligibility at each rebalance reflects the companies listed on that date. |
| Delisting treatment | Any holding that later delisted was booked out at its frozen last traded price, never dropped from the history. |
One thing deserves to be said plainly, because it is the most common mistake in backtests that carry Piotroski's name. This is a financial-strength screen built in his spirit, and it is not the nine-point F-Score itself. Only two of his nine points are plain levels, positive net income and positive operating cash flow, and a third compares operating cash flow against net income within the same year. The other six ask whether a company improved on its own prior year: return on assets, leverage, the current ratio, gross margin, asset turnover, and share count. A screen filters on levels, so the run above uses his value anchor plus level-based stand-ins for his profitability, leverage, and liquidity tests, then stops there. Read the result as evidence about that screen rather than a verdict on the published F-Score.
The result against the index
| Metric | Piotroski-style screen | S&P 500 |
|---|---|---|
| Total return | 141.45% | 297.03% |
| Annual return (CAGR) | 9.52% | 15.28% |
| Final value | $241,447 | $397,034 |
| Sharpe ratio | 0.43 | 0.81 |
| Volatility (ann.) | 19.80% | 16.18% |
| Max drawdown | -48.89% | -33.72% |
| Calmar ratio | 0.19 | 0.45 |
| Winning months | 72 of 116 | 77 of 116 |
| Total trades | 278 | n/a |
| Avg holding period | 562 days | n/a |
| Annual turnover | 121.6% | n/a |
The screen won 72 of its 116 months to the index's 77, and its losing months were deeper: an average of -4.6% against SPY's -3.7%, with a worst month of -26.0% against -16.4%. That is what a high-volatility strategy looks like from the inside, and it is why the Sharpe ratio comes in at 0.43 against the index's 0.81.
A note on survivorship. Run this same screen the way most free backtesters quietly do, against only the companies still listed today, and it reports 12.58% a year instead of 9.52%. That extra 3.06 points a year is survivorship bias, phantom performance worth $74,005 on a $100,000 stake, earned by a universe that quietly forgets its casualties. It does not change the verdict against the index this decade, but it changes how far behind the strategy looks, from nearly six points a year to under three, and on other screens the bias runs the other way entirely. We pulled the mechanism apart in detail on the midcap value backtest.
Something else is visible once you have run several strategies through the same harness. The Magic Formula screen we published earlier shows no survivorship gap at all, because the twenty largest cheap, profitable companies never included one that later delisted. This screen, which shops further down into low price-to-book territory, is overstated by 3.06 points a year in a survivor-only universe. Our midcap value screen shows the bias running the other way, with the honest run 1.51 points a year ahead of the survivor-only one, because the companies it lost were mostly bought out at a premium. The three together make one point. How much survivorship bias distorts a backtest, and in which direction, depends on how a screen's holdings left the market, and a survivor-only database has no way to tell you.
The companies it held that later left the market
Fourteen of the names this screen held went on to leave the exchange, positions a survivor-only backtest never gets to take. Most of them were bought out, which is the ordinary fate of a cheap company with a clean balance sheet, and a few were foreign issuers whose US listings ended for policy or corporate reasons. The honest run had to live through every one of those exits, because the companies were in the book on the day they delisted.
The exits carrying no reason are recent delistings whose circumstances we have not confirmed to our own satisfaction, so we have left them unlabeled rather than publish a reason we cannot stand behind. They are in the run and they are in the numbers above either way, which is the part that affects the result.
What this backtest does not prove
A page like this earns its credibility in this section, so here is what the numbers above do not establish.
The strategy lost to the index, and it was a rougher ride. Nine and a half percent a year for nearly ten years is a respectable absolute result, and it still fell well short of simply owning the S&P 500, which returned close to six points more a year with a drawdown some fifteen points shallower and lower volatility throughout. Anyone reading this page as a market-beating system should keep looking. What the run shows is that financial strength expressed as a set of thresholds, applied to cheap large companies, did not clear the index over this particular decade.
This is a levels proxy, and the F-Score is mostly a change-based score. As described above, six of the nine original points ask whether a company improved on last year, which no screen filter can express. A faithful F-Score implementation would hold a different set of companies and could land anywhere relative to this result, so treat the two as cousins.
One window, one configuration. This is a single run of nearly ten years across an era that punished value and rewarded megacap growth, the exact stretch where the S&P was hardest to beat. Different dates, market-cap bands, or thresholds will produce different results, and the size of the survivorship gap shifts with them too.
The largest-cap tiebreak is a real choice. When more than twenty companies clear all five filters, this run takes the twenty largest. That keeps the portfolio liquid and tradable, and it also tilts away from the small, deeply cheap names where Piotroski found his strongest results. A version that ranked by cheapness instead would be a different strategy with a different answer.
Frozen-price booking is conservative but imperfect. When a company delists, the run books the position out at its last traded price. For an acquisition that lands near the deal price, and for a forced delisting under sanctions the real proceeds to a retail holder can be much worse than the frozen mark. The modeled 0.10% per trade covers commissions and typical slippage at large-cap liquidity, and real execution in a stressed market runs worse than any flat assumption.
You can run this exact screen, or your own version of it, on the same survivorship-free point-in-time data and see how it holds up with the delisted names left in.
Run your own backtest · free, no card required