10,000 Trials Independent software studio

The lines that did not survive.

Every result below went the wrong way. They are published because a studio that shows you only its winners has told you nothing about how it decides anything — the losing ones are the evidence that the method is real. Each entry gives the number that killed it.

12 lines closed 58 figures published data back to 2011 longest run 15 years through August 2026

How to spot one yourself

Twelve strategies died on this page and they did not die in twelve different ways. Each one is tagged with the thing that killed it, and eight tags cover all twelve. These are the questions to ask of any claim, including ours.

  1. What was it compared against?4 of 12

    A result measured against a control that guarantees the answer. Any rule that sits in cash part of the time shows a smaller drawdown; any stock at a yearly high is in an uptrend already. The comparison has to be matched on whatever the rule changes by construction, or it is measuring the construction.

    Killed: The golden cross, Buying the 52-week high, Stop-losses, Selling into earnings

  2. Is it better than doing nothing?2 of 12

    A strategy priced on its own terms, never beside simply holding the thing. Income, premium collected and win rates all look impressive in isolation. The number that decides it is the one the index did over the same window.

    Killed: The Wheel, Short-volatility carry

  3. Does it win often, or win well?1 of 12

    A high strike rate presented as the result. Win rate and return are separate quantities and can move in opposite directions — a system can be genuinely undefeated and still be the worst available use of the money.

    Killed: Martingale averaging down

  4. When were the candidates chosen?1 of 12

    The names, pairs or parameters picked using data from the period being tested. Choose on one window and trade the next and the same rule returns nothing. If selection did not happen strictly before the tested window, the test is of the selection.

    Killed: Pairs trading

  5. Would this be true anyway?1 of 12

    A finding that restates something arithmetic. Implied volatility always falls after an announcement. A price finishing nearer a strike than it started happens about half the time. Neither is evidence of anything.

    Killed: Expiry pinning

  6. How many rows is it standing on?1 of 12

    An average carried by a handful of extreme observations. Drop a small percentage of the data and watch: if the sign flips, there was no effect, there was a tail.

    Killed: Options flow as a conditioner

  7. Is that before or after costs?1 of 12

    A gross result quoted as though it were reachable. Spreads, commission and the cost of rebalancing come out of the same number, and for small or short-dated instruments they are routinely larger than the effect.

    Killed: The published anomaly literature

  8. How many things were tried?1 of 12

    The best cell of a large search, reported as if it were the only one tested. Search enough noise and something always clears the bar — so the bar has to be raised to match the number of attempts, or the search re-run on shuffled data to see what luck alone produces.

    Killed: Searching harder

The five words this page cannot avoid

a t
One number for how far a result sits from ordinary luck. Below about 2, the result is the sort of thing chance produces all the time. Above about 3, it starts to be worth believing.
p = 0.17
The chance that pure luck would have handed us a result this good. At 0.17 that is roughly one time in six — far too often to call it a discovery.
significance
The bar a result has to clear before we will believe it is real rather than lucky. Not clearing it does not prove the idea is worthless; it means the test could not tell it apart from noise.
permuted
We shuffle the history so any real effect is destroyed, then run the identical search on the scrambled version. Whatever it finds is what searching pure noise looks like — the score a real finding has to beat.
the sample
How many independent events the number rests on. Thirty trades and thirty thousand trades can produce the same average and mean completely different things.

14 years · 2012–2026 · 9 configurations, every leg at the bid

The Wheel

Whether selling cash-secured puts, taking assignment, then selling covered calls back out — the most widely taught options income strategy there is — beats simply owning the thing.

configurations tested
9
configurations that won
0
premium collected, best case
$39,000 on $5,900 of capital
ending value
$47,019
buy-and-hold ending value
$69,513

Closed August 2026 Nine of nine lose, by 0.31 to 7.26 points of annual return.

Three roots, three strike distances, fourteen years, every leg sold at the bid and settled at expiry. The most flattering configuration collected 660% of the account in premium and still finished $22,000 behind holding the index. That gap is the whole lesson: the income is real and large and completely beside the point, because it is the price received for the upside that was sold away. Reporting premium collected as if it were profit is how the strategy is marketed, and it is exactly what a comparison against the alternative exposes.

10 years · 2016–2026 · 8,182 campaigns · one-year horizon

Martingale averaging down

Whether doubling into a falling position until the average cost is recovered — a ladder that only needs a small bounce to clear — actually works. It does. That is the problem.

campaigns simulated
8,182 across 228 names
win rate, best ladder
99.8%
its annual return
0.111%
owning the same names
15.11%
index version
100.0% win rate, 0.041%/yr

Closed August 2026 Undefeated, and the worst available use of the capital.

Every rung added to the ladder raises the win rate and lowers the return, monotonically, in every configuration tested. On indices the strongest form of the defence is literally true — across 177 campaigns it never lost a single one, and its worst case was still positive — while returning 0.041% a year against 14.37% for holding the same index with the same money. The reserve is the cost: capital committed to funding rungs that mostly never get used is capital not invested. A strategy can have a perfect record and still be the worst thing you could have done.

15 years · 2011–2026 · timing; selection from 2016

The golden cross

Whether holding while the 50-day average sits above the 200-day, and sitting in cash otherwise, captures the upside and dodges the crashes as advertised.

max drawdown, S&P, signal
−34.1%
max drawdown, S&P, buy-and-hold
−34.1%
random books beating it, S&P
75% on return, 77% on return/drawdown
small caps
drawdown deepened, −42.3% to −51.9%
roots where it beat random timing
1 of 3

Closed August 2026 On the index everyone quotes, it delivered no crash protection at all.

The control is what decides this one. Any rule that sits in cash a quarter of the time shows a smaller drawdown whether or not it means anything, so the signal has to be measured against random timing matched on time-in-market, not against buy-and-hold. Against 200 such random books it lost on two roots out of three, and on small caps it was worse than useless — it deepened the drawdown while cutting the return by two thirds. The headline number is the S&P drawdown: identical to buy-and-hold, to the tenth of a point. The signal's entire selling point, on its most-cited market, is absent.

10 years · 2016–2026 · 12,954 breakouts

Buying the 52-week high

Whether a stock making a new yearly high keeps going — the breakout trade, and the single most photographed chart pattern in retail education.

breakouts measured
12,954 at the shortest horizon
three-month return
+2.432% (t 10.23)
random days, same uptrend
+3.367%
difference
−0.935% (t −3.15)
horizons where it beat the control
0 of 3

Closed August 2026 Significantly worse than buying the same stock on an unremarkable day.

In isolation the arm looks superb: a t of 10.23 over three months is the sort of number a breakout course is built on. But a stock at a yearly high is in an uptrend by construction, so the only honest control is random days drawn from that same stock's uptrend, at the same firing rate and holding period. Those ordinary days returned more, at every horizon, with t between −3.15 and −4.10. The signal does not merely add nothing. It subtracts — you are paying attention to the chart in order to buy at a worse moment than if you had not looked.

10 years · 2016–2026 · 9,120 entries

Stop-losses

Whether a protective stop improves anything once you compare it against the same amount of risk rather than against a larger position.

entries, shared across arms
9,120
no stop, at matched volatility
14.39%
tightest stop, matched
3.35%
widest stop, matched
6.61%
worst case, no stop vs stopped
−80.5% against −82.9%

Closed August 2026 They cost more than they save, and they did not hold at the extreme either.

The two tests people usually run are both rigged. 'It reduced my drawdown' is trivially true, because a stop makes the position smaller in expectation — so does trading smaller. And grading in R is arithmetic rather than evidence: drift is fixed in percent, so drift measured in R scales as one over R and tighter stops win mechanically without predicting anything. Any study concluding that tighter stops have better expectancy in R has discovered division. Measured in percent on matched capital and then compared at matched volatility, every stop lost. The worst case got worse too, because gaps go through stops.

10 years · 2016–2026 · ~26,000 candidate pairs

Pairs trading

Whether trading the spread between two stocks that move together works — the classic version, with the pairs chosen the way you would actually have to choose them.

candidate pairs searched
~26,000
alpha, chosen honestly
−0.04% (t −0.08)
alpha, random pairs
−0.17% (t −0.11)
alpha, pairs chosen with hindsight
+2.73% (t 4.84)
Sharpe with hindsight
1.69

Closed August 2026 Zero — unless you pick the pairs knowing how they turned out.

Same rule, same threshold, same universe. The only thing that changes between the three rows is when the pairs were chosen. Formed on one window and traded on the next, the strategy returns nothing, at a t of −0.08 — not weak, not marginal, zero. Worse for the method: carefully selected pairs are statistically indistinguishable from randomly assembled ones. And chosen using the whole sample, the identical rule produces a Sharpe of 1.69 and an alpha at t = 4.84. That last row is the one that gets published, and the gap between it and the first row is the entire value of insisting selection happens before the tested window.

14 years · 2012–2026 · 108 announcements

Selling into earnings

Whether the implied-volatility collapse after a scheduled announcement can be sold. The crush is mechanical and calendar-timed, so unlike most premium it is not obviously just compensation for risk.

announcements traded
57 and 51 across two names
earnings arm
−0.345% per trade (t −2.75)
same option on a random day
−0.104%
win rate, earnings
36.8%

Closed August 2026 Worse than selling the identical option on an ordinary day.

The question that matters was never whether implied volatility falls after the announcement — it always does, that is arithmetic. It is whether the premium collected exceeds the move the event produces. It does not. Selling the put the session before and buying it back the session after lost more than three times what the same trade lost on matched non-earnings windows in the same name, at the same tenor, over the same holding length. The crush is real, it is priced, and what you are actually being paid for is standing in front of the announcement.

1 year · 2025–2026 · 50 expirations

Expiry pinning

Whether price is pulled toward the strike where option open interest is heaviest as expiration arrives — the 'max pain' claim.

expirations tested
50
closes nearer max pain
50.0% (z +0.00)
mean miss, max pain
1.647%
mean miss, guessing the open
0.527%
expiries max pain won
16.0%

Closed August 2026 Three times worse than doing nothing.

Max pain was computed from the full open-interest ladder as of the prior session, so nothing about the outcome could leak into the prediction. Price finished nearer the max-pain strike than it started on exactly half of expirations, and the average distance to it grew through the session rather than shrinking. The decisive comparison is not against chance but against the laziest available alternative: predicting the day closes where it opened is three times more accurate. Fifty expirations is a modest sample, but a predictor that loses to guessing the open on 84% of occasions is not a marginal effect waiting for more data.

8 years · 2018–2026 · 2,130 sessions

Short-volatility carry

Whether the standing premium in VIX futures — the curve sits in contango on most days — could be harvested as an edge in its own right.

sample
2,130 sessions
beta to the S&P
0.96
alpha
+1.87%/yr
alpha, t
0.22
vs buy-and-hold
12.10% against 14.36%

Closed August 2026 It was beta.

The only version a cash account can actually trade moves one-for-one with the S&P and returns less than simply holding it: worse compound return, worse Sharpe, deeper drawdown, an alpha indistinguishable from zero. The premium is real, but it is payment for crash risk — earned through calm years and handed back in the volatility events — rather than a skill anyone is being compensated for. A second finding made the test itself awkward: the condition is true on 91.3% of sessions, which is not a condition, it is 'be invested'. A matched-random control drawn against a signal that is on 91% of the time selects nearly the same days and has almost no power to reject anything, so the control it had already passed meant nothing.

52 sessions · May-August 2026

Options flow as a conditioner

Whether unusual options activity, read off the live tape, said anything about where the underlying went next.

cells tested
~440 across four lanes
survivors
0
concentration
7% of rows held 84% of the effect
top ten names removed
t −0.46
against a market-signed benchmark
+0.026% (t +0.63)

Closed August 2026 There was no baseline left to condition on.

The effect everything was being conditioned against turned out to be a fat right tail in a handful of high-volatility names rather than a mean: removing 7% of the rows sent the result negative. Leave-one-out resampling, the standard check for exactly this, was blind to it and passed. The benchmark was wrong as well. Each pattern's outcome had been signed in that pattern's own favoured direction, so bullish and bearish rows cancelled inside the control and it removed only about a third of the market's own drift. Measured against a market-signed benchmark instead, the headline moved by a factor of 4.8 — to nothing.

120 anomalies · 452 replications · 4 reviews

The published anomaly literature

Whether any of the several hundred documented cross-sectional anomalies could clear the return bar a small account has to clear.

the bar
15.85%/yr
average published effect
7.9%/yr gross, in-sample
net, post-publication, post-2005
0.96%/yr
cost of rebalancing monthly
1.08%/yr
162 anomalies, out of sample
+0.14%/mo before borrow, −0.01%/mo after

Closed August 2026 A level mismatch, not bad luck.

The average published anomaly delivers about half the required return before a single cost is charged, and every adjustment after that points downward. The decisive pair is the middle two figures: rebalancing monthly costs more than the average anomaly nets. Being small helps with exactly one of the four cost terms, price impact, and destroys the thing the effect depends on — a cross-sectional edge of a few basis points a month exists only as an average across hundreds of names, and held in the handful of positions a small account allows it is a coin flip with a thumb on the scale. Because the gap is arithmetic rather than statistical, reading further into the literature cannot close it. That is why this entry closes a source rather than a strategy.

600 permutations of the same 52 sessions

Searching harder

Whether a wider sweep would eventually turn up something that worked — the assumption underneath every parameter search, and the last thing left to try.

median best t, 20 cells searched
2.16
median best t, 188 cells searched
2.69
best t available anywhere
3.319
its p-value
0.067

Closed August 2026 The ceiling of the search space is not significant.

The data was permuted to destroy any real effect and the same sweep re-run six hundred times. A best-of-search t around 2.5 is the median outcome of that — the ordinary result of searching noise, not evidence of anything. A t of 2.8 found after sixty attempts works out to roughly p = 0.17, which is unremarkable. And the single best result obtainable anywhere in the legal search space, on the real data rather than the permuted data, still does not clear significance. That closes the strategy of looking harder: the space does not contain a finding, so no amount of further search can produce one.

Nothing here was abandoned because it got boring. Each one was measured, and the measurement said no — which is a result, and cost real time to obtain. A line that dies on a number is worth more than one that survives on a hunch. Lines still open are not listed: publishing those would be telling you where to look, not what was found.

Ten minutes here. Fourteen years the other way.

The Wheel took fourteen years of data to settle. The martingale took 8,182 campaigns. The 52-week high took 12,954 breakouts. Every one of them is answerable the other way too — one position at a time, with your own money, over a career.

That is what this page is for. If it saved you a detour, consider putting something toward the next one.

Fund the next one $5 · $15 · $50