kappi.me

ask AIs about kappi

Is My Trading Edge Real, or Is It Luck?

Thirty trades cannot tell a 55% edge from a coin flip. They can tell a 70% one. The sample size that settles the question is a property of your edge rather than of trading — 21 trades at a 70% win rate, 381 at 55% — and neither number helps if the record was assembled from memory afterwards.

How many trades does it take to know?

That depends on the size of the edge, and the gap is wider than most traders expect. The standard error on a win rate is sqrt(p(1−p)/n) and the 95% interval is roughly ±1.96 standard errors, so run it at two win rates rather than one:

Trades95% interval at p = 55%95% interval at p = 70%
3037.2% – 72.8%53.6% – 86.4%
10045.2% – 64.8%61.0% – 79.0%
20048.1% – 61.9%63.6% – 76.4%
1,00051.9% – 58.1%67.2% – 72.8%

At 55%, thirty trades are consistent with a real 37% and with a real 73% — the same sample contains a losing system and a spectacular one. At 70%, thirty trades have already excluded the coin flip.

The crossing point is the number worth knowing. Requiring the lower bound to clear 50% gives n > (1.96 × sqrt(p(1−p)) / (p − 0.5))²: 21 trades at a 70% win rate, 381 at 55%. Eighteen-fold apart, which is why "how many trades do I need?" has no general answer. A large edge is provable inside a month. A small one takes something closer to a year and a half of complete records — and the trades that would have proved it are the ones already behind you, which is the half that cannot be recovered later.

Does clearing the interval mean the strategy makes money?

No, and a high win rate is exactly where that gap hides. The interval settles whether the win rate is real. It says nothing about what the wins are worth. A trader risking $300 to make $100 breaks even at a 75% win rate — L / (W + L) — so a genuine, provable, statistically airtight 70% loses $20 a trade at that payoff. Across ten trades the seven winners bring in $700 and the three losers give back $900. The number that feels most like proof is the one best able to conceal a losing system, and win rate and payoff have to be read together or neither means anything.

What breaks the sample even when it is large?

Three things, and all three are properties of the record rather than the market:

  • Missing trades. A record that omits the trades you would rather not think about is not a small sample, it is a biased one, and more data does not fix bias. Honesty is not the failure point either — recall is. Nobody decides to omit the trade they took at 3:40pm on a Thursday; they simply do not think of it six weeks later, and nothing in the file says it is missing.
  • Edited intent. If the plan is written after the result, every trade appears to have gone roughly as intended, and the measured discipline is an artifact.
  • Selection across accounts or strategies. Running five approaches and evaluating the survivor is the same error as five traders posting only the winner's curve.

It is also why a streak carries no information. Eight wins in a row at 55% happens about 0.84% of the time on a given run — rare enough to feel like a signal, common enough that across 300 trades you will almost certainly see one, and it arrives feeling like the moment the system started working.

What should you do while the sample is small?

Size as though the edge might not be real, because it might not be. At a 10% edge, 20 units of risk gives a 1.8% chance of ruin and 10 units gives 13.4%; keeping the unit count high is how you stay solvent long enough to reach your own crossing point, whether that is 21 trades or 381.

The formal version is worse than most traders assume, and it has been quantified three times over. Bailey, Borwein, López de Prado and Zhu showed an excellent backtested Sharpe ratio is reachable after trying only a modest number of strategy configurations, so an unreported trial count is itself the finding.[1] Reviewing hundreds of published return predictors, Harvey, Liu and Zhu argued that a newly claimed factor should clear a t-statistic above 3.0 rather than the conventional 2.0[2] — your own strategy search is the same problem with a smaller sample. And following Taiwanese day traders from 1992 to 2006, under 1% predictably earned positive abnormal returns net of fees.[3] An edge is not impossible; it is rare enough that the null deserves the benefit of the doubt.

kappi is a trade recorder: you commit a trade before the fact, it is sealed on your device for a time-capsuled delay you choose, then kappi publishes it on a Merkle-anchored log. The seal is the whole mechanism — once a stop, a target and a thesis are sealed, revising them is no longer a private edit. $15/month, no free tier.

An append-only log has a property a spreadsheet cannot: what is absent is detectable. Every commit is in there, including the ones that went badly, because at the moment of sealing there was no way to know which those would be.

Sources

  1. Bailey, Borwein, López de Prado & Zhu, 'Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance', Notices of the AMS 61(5), 2014, 458 read 2026-08-16
  2. Harvey, Liu & Zhu, '… and the Cross-Section of Expected Returns', Review of Financial Studies 29(1), 2016, 5–68 — argues a newly claimed factor should clear a t-statistic above 3.0 read 2026-08-16
  3. Barber, Lee, Liu & Odean, 'Do Day Traders Rationally Learn About Their Ability?' — of Taiwanese day traders 1992–2006, under 1% predictably earned positive abnormal returns net of fees read 2026-08-16

Frequently asked questions

How many trades before I know my edge is real?

It depends on the edge. A 70% win rate clears a coin flip after 21 trades; a 55% win rate needs 381. At 55%, thirty trades give a 95% interval of 37.2% to 72.8% — it contains a coin flip and a spectacular system at once.

Is a 70% win rate proof of an edge?

It is provable far faster — the 95% interval at 70% already excludes 50% after 30 trades, running 53.6% to 86.4%. But it proves the win rate, not the profit: risking $300 to make $100 needs 75% to break even, so a real 70% still loses $20 a trade at that payoff.

Does a winning streak mean anything?

Very little. Eight wins in a row at a 55% win rate happens about 0.84% of the time on a given run, but across 300 trades you are likely to see one somewhere. Streaks are what randomness looks like.

Can more data fix a biased record?

No. Missing trades, intent written after the fact, and evaluating whichever strategy survived are all biases, not sampling error, and they do not shrink with sample size.

How should I size while I am still unsure?

As though the edge might not exist. At a 10% edge, 20 units of risk carries a 1.8% risk of ruin against 13.4% at 10 units — staying solvent is what buys you the sample size to find out.

Let's set some records

Broker-import journals prove what you did after the fact, from data you control. kappi timestamps what you said you would do, before you knew how it would turn out, on a record you cannot edit.

Start a verified track record — $15/mo

No free tier. Cancel any time.

Related