Methodology

What this site does

I take trading claims that circulate widely — the ones repeated in courses, threads and chat rooms until they sound like facts — and I test them. Each test is specified in advance, run once against the specified data, and written up. The result is published whichever way it falls. A claim that survives and a claim that collapses get the same treatment here, because I decide what to publish before I know which one I am holding.

Pre-registration

Before a test runs, I publish its specification with a date. That date is on the claim page and it does not move.

The specification fixes six things:

  • the instrument
  • the sample period
  • the entry and exit definitions
  • the benchmark
  • the statistic that decides the outcome
  • what result counts as each verdict

This is the part that matters most, so I want to be direct about why. A backtest has many small decisions in it. Which sessions to exclude. Where to put the stop. Which years to treat as representative. Each one is defensible on its own. Made after seeing the results, they stop being research decisions and become steering. The specification is not adjusted once the numbers are in. If I got the specification wrong, that is a correction, dated and visible, not a quiet edit.

The three verdicts

Holds
The effect appeared, at the size and significance the specification named in advance.
Does not hold
The effect did not appear, or appeared too small to matter after costs.
Inconclusive
The data cannot separate the two. The sample is too small, the estimate too noisy, or the result too unstable across subsamples to call.

Inconclusive is a published result, not a withheld one. It gets a page, a date and a write-up like any other verdict. It means the test ran and the answer is that we do not know.

I expect most tests to land there. Financial data is noisy and samples are short. Anyone testing honestly across many claims and reporting a clean verdict every time is either lucky or is deciding what counts as a result after the fact.

Testing standards

These apply to every test on the site.

  • Costs are modelled. Commission, spread and slippage are subtracted, and the assumptions behind them are stated on the claim page. Plenty of published edges are smaller than the cost of trading them.
  • Confidence intervals, not point estimates. A hit rate of 58% means nothing without a range around it. The interval is reported and the verdict depends on it.
  • Subsample stability. Results are broken out across time periods and across regimes. An effect that lives in two years out of ten is reported that way, not averaged into a single number.
  • Multiple testing is disclosed and adjusted for. If I test twenty variants, I say so and correct the thresholds. Reporting the best of twenty as though it were the only one is how noise gets published.
  • Benchmarked against buy-and-hold. The comparison is holding the instrument over the same period, not zero. Beating zero in a rising market is not an edge.
  • Effect size is reported. Significance says an effect is probably there. Size says whether it is worth acting on. Both are on the page.

Baselines

Every conditional statistic is reported against its unconditional equivalent. A number about what happens after a setup is meaningless without the number for what happens anyway.

Suppose 65% of swept levels reverse. That sounds like a strong result, and it is the form most claims take.

Now measure every level over the same horizon, swept or not, and suppose 63% of those reverse too. The edge is 2 percentage points, not 65. The setup is carrying almost none of the result. The market is.

Most retail claims omit the baseline. That omission is usually not dishonest — the conditional number is the interesting one to look up, and the unconditional one takes extra work to compute. But without it, a statistic cannot be read at all. Where a baseline is impractical to compute, I say so and treat the result as inconclusive rather than quoting the conditional number alone.

Code and corrections

The code for each test is published with it, along with the data source. If you want to check whether I did what the specification said, you should not have to take my word for it.

I will get things wrong. When I do, the correction goes on the claim page as a dated entry in a visible log, and the claim page says it has been corrected. The original error stays legible. Pages are not silently amended, because a record that can change without trace is not a record.