Trading Auditor

The six-control protocol

How the engine separates a technique that works from a statistical illusion — and what it refuses to claim.

We audit the technique, never the person. A number with no control is not a result: it is a drawing on a chart.

What the battery has done so far

These numbers come out of the archive itself on every publication — they are not typed into this page. 325 distinct claims, measured between 2017 and 2026.

6
passed
228
failed
91
inconclusive

A battery that passes almost everything is not auditing anything. Neither is one that fails everything — which is why the ones that passed stay published, by name and by number.

Before the six: the prediction comes first

Every audit starts with a hypothesis written and dated before the number exists. It is published on the card, with the filing date, next to the result. That is what separates an audit from a backtest: whoever writes the prediction after seeing the result is always right.

When the prediction is wrong, it stays published wrong. An archive in which the auditor never misses is an archive where he wrote it afterwards.

1. Invariant — where the result came from

Before looking at profit, the engine asks whether it is arithmetically plausible. What fails most here is not a bug: it is concentration. If half the gross profit comes from a handful of trades, the result is not an edge — it is a lottery that happened to pay inside the sample. Nobody can follow a technique like that, because the few trades that pay for everything are indistinguishable from the rest at the moment of entry.

It is the control that fails most in the archive: 90 of the 228 failed techniques die here, and 90 of them by concentration. The engine also checks outright impossibilities — more trades than bars, a mean dominated by a single outlier — but no technique in the archive has failed on those.

2. Cost — the toll the theory ignores

The broker's fee enters on both sides, plus the cost of carrying a position overnight. It is where most paper results evaporate: the edge exists, but it is smaller than the toll for capturing it.

Cost is also the dominant cause of the inconclusive verdict: in 87 of the 91 cases the engine could not show that the edge survives cost — nor that it dies. Saying 'I don't know' there is more honest than rounding to either side.

3. Placebo — the technique, or the tide?

The engine draws thousands of sham dates and trades on them, in the same assets over the same period. If the draw yields the same, the profit did not come from the rule: it came from the market having gone up (or down) for everyone.

It is the archive's second largest gravedigger. The placebo is the control that separates 'the technique works' from 'this year was good' — and the second explanation covers most of the pretty results in circulation.

4. Benchmark — the lazy alternative

Every technique's competitor is doing nothing: buying and holding the same assets over the same period. Beating zero is no achievement. A technique that demands daily attention and yields less than sleeping has failed, even while making money.

Failing here is the most uncomfortable case to explain, because the technique's own number is positive. The card publishes both side by side, so the reader can see the difference instead of taking the comparison on trust.

5. Out of sample — does it replicate where it was not tuned?

The measurement is split into halves the technique did not choose: half the assets against the other, the liquid against the illiquid, and four slices of time. A real effect shows up across the partitions; an effect tuned to fit shows up only where it was tuned.

It is also where decay shows: techniques that worked in the older slices and stop in the recent ones. The card publishes the slices one by one, with each one's number of episodes — because a slice with six episodes supports no conclusion at all.

6. Multiple testing — the price of trying many times

Whoever tests fifty variations finds one that 'works' by chance. The engine applies the Benjamini-Hochberg correction over the whole archive, not just over one technique's variations.

The consequence is uncomfortable and it is declared: the bar rises with every claim published. Today a technique in crypto has to deliver about 4.5% per trade to survive this correction. Publishing more measurements makes asserting anything more expensive — including for us.

What FAILED does not mean

FAILED is not 'loses money'. It is 'there is no effect above what this instrument can detect' — and that number is measured and published: 0.694% per trade in crypto, 0.092% in forex. A smaller effect may exist; we simply cannot assert it, and we do not.

You can check it, and that is what it is for

Each report's identifier is the hash of its own content: any edit after publication breaks the verification. Every card offers the complete record for download — the parameters, the universe, the period, the six controls and the dated hypothesis. Checking it does not depend on trusting us.

The three verdicts

  • PASSEDIt survived all six. This is not investment advice nor a promise of profit: it is the record that the claim withstood the most hostile battery we know how to build, over the period and universe declared on the card.
  • FAILEDIt failed at least one control. The card says which one and by how much. It does not mean the technique always loses money — it means that if there is an edge, it is smaller than the declared detection floor.
  • INCONCLUSIVEThe engine could not decide. Almost always because the edge and the cost are of similar size, or because the independent episodes are too few. It is the answer an honest instrument gives more often than we would like.

No seal is for life

  • The bar rises. When we find a new statistical illusion and tighten a control, the archive is re-audited. Techniques that passed before can fall — and they fall publicly.
  • The market changes. A technique that worked from 2017 to 2020 may have stopped. We re-test with new data in cycles.
  • The history stays. A re-test does not erase the earlier result: it is born next to it, and the card shows when the verdict changed. Finding out that a technique stopped working is the product, not a defect of it.