Skip to content
← Back to Benchmark Review

Operator brief · 443

The benchmark is built from your own rules, so beating it is not a discovery.

The key idea

The distinction

This benchmark is a mirror, not a rival.

An index benchmark measures a manager against an alternative use of the capital, and beating it is the entire objective. The MARS benchmark does nothing of the kind. It is generated by simulating the operator's own configuration — the same branch probabilities, the same tier structure, the same gate ladder — across many thousands of paths. The comparison it supports is between the live account and the distribution of outcomes that this exact system produces. Nobody else is in the picture.

FigureLive equity inside a self-generated envelope
Simulated P75Simulated medianSimulated P25Live accountweeksequity

Schematic. The bands are the operator's own rules, simulated. The line is the live account.

What outperformance means here

Sitting high in the fan has three explanations, and only one is good.

A live curve running above the simulated median is not evidence of skill in the way an index-beating return would be. It has three candidate causes. The account may have had a favourable draw, which is expected — roughly half of all paths sit above the median by construction. The configuration may have deviated from what was simulated, taking more risk than the model assumed. Or the model's inputs may have been conservative, and the live edge is genuinely better than the branch probabilities that were fed in. Only the third is a finding, and it is the least likely of the three.

The real question

Placement asks whether the live account still belongs to this distribution.

Once outperformance is understood as uninformative, the benchmark's actual purpose becomes clear: it is a membership test. The question is not how far above or below the median the account sits, but whether it is behaving like a path drawn from this distribution at all. A curve wandering inside the fan is a system running as modelled, wherever in the fan it happens to be. A curve outside the adverse bands, or one whose drawdown behaviour has no counterpart anywhere in the simulated set, is telling you the live system has stopped being the system that was simulated.

The asymmetry

Underperformance is informative in a way outperformance is not.

The two directions do not carry equal weight, and the reason is structural. A live curve far below the simulated bands is difficult to explain by luck alone once enough evidence has accumulated, and points at something real, and it does so with a force that a single month never has: execution decay, friction growth, a regime the inputs did not anticipate, or branch probabilities that were estimated optimistically. A curve far above is explained by luck very easily, since favourable draws exist in the distribution by construction. The benchmark is therefore a more sensitive instrument for detecting trouble than for confirming success — which is the correct bias for a governance tool.

The circularity to watch

A mirror inherits every flaw in what it reflects.

Because the benchmark is generated from the operator's own stated configuration, an error in those inputs propagates straight into the ruler. Optimistic branch probabilities produce an optimistic envelope, against which a mediocre live account will place comfortably and look healthy. The benchmark cannot detect this, because it has no independent view of the truth — it is describing the model it was given. This is why the inputs are validated separately, and why a favourable placement is never on its own sufficient evidence that a system is working. The circularity is manageable precisely because it is known: the inputs are audited on their own terms, by instruments that do not derive from them.

What it is worth

A self-referential ruler is still the only one that fits.

None of this diminishes the instrument. A discretionary trading system has no meaningful peer group and no index that describes what it is attempting, so an external benchmark would be measuring against something irrelevant. Simulating the operator's own rules produces the only comparison that is actually about their system — provided everyone understands that its output is a membership verdict rather than a competitive one, and that placing well is the absence of a warning rather than the presence of an achievement. Read that way, a benchmark review is a short exercise most weeks and a genuinely important one occasionally — the correct distribution of attention for a governance instrument, and the opposite of how performance comparisons are usually consumed.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.