Skip to content
← Back to MON — Monte Carlo

Operator brief · 391

What has to be true before a simulation may be called a benchmark.

The key idea

The distinction

A simulation is an output. A benchmark is an authority.

Generating paths is arithmetic and almost free; the moment a distribution is allowed to determine what may be risked, it has acquired authority over capital, and authority requires qualification. MARS therefore distinguishes sharply between a simulation someone ran and a benchmark the system answers to. The first is a picture. The second is a control input, and it has to earn that role by satisfying conditions the first does not. The elevation from picture to authority is worth marking explicitly, because it happens silently in most workflows: a chart gets generated, then consulted, then relied upon, and nobody ever asked whether it qualified.

The conditions

Four requirements, and all of them are refusable.

A distribution qualifies as a benchmark only when each of the following holds. Failing any one does not produce a worse benchmark — it produces something that is not one:

  • It resamples the operator's OWN recorded trades, not assumed parameters.
  • It is gate-aware — the allocator applies the same authority ladder the live account runs on.
  • Its sample is deep enough that the input distribution is not itself mostly noise.
  • Its window is current enough that the system it describes still exists.

Why gate-aware is the hard one

A fixed-risk simulation describes a system nobody is running.

The third condition is the one most simulations fail, and the failure is easy to miss. A conventional Monte Carlo applies constant risk across every path, which produces a clean distribution of a system that does not exist — because the live account compresses authority as drawdown deepens. Gate-aware allocation changes the shape materially: adverse paths de-risk as they descend, so the deep tail is thinner and the recovery profile is different. Benchmarking a governed account against an ungoverned simulation compares it to a stranger.

FigureFixed-risk simulation against gate-aware
15Fixedrisk13Gate-awareMedian path31Fixedrisk24Gate-awareP10 path44Fixedrisk29Gate-awareDeep taildrawdown depth (%)

Schematic. The same trade distribution, simulated with constant risk and with the live gate ladder applied.

The self-referential requirement

The benchmark has to be built from your trades, not a plausible trader's.

A simulation parameterised by assumed win rates and payoff ratios describes a hypothetical operator, and comparing live results against it measures the distance between you and an invention. The first condition exists to close that gap: the input is the recorded trade distribution, with its real hit rate, real magnitudes, real branch mix, and real friction. The benchmark then answers a question worth asking — how is this system performing against itself — rather than how it compares to a set of numbers someone typed in.

Staleness

A benchmark has a shelf life and nothing announces its expiry.

The fourth condition is the one that decays silently. A benchmark built from a window that no longer represents the system — because the method changed, the instrument changed, or the edge drifted — keeps producing confident bands describing a machine that has been retired. Nothing in the output signals this. It is why benchmark regeneration is scheduled rather than event-driven, and why structural drift readings are consulted alongside the fan: the diagnostics are what tell you the ruler needs remaking. A practical consequence is that the regeneration schedule should be written down alongside the benchmark itself, since a ruler with no stated expiry will be trusted indefinitely by an operator who has no reason to doubt it.

What it is not for

The benchmark sets expectations. It does not set targets.

A final boundary. The distribution describes what the system should reasonably produce; it is not a goal to be beaten, and treating the P90 path as an objective inverts the entire instrument. That inversion is common and expensive — it converts a tool for staying calm below median into a source of pressure to reach the top decile, which reliably produces exactly the oversizing the gate ladder exists to prevent. The ruler measures. Asking it to motivate breaks it. The same inversion explains a common complaint that the benchmark is demotivating: it is not motivating or demotivating, and an instrument being asked to do a job it was not built for will always seem to be doing it badly.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.