Skip to content
← Back to Variant Comparison

Operator brief · 438

A variant comparison is only as fair as what was held constant.

The key idea

The premise

A comparison is a controlled experiment or it is nothing.

The purpose of comparing variants is to isolate one thing: the effect of the management structure itself. That isolation only exists if every other input is identical. The moment the two runs differ in some second way — a different random draw, a different stretch of history, a different number of trades — the measured gap contains both the variant's effect and that second difference, mixed together and impossible to separate afterwards. The output still looks like a comparison. It just no longer answers the question it appears to answer.

FigureWhat must match, and what the mismatch costs
InputHeld constantIf it drifts
SeedSame draw both runsGap is partly luck
Trade sampleIdentical evidenceDifferent questions
PeriodSame windowRegime, not variant
Path countEqual run depthUnequal precision

The seed

Two variants on two different draws are two anecdotes.

Monte Carlo output is generated from a random draw, and the draw carries its own luck. Run variant A and variant B against different seeds and part of the gap between them is the difference in the draws rather than the difference in the structures. On the same seed both variants meet the same sequence of outcomes, so whatever separates them is attributable to how each one managed those outcomes. This is the single cheapest control in the entire exercise and the one most often skipped, because skipping it produces no error message. The run either was or was not seeded identically, and the output looks the same either way — which is why the seed belongs in the run log rather than in anyone's memory of what they did that afternoon.

The subtler drift

The sample changes underneath you without anyone deciding to change it.

Seed discipline is well understood; sample drift is not, and it is more common. Variants are frequently compared across weeks — one profiled in March, another in June — and the trade evidence has grown in between. The second variant is then evaluated against a larger and partly different record, which quietly advantages or penalises it for reasons that have nothing to do with its management rules. If a comparison spans time, the older variant must be re-run on today's evidence before the two results are placed beside each other. This is the most common way a comparison silently decays: nothing was done wrong at either point, and the invalidity is created entirely by the gap between them.

What may legitimately differ

Exactly one thing: the variant's own rules.

The rule has a clean statement. Everything that describes the world is held identical; everything that describes the variant is allowed to differ. Partial-taking behaviour, break-even movement, runner management, and the quota structure are the variant, and those are the whole point of the exercise. The seed, the trade sample, the period, the path count, and the tier assumptions describe the world both variants live in. Any of those moving invalidates the reading, and the invalidation is silent — the numbers arrive looking exactly as authoritative as they would have.

Recording it

A comparison that was not written down cannot be trusted later.

The controls have to be recorded alongside the result, because a comparison is consulted long after the conditions of the run have been forgotten. Six weeks on, a table showing that variant B outperformed by some margin is worthless without the seed, the sample size, and the period beside it — nobody can tell whether it was a controlled finding or two loose runs. The run log exists for exactly this: it makes a stale comparison identifiable as stale rather than quietly reusable.

The honest limit

Even a perfect comparison ranks two candidates on one draw.

Controls make the comparison valid; they do not make it conclusive. A fair test still tells you how these variants handled this particular set of futures, and a small margin between them is well within the range that a different seed could reverse. The practical reading is that a large, stable gap across a well-controlled run is a finding, and a narrow one is a tie that should be broken on grounds other than the number — usually on which structure the operator can actually execute. Reading a narrow margin as a decisive result is how a comparison stops being evidence and becomes a justification for a preference that was already held.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.