Skip to content
← Back to The Ruler Doctrine

Operator brief · 173

A benchmark measuring a system you no longer run.

The key idea

The failure being prevented

A stale ruler converts drift into a story about performance.

Suppose the live branch mix has slid away from the benchmark's modelled weights over two quarters. Nothing dramatic happened; trade selection simply drifted. The account is now a different system from the one the bands describe, and every comparison against those bands attributes the resulting difference to performance. Below-median reads look like decay when they are composition. Above-median reads look like skill when they are a heavier weighting toward a branch the model under-represented. The comparison remains arithmetically correct and has stopped being informative, and the only defence is a schedule that re-derives the ruler before the gap between model and system grows large enough to swallow the signal.

The monthly tier

The bands that carry live comparisons refresh most often.

Four statistic families refresh monthly, and they are exactly the ones a monthly review touches: equity bands, return bands, drawdown bands, and gate dwell together with worst gate reached. Tier usage distribution, smart exposure compression, and override stress metrics sit on the same monthly rhythm. The logic is proportional — anything read every month must be current every month, or the review is comparing this month's live behaviour against a description of the system as it stood some indefinite time ago. These are also the cheapest statistics to regenerate, which is why the schedule can afford to keep them fresh without turning benchmark maintenance into a second job.

FigureThe regeneration loop — the ruler is an output, not a fixture
Compare live to bandsagainst the current rulerTest the assumptionshas mix or cost drifted?Regenerate on cadencemonthly and quarterly tiersVersion the workbookand update the standardLog the handoverwhich ruler graded whichRE-MARK

The loop closes at the top: a re-derived benchmark becomes the reference the next period is graded against, and the log records the handover so a later reviewer knows which ruler produced which verdict.

The quarterly tier

Survival statistics move slowly and are checked slowly.

Completion probability, lock rate, and hit-week percentiles refresh quarterly or after a major model change. The slower cadence is not neglect — it reflects what these figures are. Completion and lock rates are properties of a full-year horizon, and re-deriving them from a marginally different configuration every month would produce a series of small movements that invite interpretation and carry none. Quarterly regeneration answers a genuinely quarterly question: is the system still aligned with the benchmark model at all, or has enough accumulated that the ruler needs rebuilding rather than adjusting? Throttle efficiency straddles both tiers, checked monthly or quarterly depending on how actively deployment behaviour is moving.

The event trigger

Approved model changes refresh the benchmark regardless of the calendar.

The third trigger is not temporal. Any approved change to the model — revised branch weights, a different stop policy, a meaningfully shifted evidence window — obsoletes the bands immediately, and the refresh follows the change rather than the schedule. The direction of the rule is worth noticing: the benchmark updates after approved model changes, never in response to a live week that felt unrepresentative. That asymmetry is the entire protection. A ruler that could be re-derived because the current one was unflattering would measure nothing, and the temptation to do so arrives precisely when the measurement is most worth having.

The versioning obligation

Re-marking creates a new instrument, and the record has to say so.

The standard requires that assumption changes be versioned in the workbook and reflected in the operating standard itself — the document and the numbers move together. This is heavier than it first appears, because it means a benchmark refresh is a release rather than a recalculation. The reason is downstream: twelve months of review conclusions are only comparable to each other if a reader can tell which of them were graded against which ruler. Without that record, a run of within-benchmark verdicts spanning a refresh looks like consistency and may be two different systems each passing its own exam. The regeneration log is what keeps the historical series interpretable.

The key idea

The instrument is maintained on a schedule, or the measurements decay without notice.

Everything the ruler doctrine claims — that the median is the expected centerline, that the bands define normal, that persistent excursions require diagnosis — is conditional on the bands still describing the system in the account. That condition is not self-maintaining. It is held by a three-speed refresh schedule, an event trigger that fires on approved change and never on emotion, and a version record that lets a reviewer reconstruct which ruler was in force. A benchmark nobody re-marks does not become useless loudly. It becomes useless silently, while continuing to produce percentiles.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.