Skip to content
← Back to Risk-Tier Performance

Operator brief · 41

Spotting when a tier stops behaving like its simulation.

The key idea

The signature concept

Every tier has a simulated fingerprint to drift from.

The benchmark doesn't just expect the ladder to be used — it expects each rung to be used in a particular way: how often, in which gates, with what deployment intensity, producing what contribution. Together those expectations form a per-tier signature, drawn from the tier-usage, deployment, and gate-conditioned tables of the 50,000-path run. Drift detection treats the signature as the reference and asks a narrow question on cadence: is this rung's live behavior still recognizably the behavior the simulation modeled? The question deliberately precedes any judgment about good or bad — a tier can drift toward better numbers and still be drifting, because the issue isn't quality. It's whether the live ladder and the simulated ladder are still the same instrument.

The two drift families

Usage drift and conversion drift are different diseases.

Usage drift is frequency moving off-signature: a rung firing far more or less often than the benchmark's distribution, or firing in gates where the simulation rarely placed it. Its causes split cleanly — either the inputs feeding the allocator have shifted (evidence quality, gate dwell, market regime), or the operator has, through overrides and hesitations that bend resolved tiers toward preferred ones. Conversion drift is the signature's output side eroding: the rung fires on schedule but converts deployed risk into return progressively worse than its history and its benchmark expectation. Its usual suspects are execution decay under that tier's conditions, cost structure creeping, or the conditions themselves changing under a stable trigger. The families matter because their investigations diverge immediately — usage drift interrogates the allocator's inputs and the override log first; conversion drift interrogates execution quality and the attribution slices.

FigureTwo drift families against a stable benchmark signature
benchmark expectationusage driftconversion driftreview quartersvs. benchmark signature

Schematic: usage frequency sliding off the benchmark's expected rate, and conversion efficiency eroding against a stable expectation. Either alone triggers investigation; both together suggest the tier's conditions have genuinely changed.

The detection mechanics

Drift is a trend across reviews, never a reading in one.

Because drift is by definition gradual, no single review detects it — the instrument is the accumulated record. The quarterly attribution reads, each individually unremarkable, are laid side by side: usage share per tier across quarters, conversion per tier across quarters, gate-conditioned placement across quarters, each against its benchmark band. The classification borrows the deviation doctrine wholesale — a reading that wanders and reverts is noise; a sequence trending one direction is drift; a step-change that holds is a break in the tier's conditions. The thin-sample rungs get proportionally patient treatment: drift verdicts on rarely-used tiers ripen on annual timelines, and the honest interim output is a directional watch-note, not a finding. What makes any of this possible is the closed-vocabulary review record — drift detection is, mechanically, just reading that record with a ruler.

  • Lay quarters side by side; the trend is the signal, the individual reads are just samples of it.
  • Noise reverts, drift trends, breaks hold — the same taxonomy that governs benchmark gaps governs tier signatures.
  • A tier drifting toward better numbers still gets flagged. Unmodeled improvement is still unmodeled.

What a confirmed drift triggers

Diagnosis first, then the sandbox — and sometimes a new ruler.

A confirmed drift opens with attribution's standard triage: is the gap execution, conditions, or structure? Execution findings route to the discipline layer — adherence, override behavior, the mechanics of trading that rung's conditions. Structural findings — the tier's trigger conditions no longer matching where the edge lives — become scenario candidates, entering the sandbox with explicit expected behavior and surviving the persistence window before any promotion touches the production ladder. And one outcome is unique to drift detection: if the investigation concludes the live system has legitimately, deliberately evolved past the benchmark's assumptions, the finding isn't a tier problem at all — it's a stale-ruler problem, and the Operating Standard's refresh rule applies. Drift detection thus protects both directions of the comparison: the ladder from unnoticed decay, and the benchmark from measuring a system that no longer exists.

The key idea

The simulation is only a mirror while the resemblance holds.

Every comparison in the simulation layer — bands, dwell, deployment, attribution — silently assumes the live ladder still is the simulated ladder. Drift is that assumption failing in slow motion, and drift detection is the maintenance contract: per-tier signatures, checked on cadence, classified by the standard taxonomy, routed to diagnosis, sandbox, or refresh. It's unglamorous work that mostly confirms nothing has changed. The quarters where it finds otherwise are the quarters it pays for years of itself.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.