Skip to content
← Back to Live-Trade Benchmark Comparison

Operator brief · 31

Table K: live results against the simulated envelope, by the book.

The key idea

The precondition

No comparison until the evidence stack is current.

The workflow's first step happens outside the benchmark workbook entirely: Compliance Panel 3 must be fully updated for the period — Journal, Weekly Summary, Gate/Brake state, EV Scorecard, Capital Dynamics, Dashboard. The comparison consumes CP3's outputs as its live inputs, and a comparison built on a stale or half-reconciled journal isn't conservative, it's fiction with a percentile attached. Only once the production record is closed do the three live numbers get collected: month-end equity, monthly return percent, and max drawdown percent, drawn from Capital Dynamics and the Gate/Brake state record.

The entry and the read

Three numbers meet seven bands.

The live values enter the comparison template's input column and land against the simulated P5, P10, P25, Median, P75, P90, and P95 bands for the equivalent point in the year. The read is a classification, not a score: above P75, near median, below median, P10–P25, P5–P10, or below P5. The Operating Standard's framing is worth quoting in spirit — the goal is not to always beat median; the goal is to know whether live behavior is within expected variance and whether risk is being paid for. A P40 month is an answer, not a problem. Persistent P10 months are a different answer. Below P5 is a formal audit trigger regardless of how explainable it feels.

FigureThe band ladder — where a live month can land
Above P90100confirm it isn't excess riskP75–P9086strong — document what's workingP50–P7570healthy, no forced aggressionP25–P5052normal — monitorP10–P2534review branch quality, EV, feesP5–P1020formal diagnostic reviewBelow P58full audit trigger

Schematic of the comparison bands. Position is a location fact; the response escalates only with adverse persistence, never with a single month's placement.

The cross-check battery

Equity placement alone proves nothing — five checks make it a verdict.

A strong equity read with abnormal drawdown is not clean alpha, so the workflow forces the placement through a battery before any conclusion is recorded. Drawdown against its own distribution: is the risk cost of the result normal? Gate dwell: is the live account spending benchmark-normal time in defensive states, or camping in Buffer and Floor? Tier usage and deployment: were higher tiers earned and gate-authorized, or reached by enthusiasm? Open exposure: did carryover risk compress fresh deployment the way the model expects? Overrides: were any Force Full Tier events logged deliberately, or accumulating as habit? Each check can overturn the equity read's first impression — and that's their job.

  • Above-median equity + worse-than-band drawdown = fast-but-risky, not outperformance. Audit sizing and overrides before celebrating.
  • Below-median equity + clean risk behavior = normal variance until persistence says otherwise.
  • Excess defensive dwell with normal equity is an early structural warning most operators would otherwise miss entirely.

The record

The workflow ends in a written classification.

The final step is documentary: record the month's conclusion in the monthly review as one of the standard classifications — within benchmark, outperforming, underperforming, drawdown-adverse, throttle-inefficient, or model mismatch. The vocabulary is fixed on purpose. Free-text conclusions drift with mood; a closed classification set makes twelve months of reviews comparable to each other, and makes persistence — the thing every escalation decision actually depends on — visible at a glance. Model mismatch deserves its own mention: if live assumptions have drifted from the benchmark's register, the honest conclusion is that the ruler needs refreshing, not that performance changed.

The key idea

Narrow inputs, wide cross-checks, closed vocabulary — that's the whole method.

Everything about the comparison workflow fights the same enemy: interpretation drift. Three live inputs mean there's no cherry-picking which numbers to compare. Seven fixed bands mean placement isn't negotiable. Five cross-checks mean a flattering placement can't skip inspection. A closed verdict vocabulary means this month's conclusion can be laid against last month's. The cockpit is small because rulers should be — the judgment lives in what you do with a persistent read, and that belongs to the deviation-classification doctrine one brief over.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.