Skip to content
← Back to Dynamic 7-Tier Benchmark

Operator brief · 29

The only two questions the benchmark answers.

The key idea

Question one

Is live MARS performance inside the benchmark envelope?

This is a location question, and it deliberately replaces the questions traders actually ask — 'am I doing well?', 'am I behind?' — with one that has a measurable answer. The envelope is the P5-to-P95 span of governed simulated outcomes for equity, return, and drawdown at your current point in the year. Inside it, the system is behaving as modeled, whether that feels good or not: a month at P35 is below median and completely normal, because 35% of governed futures with your exact edge look like that or worse. The question's power is what it refuses to do — it refuses to treat below-median as failure or above-median as skill. Location first; judgment only if location warrants it.

Question two

If the difference is real — what species of difference is it?

Only when live results sit persistently outside normal bands does the second question activate, and it is a classification, not a celebration or a panic. The Operating Standard names the species: favorable alpha — genuine outperformance with normal risk behavior. Normal variance — a deviation that time will reabsorb. Execution drag — the model is fine but fees, slippage, or adherence are leaking. Excessive drawdown — returns bought with risk the simulation says you didn't need to spend. Governance failure — the numbers differ because the rules weren't followed. Each species has a different correct response, which is exactly why the classification must precede any action.

  • Above the envelope with worse-than-band drawdown is usually over-risk wearing alpha's clothes.
  • Above the envelope with heavy override counts is governance leakage, not edge.
  • Below the envelope with clean risk behavior most often traces to execution drag or branch-mix drift — checkable, fixable, boring.

Why only two

The discipline of a small question set.

Benchmark workbooks fail in practice not from missing data but from unbounded interrogation — an operator with dozens of tables and no fixed questions will find whatever mood they brought. Restricting the benchmark to two questions makes every table's role legible: the equity, return, and drawdown distributions serve Question One by locating you; gate dwell, tier usage, exposure, override stress, and throttle efficiency serve Question Two by classifying the gap. A table that can't be traced to one of the two questions has no vote in the review. That discipline is what separates a benchmark review from a scroll through charts.

FigureEvery benchmark table serves one of two questions
Q1 · Locationinside the envelope?· Equity distribution· Return distribution· Drawdown distribution· Live-compare cockpit· Completion & timingQ2 · Classificationwhat kind of gap?· Gate dwell & regime· Tier usage & deployment· Smart exposure· Override stress· Throttle efficiency

The workbook's table alphabet, sorted by the question it answers. Location tables place you; classification tables explain persistent gaps.

The cadence

The questions repeat on a fixed rhythm, at fixed depth.

Weekly, the questions get a light touch: is the current week inside normal variance, or creating early stress? Monthly, the full treatment: month-end equity, return, and max drawdown against the bands, then the classification battery if anything sits outside. Quarterly, the meta-version: is the system still aligned with the benchmark model at all, or has enough drifted — branch mix, rhythm, cost structure — that the ruler itself needs refreshing? The escalating depth matters because the two questions have different evidentiary appetites: location can be read off one month, but classification demands persistence before it means anything.

The key idea

A benchmark review that ends without answering both questions didn't happen.

The complete monthly output is two sentences. First: 'Live performance is at [position] relative to the governed envelope.' Second, only if warranted: 'The persistent difference classifies as [species], and the response is [action].' If a review session produces pages of observations but neither sentence, the tables were visited and the benchmark was not used. Two questions, asked the same way every month, compound into the thing the whole simulation layer exists to provide: an evidence trail showing whether MARS is actually converting expectancy into alpha — or just into motion.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.