Skip to content
← Back to EV & Edge Validation

Operator brief · 435

Validation puts the model on trial, not the week.

The key idea

Two activities, one set of numbers

Computing expectancy is not the same as validating it.

CP3's EV Scorecard computes expectancy across its layers; that is measurement, and it happens whether or not anyone is validating anything. Validation is a separate act performed on top of the measurement: it takes the expectancy model that justified deployment and asks whether live evidence still supports it. The distinction matters because measurement always yields a number and validation is allowed to yield a verdict of no. A desk that only measures will never find out that its thesis has stopped being true, because measurement has no concept of a thesis.

What the thesis has to contain

You cannot validate a claim that was never stated.

Deployment was justified by something specific: branch probabilities in a range, a payoff structure, an expected blend, a net-of-fees figure that survives friction. If those commitments were never written down, validation has nothing to test against and degrades into looking at the current number and deciding whether it feels acceptable. That is not validation; it is reaction with arithmetic attached. The thesis has to exist as a record made before the evidence arrived, which is the only condition under which the evidence can contradict it.

FigureMeasurement and validation, side by side
AspectMeasurementValidation
Question askedWhat did it produce?Does the claim hold?
Reference pointThe current dataThe prior thesis
Possible outputsAlways a numberHold, or falsified
Fails silently whenInputs are staleNo thesis exists

The layered defendant

The thesis is several claims, and they can fail one at a time.

Deployment rests on a stack of commitments rather than a single number, and validation is more useful when it is run against each. Branch-level expectancy claims can fail while blended expectancy holds, because one strong branch can carry a blend that contains something quietly broken. A gross claim can hold while the net claim fails, because friction grew. Volatility around the estimate can widen while the estimate itself stays put, which changes how much the figure deserves to be believed without changing the figure. A single verdict on the blend hides all three.

Why the thesis usually erodes rather than breaks

Falsification arrives slowly and looks like noise the whole way.

Theses rarely fail in a manner that announces itself. The far more common path is gradual: the live figure drifts a little below the claimed range, then a little further, and each individual week is comfortably inside what variance can explain. Validation handles this by testing persistence rather than severity — how long the disagreement has held, not how bad the worst week was. A modest deviation sustained across many weeks is stronger evidence against a thesis than a single dramatic one, and it is precisely the pattern that intuition discounts.

The uncomfortable outcome

Falsified is not the same as fixable.

When validation returns a negative, the reflex is to locate the cause and repair it. Sometimes that is right and the diagnostic layer will name something specific. But validation is not obliged to produce a repair, and it should not be pressed into inventing one. A thesis can simply turn out to have been wrong — the branch probabilities were estimated optimistically, or the regime that supported them has ended. In that case the correct output is that the claim is no longer supported, handed to the evaluation layer for a keep-adjust-retire verdict rather than absorbed as a problem to be solved in place.

The cadence

Validate on a schedule, so the schedule is not chosen by the results.

The timing of validation has to be fixed in advance for the same reason its thresholds do. An operator who validates when performance feels wrong will systematically validate after losses and skip the exercise after gains, producing a record that is thoroughly examined in one direction and unexamined in the other. Running it on the weekly rhythm regardless of how the week went keeps the evidence base symmetric — and means a green stretch gets the same scrutiny as a red one, which is where quiet decay is most likely to be sitting unnoticed. The scheduling discipline costs almost nothing and removes an entire class of selection bias from the record, which is a rare ratio in this kind of work.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.