Skip to content
← Back to Promotion Workflow

Operator brief · 193

Resampled evidence and forward evidence are not the same evidence.

The key idea

The gap the sandbox leaves

Resampling reuses a fixed pool of outcomes, however many times it shuffles them.

Every simulated path is built from recorded evidence — actual outcomes, actual hit rates, actual branch mix, net of actual friction. Reordering that pool tens of thousands of times produces an enormous number of sequences and exactly one underlying sample, and a candidate tuned against it can fit the sample's particular features without anyone intending to overfit. Seed variation and adjacent-parameter testing narrow this, and they narrow it within the same evidence. The one thing they cannot supply is an outcome the candidate has never been exposed to, and forward observation is definitionally made of those.

FigureA candidate tracked in shadow against its own predicted envelope
the Standard baseline the candidate must beatpredicted upperobservedpredicted lowerobservation weekscandidate performance vs baseline

Schematic: the contract's predicted band, drawn before observation began, with forward-observed behaviour inside it. The prediction's timestamp is what makes the tracking meaningful.

What shadow tracking is

The candidate is scored on live evidence without being given authority over it.

Shadow tracking runs the candidate's logic against evidence as it arrives, recording what the candidate would have done and what that would have produced, while production continues to operate under the existing Standard. The candidate is being measured on genuinely new outcomes and is deciding nothing. That separation is the whole design: it delivers the forward evidence that resampling cannot, without exposing capital to a change that has not finished proving itself. It is the sandbox boundary maintained through the one phase where the temptation to relax it is strongest, because the candidate at this point has cleared every earlier gate.

The comparison that matters

Against the contract's prediction, not against whether it made money.

The contract stated expected behaviour before the run, and the shadow period's job is to check that expectation against forward outcomes. This is a narrower and more useful test than asking whether the candidate performed well. A candidate that outperformed the baseline for reasons its contract never predicted has not been validated — it has produced a favourable result through an unexplained mechanism, which is the profile of something that will reverse without warning. The verdict language is therefore confirmation or violation of the stated behaviour, and a confirmed prediction on a modest margin is stronger evidence than an unexplained large one.

Why it is the last gate

It is the only gate the candidate's author cannot influence.

Each earlier gate has some surface an enthusiastic operator can work: the sandbox run can be re-run, the parameters can be nudged, the persistence window can be chosen, the verdict can be written persuasively. Forward observation offers none of that, because the evidence arrives after the prediction and is generated by the market rather than by the workbook. That immunity is exactly why it sits last and why it cannot be substituted with more sandbox work, however extensive. More resampling produces more confidence about the same sample; a shadow period produces less confidence about a larger claim, which is usually the honest direction.

The cost, acknowledged

This is the slowest gate, and its slowness is not recoverable.

A shadow period cannot be accelerated. It takes as long as it takes for enough forward evidence to accumulate, which on a sixteen-trade weekly rhythm means real weeks, and during all of them production runs the older rule while the better candidate waits. That is a genuine cost paid in foregone improvement, and the system pays it deliberately — the alternative is promoting on backward evidence alone and discovering the difference with capital. The lag is not a defect of the workflow. It is the workflow's product, and the thing being purchased is the difference between a change that looked right and a change that has been right.

  • Shadow evidence is generated after the prediction, which is what makes it immune to tuning.
  • The verdict is confirmation or violation of the contract's stated behaviour, not profitability.
  • No amount of additional resampling substitutes for a shadow period.

The key idea

The last question is whether the model predicted something it had not already seen.

Every gate before this one asks whether the candidate performs. This one asks whether the contract that described it was actually right about a future it did not have access to — and that is the only property that transfers to production, because production is entirely made of futures the candidate has not seen. A promotion that clears a shadow period rests on a prediction that was made in advance and held. Everything earlier rests on a prediction that was made about material already in hand.

Connected inside MARS

Every brief documents the same shipped system.

The complete MARS package — eleven workbooks, three TradingView indicators, the full manual library — $497.