The premise
A comparison is a controlled experiment or it is nothing.
The purpose of comparing variants is to isolate one thing: the effect of the management structure itself. That isolation only exists if every other input is identical. The moment the two runs differ in some second way — a different random draw, a different stretch of history, a different number of trades — the measured gap contains both the variant's effect and that second difference, mixed together and impossible to separate afterwards. The output still looks like a comparison. It just no longer answers the question it appears to answer.
| Input | Held constant | If it drifts |
|---|---|---|
| Seed | Same draw both runs | Gap is partly luck |
| Trade sample | Identical evidence | Different questions |
| Period | Same window | Regime, not variant |
| Path count | Equal run depth | Unequal precision |

