The problem with ad-hoc
A blank test surface returns the operator's own agenda.
Given an empty configuration and no prompts, an operator tests the question currently on their mind — usually a variation on something recent, and usually in the direction they suspect is favourable. Nothing about that is dishonest, and it produces a research record with a systematic hole in it: the adverse conditions nobody was thinking about go untested, indefinitely, because nothing prompts them. Hostile scenarios in particular are almost never generated spontaneously. They have to be sitting there already, as items in a list, waiting to be run.
The first three are inputs the operator can reach for. The fourth is an output most research processes discard, and it is the one that prevents the same ground being re-tested every year.

