Keep some evidence out of the workshop
A researcher uses an in-sample period to explore, build and refine an idea. Before opening a separate holdout period, the rule is frozen. The holdout is then used once to see how the finished rule behaves. Because those observations did not shape the method, they provide a cleaner challenge.
The boundary applies to human choices as well as computer training. If somebody repeatedly checks the holdout and adjusts the rule, the holdout has joined the workshop. It is no longer a fresh test.
Later is not automatically unseen
A chronological split is often sensible for markets because it respects time. But a later date is not enough. If the researcher had already inspected the entire archive before choosing the rule or split, the later observations may already have influenced the idea.
That later slice can still reveal instability and provide useful descriptive evidence. It should be labelled as an exposed replay rather than promoted to out-of-sample validation.
One clean test is not the finish line
A rule can pass unseen data by chance, especially when many rules were developed. The test period may also represent only one market environment. Robustness checks, independent replication and prospective records add different kinds of pressure.
Out-of-sample performance should be judged with the same outcome, costs and exclusions frozen in advance. Changing the target after seeing the test creates a new version that needs its own validation.
