A threshold applied to a test
Researchers often choose a significance level before analysing the result. A common convention is 5%. If the resulting p-value falls below that threshold, the test is called statistically significant. The phrase sounds stronger and broader than it really is.
What has happened is limited: the observed data crossed a line under a defined calculation. A p-value just above the line is not fundamentally different from one just below it. Treating 0.049 as discovery and 0.051 as nothing can create false precision.
Significant is not the same as sizeable
With enough observations, a very small difference can become statistically detectable. Imagine a strategy improving a rate by a fraction of a percentage point. The test may call that difference significant while trading costs make it economically useless.
The reverse can also happen. A potentially important difference may fail to cross the threshold because the sample is small and uncertainty is wide. Absence of statistical significance is not proof that no effect exists.
Good design still comes first
A biased or leaky study can produce tiny p-values. Significance does not repair look-ahead bias, selective reporting, dependent observations or a badly chosen baseline. Nor does it prove that one factor caused another.
A threshold is especially fragile after a wide search. If many industries, filters and time windows were inspected, the final label needs that history. A result selected because it crossed the line has already had an easier test than a rule frozen in advance.
