Learn category
Reading the evidence
Plain-English statistics for judging patterns and claims.
What are you interested in?
Search the Learn library
Search ordinary words or technical terms, then narrow the library by category or level. The seven families stay folded until you choose one.
Reading the evidence
Why sample size matters
A perfect-looking rate can still be weak evidence when it rests on very few observations.
Reading the evidence
A result needs a baseline
A result becomes meaningful only when it is compared with what would ordinarily happen under a relevant definition.
Reading the evidence
Uncertainty is part of the result
A result is more useful when you can see the range of values still compatible with the available evidence.
Reading the evidence
Backtesting without time travel
A backtest replays an idea on old data. It is useful only when the replay respects what could have been known at the time.
Reading the evidence
Probability
Probability represents uncertainty on a scale from impossible to certain under a defined model or evidence base. It describes possible outcomes, not a hidden guarantee about one event.
Reading the evidence
Chance and randomness
Randomness is variation not predictably explained by the information or model being used. It can produce streaks, clusters and convincing patterns even when no durable signal exists.
Reading the evidence
What is a p-value?
A p-value asks how unusual the observed result would be if a particular chance-only model were true.
Reading the evidence
What does statistical significance mean?
Statistical significance means a result crossed a pre-chosen threshold under a particular test. It does not mean important or profitable.
Reading the evidence
What is a confidence interval?
A confidence interval shows a range of values compatible with the estimate under a stated statistical method.
Reading the evidence
False positives
A false positive occurs when a test signals an effect or condition that is not present under the relevant truth definition. The rate depends on thresholds, prevalence and the testing process.
Reading the evidence
False negatives
A false negative occurs when a test fails to detect an effect or condition that is present. Small samples, noisy measurement or strict thresholds can make detection unlikely.
Reading the evidence
Multiple testing
Multiple testing occurs when many hypotheses, groups, thresholds or outcomes are examined. Even valid individual tests can produce an impressive-looking winner by chance across the wider search.
Reading the evidence
What is overfitting?
Overfitting happens when a rule learns the accidents in old data so closely that it struggles with new data.
Reading the evidence
Look-ahead bias
Look-ahead bias lets information unavailable at a historical decision time influence that decision or its measured outcome. It gives the backtest a form of time travel.
Reading the evidence
Data snooping
Data snooping is repeated exploration of the same dataset until a favourable pattern is found, often without carrying the full search into the reported uncertainty.
Reading the evidence
In-sample testing
In-sample testing evaluates an idea on data available for developing, selecting or tuning it. It shows fit to the workshop material, not independent validation.
Reading the evidence
What does out of sample mean?
Out-of-sample testing applies a frozen rule to data that did not help create or tune it.
Reading the evidence
Prospective testing
Prospective testing records a rule, inputs and expectation before future outcomes exist, then attaches those outcomes without rewriting the original record.
Reading the evidence
Correlation and causation
Correlation shows that variables move together under a sample; causation means changing one would change the other under specified conditions. Shared causes, selection and chance can create correlation without causation.
Reading the evidence
Survivorship bias
Survivorship bias occurs when analysis includes entities that remain visible while omitting those that failed, closed or left the dataset. The survivors can make history look safer or stronger.
Reading the evidence
Selection bias
Selection bias arises when inclusion in a sample depends on factors related to the outcome, making the observed group unrepresentative of the intended population.
Reading the evidence
Publication bias
Publication bias occurs when results are more likely to become visible because they are positive, striking or statistically significant. The available record then overstates effects.
Reading the evidence
Confirmation bias
Confirmation bias is the tendency to seek, interpret and remember information in ways that support an existing belief while discounting disconfirming evidence.
Reading the evidence
Regression to the mean
Regression to the mean is the tendency for an extreme noisy observation to be followed by one closer to the typical level, even without a causal intervention.
Reading the evidence
Independent observations
Observations are independent when knowing one does not change the relevant probability distribution of another under the model. Market events often share firms, dates or shocks and are not fully independent.
Reading the evidence
Effect size
Effect size describes the magnitude of a difference or relationship in meaningful or standardised units. It complements uncertainty and helps distinguish detectable effects from useful ones.
Reading the evidence
Statistical power
Statistical power is the probability that a specified test rejects its null when a specified alternative effect is true. It depends on effect size, sample, variability and threshold.
Reading the evidence
Robustness
Robustness is the degree to which a conclusion survives reasonable changes in assumptions, definitions, samples and methods. It is accumulated evidence, not one favourable alternative specification.
Reading the evidence
Replication
Replication repeats a study or test using the same method, independent implementation, new data or a new setting. Each form checks a different source of error.
Reading the evidence
What is a hypothesis?
A hypothesis is a precise, testable statement about a relationship or outcome. It turns a broad observation into something that data could support, weaken or leave unresolved.
Reading the evidence
Falsifiability
Falsifiability means a claim permits some possible evidence to count against it. A claim that explains every outcome after the fact cannot be meaningfully tested.
Reading the evidence
Signal and noise
Signal is repeatable information relevant to the question; noise is variation that obscures or imitates it under the chosen model. The separation is uncertain and context-dependent.
Reading the evidence
Win rate and profitability
Win rate is the proportion of positive outcomes; profitability depends on the size of wins and losses, costs, exposure and sequence. A strategy can win often and still lose money.
Reading the evidence
Mean, median and outliers
The mean adds values and divides by their count; the median is the middle ordered value. Outliers and skew can separate them, so each describes a different centre.
Reading the evidence
Reading a distribution
A distribution shows which values occurred or are possible and how frequently or plausibly they appear. Shape reveals spread, skew, clusters and tails hidden by one average.
Reading the evidence
Sampling error
Sampling error is the variation between a sample estimate and the population quantity caused by observing only part of the population. It differs from bias and measurement error.
Reading the evidence
Pre-registering a test
Pre-registration records the hypothesis, data, exclusions and analysis plan before outcomes are examined. It separates planned confirmation from later exploration and makes deviations visible.
Reading the evidence
Data quality and missing information
Data quality is fitness for a specific use across completeness, accuracy, consistency, timing, provenance and meaning. Clean formatting does not establish that observations represent the intended facts.
