Stats & Epi Dictionary

A comprehensive glossary of terms for field epidemiologists

69 Β· EN / RU / KK

Alpha (Ξ±)

Statistics

The probability of rejecting the null hypothesis when it is actually true (Type I error). Typically set at 0.05.

Beta (Ξ²)

Statistics

The probability of failing to reject the null hypothesis when it is actually false (Type II error). Statistical power = 1 - Ξ².

Confidence Interval

Statistics

A range of values that is likely to contain the true population parameter with a specified level of confidence (e.g., 95%).

Try it in IRBIS

Confidence Level

Statistics

The probability that a confidence interval contains the true population parameter. Common levels are 90%, 95%, and 99%.

Degrees of Freedom

Statistics

The number of independent values that can vary in a statistical calculation. Often calculated as n-1 for a single sample.

Effect Size

Statistics

A quantitative measure of the magnitude of a phenomenon, independent of sample size. Common measures include Cohen's d and Pearson's r.

Margin of Error

Statistics

The range of values above and below a sample statistic in a confidence interval. Represents the maximum expected difference between the sample and population parameter.

Mean

Statistics

The arithmetic average of a set of values, calculated by summing all values and dividing by the count.

Median

Statistics

The middle value in a sorted dataset. Less affected by outliers than the mean.

Normal Distribution

Statistics

A symmetric, bell-shaped probability distribution defined by its mean and standard deviation. Many statistical tests assume normality.

Null Hypothesis (Hβ‚€)

Statistics

The hypothesis that there is no significant difference or effect. Statistical tests aim to reject or fail to reject this hypothesis.

P-Value

Statistics

The probability of obtaining test results at least as extreme as the observed results, assuming the null hypothesis is true. Values < 0.05 are typically considered statistically significant.

Power (Statistical)

Statistics

The probability of correctly rejecting a false null hypothesis. Calculated as 1 - Ξ². Power of 80% or higher is generally desired.

Try it in IRBIS

Sample Size

Statistics

The number of observations or units included in a study. Larger sample sizes generally provide more precise estimates.

Try it in IRBIS

Standard Deviation

Statistics

A measure of the amount of variation in a set of values. Low values indicate data points are close to the mean.

Standard Error

Statistics

The standard deviation of the sampling distribution of a statistic. Decreases as sample size increases.

Variance

Statistics

The average of the squared differences from the mean. The square of the standard deviation.

ANOVA (Analysis of Variance)

Statistics

A statistical test used to compare means among three or more groups. Tests whether group means differ significantly from each other.

Try it in IRBIS

Chi-Square Test

Statistics

A statistical test for categorical data to determine whether the observed frequencies differ significantly from expected frequencies.

Try it in IRBIS

Cohen's Kappa (ΞΊ)

Statistics

A statistic measuring inter-rater agreement for categorical items, accounting for agreement by chance. Values range from -1 to 1.

Correlation Coefficient (r)

Statistics

A measure of the linear relationship between two variables. Ranges from -1 (perfect negative) to +1 (perfect positive).

Try it in IRBIS

D'Agostino's KΒ² Test

Statistics

A normality test that combines skewness and kurtosis to determine if data follows a normal distribution. Good for samples n β‰₯ 20.

F-Statistic

Statistics

The ratio of between-group variance to within-group variance in ANOVA. Larger values indicate greater differences between groups.

Jarque-Bera Test

Statistics

A normality test based on skewness and kurtosis. Best for large samples (n > 30). Common in econometrics.

Kruskal-Wallis Test

Statistics

A non-parametric alternative to one-way ANOVA for comparing medians of three or more independent groups.

Try it in IRBIS

Linear Regression

Statistics

A statistical method that models the relationship between a dependent variable and one or more independent variables using a linear equation.

Try it in IRBIS

Logistic Regression

Statistics

A regression method for binary outcomes (0/1). Estimates the probability of an event occurring based on predictors.

Try it in IRBIS

Mann-Whitney U Test

Statistics

A non-parametric test comparing two independent groups. Alternative to independent t-test when normality assumptions are violated.

Try it in IRBIS

Shapiro-Wilk Test

Statistics

A normality test that examines the correlation between data and normal scores. Most powerful for small samples (n < 50).

Try it in IRBIS

T-Test

Statistics

A statistical test comparing the means of one or two groups. Includes one-sample, independent, and paired variants.

Try it in IRBIS

Two-Way ANOVA

Statistics

ANOVA with two independent variables (factors). Tests main effects of each factor and their interaction effect.

Wilcoxon Signed-Rank Test

Statistics

A non-parametric test for paired samples. Alternative to paired t-test when normality assumptions are violated.

Try it in IRBIS

AUC (Area Under Curve)

Statistics

The area under the ROC curve. Ranges from 0.5 (no discrimination) to 1.0 (perfect discrimination).

Likelihood Ratio (LR)

Statistics

The ratio of the probability of a test result in people with the disease to the probability in people without. LR+ > 1 indicates positive association.

Negative Predictive Value (NPV)

Statistics

The proportion of negative test results that are true negatives. Probability of not having the disease given a negative test.

Positive Predictive Value (PPV)

Statistics

The proportion of positive test results that are true positives. Probability of having the disease given a positive test.

ROC Curve

Statistics

Receiver Operating Characteristic curve. A plot of sensitivity vs (1-specificity) at various thresholds, used to evaluate diagnostic test performance.

Sensitivity

Statistics

The proportion of truly diseased persons correctly identified by the test (True Positive Rate). Sensitivity = TP / (TP + FN).

Specificity

Statistics

The proportion of truly non-diseased persons correctly identified by the test (True Negative Rate). Specificity = TN / (TN + FP).

Youden's Index

Statistics

A single statistic summarizing diagnostic test performance: Sensitivity + Specificity - 1. Used to find optimal cutoff points.

Kurtosis

Statistics

A measure of the 'tailedness' of a distribution. Positive kurtosis indicates heavy tails; negative indicates light tails.

Outlier

Statistics

A data point significantly different from other observations. May indicate measurement error or genuine extreme values.

Skewness

Statistics

A measure of asymmetry in a distribution. Positive skew means tail extends right; negative skew means tail extends left.

Attributable Fraction (AF)

Epidemiology

The proportion of disease cases among the exposed that can be attributed to the exposure. AF = (RR - 1) / RR.

Attack Rate

Epidemiology

The proportion of a population that develops a disease during a specific time period, often used during outbreaks.

Incidence

Epidemiology

The number of new cases of a disease occurring during a specified period in a population at risk.

Incidence Rate

Epidemiology

The rate at which new cases occur in a population over time. Calculated as new cases divided by person-time at risk.

Number Needed to Treat (NNT)

Epidemiology

The number of patients who need to be treated to prevent one additional adverse outcome. NNT = 1 / Absolute Risk Reduction.

Odds Ratio (OR)

Epidemiology

The ratio of the odds of exposure in cases to the odds of exposure in controls. Used in case-control studies.

Try it in IRBIS

Poisson Rate

Epidemiology

A rate used for rare events occurring in a given time or space interval. Assumes events occur independently.

Prevalence

Epidemiology

The proportion of a population with a disease at a specific point in time (point prevalence) or over a period (period prevalence).

Relative Risk (RR)

Epidemiology

The ratio of risk in the exposed group to risk in the unexposed group. RR > 1 indicates increased risk with exposure.

Risk Difference (RD)

Epidemiology

The absolute difference in risk between exposed and unexposed groups. Also called Attributable Risk or Absolute Risk Reduction.

Bias

Study Design

Systematic error in study design, conduct, or analysis that leads to incorrect estimation of the association between exposure and outcome.

Block Randomization

Study Design

A randomization method ensuring equal allocation to groups within blocks of participants. Maintains balance throughout recruitment.

Case-Control Study

Study Design

An observational study comparing people with a disease (cases) to those without (controls) to identify exposures associated with the disease.

Cohort Study

Study Design

An observational study following a group of people over time to determine how exposures affect outcomes.

Confounding

Study Design

When a third variable is associated with both exposure and outcome, distorting the apparent relationship between them.

Cross-Sectional Study

Study Design

A study examining a population at a single point in time to assess prevalence of outcomes and exposures simultaneously.

External Validity

Study Design

The extent to which study results can be generalized to other populations, settings, or conditions.

Internal Validity

Study Design

The extent to which a study establishes a trustworthy cause-effect relationship, free from systematic errors.

Randomization

Study Design

The process of randomly assigning study participants to different groups to minimize selection bias and confounding.

Selection Bias

Study Design

Bias arising from the way participants are selected or assigned, resulting in systematic differences between comparison groups.

Simple Randomization

Study Design

Basic random assignment where each participant has equal probability of being assigned to any group, like flipping a coin.

Stratified Randomization

Study Design

Randomization within strata defined by baseline characteristics to ensure balanced distribution of important variables.

Histogram

General

A graphical representation of data distribution using bars where height represents frequency of values in each interval.

Box Plot

General

A graphical display showing data distribution through quartiles, median, and outliers. Also called box-and-whisker plot.

Parametric Test

General

A statistical test that makes assumptions about the underlying distribution of data (usually normal distribution).

Non-Parametric Test

General

A statistical test that makes no assumptions about the underlying distribution of data. Used when normality cannot be assumed.