hasibai

Math

Standard deviation calculator

The average distance from the mean. The only real decision is whether your numbers are the whole population or a sample drawn from one.

Dispersion
Enter your figures to see the breakdown.

Why sample and population differ

Population:  σ = √( Σ(x − μ)² ÷ N )
Sample:      s = √( Σ(x − x̄)² ÷ (n − 1) )

Only the divisor changes. The reason is that a sample's own mean sits, by construction, in the middle of that sample — closer to those particular values than the true population mean would be. Squared deviations measured from the sample mean are therefore systematically too small.

Dividing by n − 1 instead of n corrects the bias. It is called Bessel's correction, and the intuition is that one degree of freedom was used up estimating the mean: once you know the mean and all but one value, the last value is determined.

The difference matters most for small samples. With n = 5 it inflates the variance by 25%; with n = 100 by about 1%. When in doubt, use the sample formula — if the data came from measurements or a survey rather than a complete census, it is a sample.

Why square the deviations at all

Deviations from the mean always sum to zero, so averaging them directly is useless. Squaring removes the sign, and taking the square root at the end returns the result to the original units — a standard deviation of pounds is in pounds, not square pounds.

Squaring also weights large deviations more heavily than small ones, which is deliberate: a single value far from the mean says more about spread than several nearby ones. The alternative, mean absolute deviation, treats all distances equally and is more robust to outliers but has worse mathematical properties.

The 68-95-99.7 rule

For roughly normal data, about 68% of values fall within one standard deviation of the mean, 95% within two, and 99.7% within three. This is why standard deviation is such a useful summary: it converts a raw spread into a statement about where values are likely to sit.

The rule fails for skewed or heavy-tailed distributions. The comparison rows above show what proportion of your data actually falls in each band, which is a quick check on whether the assumption holds.

Standard deviation versus standard error

These get confused constantly. Standard deviation describes the spread of individual values. Standard error describes the precision of the estimated mean, and it shrinks as the sample grows — dividing by √n.

So a large sample of variable data has a high standard deviation and a low standard error: individual observations are scattered, but you know the average well. Reporting the wrong one makes results look considerably more or less precise than they are.

Common questions

Is a high standard deviation bad?

Not inherently — it depends on the context. High variability in manufacturing tolerances is a problem; high variability in the distribution of species in an ecosystem may be healthy. The coefficient of variation, which expresses the standard deviation as a percentage of the mean, is useful for comparing spread across data with different scales.

Can standard deviation be negative?

No. It is a square root of a sum of squares, so it is always zero or positive. Zero means every value is identical.

How does an outlier affect it?

Substantially, because deviations are squared. A single value far from the mean can dominate the calculation. If one point is driving the result, check whether it is a genuine observation or a data-entry error before drawing conclusions.

What is variance for, if standard deviation is more interpretable?

Variance adds cleanly: the variance of a sum of independent variables is the sum of their variances, which is not true of standard deviations. That property makes variance the natural quantity in the underlying mathematics, while standard deviation is the one you report.