About the Covariance Calculator
Covariance measures how two variables move together. When one tends to be above its average at the same time as the other, the covariance is positive; when one tends to be above average while the other is below, it is negative; and when there is no consistent pattern, it is close to zero. Covariance is the building block of correlation, regression and portfolio theory, even though it is rarely reported on its own.
This covariance calculator finds the sample covariance (dividing by n − 1) or the population covariance (dividing by n) for two lists of paired data. It explains the direction of the relationship, shows the other version of the covariance for comparison, and gives the correlation coefficient, the unit-free version that is easier to interpret.
How to Use the Covariance Calculator
Enter the X values and Y values in the same order, so that each X is paired with its Y.
Choose whether the data is a sample or the whole population.
The covariance appears with the working and the matching correlation.
The Formula
sample covariance cov = Σ (x − x̄)(y − ȳ) ÷ (n − 1)
population covariance cov = Σ (x − x̄)(y − ȳ) ÷ n
correlation r = cov ÷ (sₓ × s_y)
Each pair contributes the product of its two deviations from the means. The product is positive when both are above or both below their means, and negative when one is above and the other below.
Step-by-Step Example
X = 2, 3, 5, 7, 9 and Y = 65, 70, 75, 85, 90.
x̄ = 5.2, ȳ = 77
x y x − x̄ y − ȳ product
2 65 −3.2 −12 38.4
3 70 −2.2 −7 15.4
5 75 −0.2 −2 0.4
7 85 1.8 8 14.4
9 90 3.8 13 49.4
─────
sum = 118
sample covariance 118 ÷ 4 = 29.5
population covariance 118 ÷ 5 = 23.6
Every product is positive, so the covariance is positive: higher X goes with higher Y. The correlation is 29.5 ÷ (2.864 × 10.368) = 0.994.
Why Covariance Is Hard to Read
The size of a covariance depends on the units of both variables. If the scores in the example were out of 1,000 instead of 100, every Y deviation would be ten times larger and so would the covariance — 295 instead of 29.5 — even though the relationship is exactly the same. There is no natural scale on which to call a covariance large or small. That is why the correlation, which divides the covariance by both standard deviations, is usually reported instead: it always lies between −1 and 1.
Sample or Population?
The sample formula divides by n − 1 rather than n for the same reason as the sample variance: the deviations are measured from the sample means, which fit the data a little too well, and dividing by n − 1 corrects the resulting underestimate. Use the sample formula for data drawn from a larger group — the usual case — and the population formula only when you have the entire group. With large samples the difference is negligible.
A Negative Example
Suppose five cars have engine sizes of 1.0, 1.4, 1.6, 2.0 and 3.0 litres and fuel economies of 60, 52, 48, 42 and 30 miles per gallon. Larger engines go with lower economy, so the products of deviations are mostly negative and the covariance is negative. The calculator confirms the direction and gives a correlation close to −1, showing a strong inverse relationship: the sample covariance is −8.5 and r = −0.991. Swapping the units to kilometres per litre would change the covariance but leave the correlation exactly the same, which is another reminder of why the correlation is the more useful summary.
Covariance in Finance
Investors use covariance to understand diversification. The risk of a portfolio of two assets depends not only on each asset's variance but on their covariance. If two shares tend to rise and fall together, combining them does little to reduce risk; if their returns have low or negative covariance, losses in one are partly offset by gains in the other. The variance of a two-asset portfolio with weights w₁ and w₂ is w₁²σ₁² + w₂²σ₂² + 2w₁w₂ cov₁₂, so the covariance term can raise or lower the total risk.
The Covariance Matrix
With more than two variables, the covariances between every pair are arranged in a covariance matrix, with the variances along the diagonal, since the covariance of a variable with itself is its variance. This matrix is central to multivariate statistics, principal component analysis, portfolio optimisation and machine learning, where it describes how a whole set of features varies together.
Understanding Your Result
The headline is the covariance, sample or population as chosen.
The direction line explains what the sign means.
The other formula line gives the covariance with the other divisor.
The correlation line gives the unit-free correlation coefficient.
The worth knowing line explains why correlation is easier to interpret.
When Should You Use This Calculator?
Use it for statistics homework on covariance and correlation.
Use it to calculate inputs for portfolio risk calculations.
Use it to check the direction of a relationship between two variables.
Use it as a step towards regression slopes, since the slope equals the covariance divided by the variance of X.
Common Mistakes
Comparing covariances across different units. Use correlation instead.
Mismatching the pairs. Each X must line up with its own Y.
Using the population formula for a sample. Divide by n − 1 for samples.
Reading zero covariance as independence. It rules out only a linear relationship.
Forgetting that outliers dominate. A single extreme pair can swing the sign.