About the Linear Regression Calculator
Linear regression finds the straight line that best fits a set of paired data, so you can describe how one variable changes with another and make predictions. How many extra marks does each hour of study bring? How much does monthly revenue rise for every extra 1,000 spent on advertising? How does fuel use grow with distance? A regression line turns scattered points into a simple equation, ŷ = a + bx.
This linear regression calculator fits the least-squares line to your data and gives the slope and intercept, R² (how much of the variation the line explains), a significance test for the slope, the residual standard error (the typical size of a miss), and an optional prediction at any X value, with a warning if you are extrapolating beyond the data.
How to Use the Linear Regression Calculator
Enter the X values (the explanatory variable) and the Y values (the variable you want to predict), in the same order.
Optionally, enter an X value to predict Y for.
The regression equation and statistics appear below.
The Formulas
slope b = Sxy ÷ Sxx = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)²
intercept a = ȳ − b × x̄
line ŷ = a + b x
R² = 1 − SSres ÷ SStot
SE(b) = √( SSres ÷ (n − 2) ÷ Sxx ), t = b ÷ SE(b)
Step-by-Step Example
Hours studied (2, 3, 5, 7, 9) and scores (65, 70, 75, 85, 90).
x̄ = 5.2, ȳ = 77
Sxy = 118, Sxx = 32.8
b = 118 ÷ 32.8 = 3.598
a = 77 − 3.598 × 5.2 = 58.29
ŷ = 58.29 + 3.598x
Each extra hour of study is associated with about 3.6 more marks. The intercept, 58.29, is the predicted score for zero hours — a value outside the data, so it should be read with caution.
Prediction at x = 6: 58.29 + 3.598 × 6 = 79.88
R² = 0.987
SE(b) = 0.236, t = 15.23, p = 0.0006
Residual standard error = 1.35
The line explains 98.7 percent of the variation in scores, and the typical point lies about 1.35 marks from the line.
What Least Squares Means
For any line through the data, each point has a residual: the vertical distance between the actual Y and the value the line predicts. Least squares chooses the line that makes the sum of the squared residuals as small as possible. Squaring makes positive and negative misses count equally and penalises big misses more heavily. The resulting line always passes through the point (x̄, ȳ), the centre of the data.
Interpreting the Slope and Intercept
The slope is the average change in Y for a one-unit increase in X, in Y's units per X unit — marks per hour, dollars per unit sold, litres per kilometre. The intercept is the predicted Y when X is zero. It is meaningful only if zero is a sensible, observed value of X; often it is merely where the line happens to cross the axis. A slope test with a small p-value suggests the relationship is not simply chance.
Prediction and Extrapolation
Predictions within the range of the observed X values — interpolation — are usually reasonable. Predictions outside it — extrapolation — assume the straight line continues, which it often does not. The study example predicts 94.3 marks for 10 hours, which seems fine, but 130 marks for 20 hours, which is impossible on a test out of 100. The calculator flags any prediction outside the data range.
Checking the Model
A straight line is an assumption, not a fact. Plot the data and look at the residuals: they should scatter randomly around zero with no pattern. A curved pattern in the residuals suggests the relationship is not linear; residuals that fan out as X grows suggest the variation is not constant; one or two huge residuals indicate outliers that may be pulling the line. A high R² alone does not guarantee a good model.
Regression in Business
Businesses use simple regression constantly. A shop can regress monthly sales on advertising spend to estimate the return on each extra unit of spending; a factory can regress energy use on output to split its bill into fixed and variable parts, where the intercept estimates the fixed cost and the slope the cost per unit produced. In each case the line is a summary of past data, so it works best when conditions stay similar.
Regression Versus Correlation
Correlation measures how strongly two variables are linearly related, treating them symmetrically. Regression goes further, giving an equation to predict Y from X, and it is not symmetric: regressing X on Y gives a different line. The two are linked: the slope equals r × (SD of Y ÷ SD of X), and for simple regression R² is the square of r.
Understanding Your Result
The headline is the regression equation.
The slope line explains the slope in words.
The fit line gives R².
The slope test line gives the standard error, t statistic and p-value.
The residual spread line gives the typical distance of points from the line.
The prediction line gives ŷ at your chosen X, with an extrapolation warning.
When Should You Use This Calculator?
Use it to find a trend line for sales, costs, measurements or scores.
Use it to estimate how much Y changes per unit of X.
Use it to make predictions within the range of your data.
Use it for statistics coursework on least-squares regression.
Common Mistakes
Extrapolating far beyond the data. The line may not hold there.
Reading the slope as proof of cause. Regression shows association.
Ignoring curves in the residuals. Try a transformed or curved model.
Swapping X and Y. Put the variable you want to predict in Y.
Trusting the intercept when X = 0 is meaningless. It may be a mathematical artefact.