
How to Calculate Variance: Formula, Meaning, and Output Interpretation
You open your data file and see Variance in the output, but you are not yet sure what that number says about your survey sample. Variance measures how far values are spread around the mean. The larger the variance, the more widely the data are spread; the smaller the variance, the closer the observations are to the mean.
Two tasks are often mixed together: calculating variance from a set of numbers and reading the variance of a variable in SPSS. This article moves from the formula and a numerical example to checking the result in a .sav file, so you can compare it directly with the output currently open on your screen.
What is variance and how do you calculate it?
For a set of numbers, variance is calculated by finding each value's deviation from the mean, squaring those deviations, adding them together, and dividing by an appropriate denominator. Squaring is used so that positive and negative deviations do not cancel each other out.
If you are calculating variance for an entire population, the formula is:
σ² = Σ(xᵢ − μ)² / N
Here, xᵢ is each observation, μ is the population mean, and N is the number of observations in the population.
If the data come from a sample used to make inferences about a population, the usual formula is:
s² = Σ(xᵢ − x̄)² / (n − 1)
Here, x̄ is the sample mean and n is the sample size. Dividing by n − 1 reflects the fact that the sample mean is estimated from the same data. In your thesis, state clearly whether you are describing a population or a survey sample. For student survey data, the descriptive output in SPSS usually reports sample variance.
For example, suppose the scores are 2, 4, 4, 6. The mean is 4. The deviations from the mean are −2, 0, 0, 2, and the squared deviations are 4, 0, 0, 4. The sum of squared deviations is 8. If these values are treated as a population, the variance is 8 / 4 = 2. If they are treated as a sample, the variance is 8 / 3 ≈ 2,67.
You can also see the variance formula if you need to check each symbol and each step of the manual calculation.
What variance means in quantitative research
In quantitative research, variance helps you see whether an item or composite variable has dispersion. Two variables can have the same mean and still have very different levels of dispersion. For example, the mean score of two groups may both be 4, but one group may be concentrated around 4 while the other has many respondents choosing 1 and 7. The second group has the larger variance.
Variance appears in descriptive statistics, regression, ANOVA, and many other analyses. It helps you check whether a variable has enough variation for analysis. If every respondent chooses the same response option, the variance is 0. That variable provides no differences between observations and may create problems when entered into a model.
Variance has squared measurement units. If a variable is measured on a Likert scale, its variance is expressed in squared score units, so direct interpretation can sometimes be difficult. Descriptive reporting therefore usually presents Variance together with Mean and Std. Deviation. Standard deviation is the square root of variance and keeps the original measurement unit. You can read the article on standard deviation to understand why these two statistics usually appear next to each other.
Variance can also help you identify an item that may have been answered in a fixed pattern. If a variable has a high Mean but very low Variance, most respondents may have selected the same response option. This alone is not enough to conclude that the item should be removed. You need to examine the distribution, the item wording, and the purpose of the scale.
What variance value is acceptable?
Variance has no universal cutoff of the type “above what value is acceptable.” The appropriate value depends on the scale, score range, sample size, and purpose of the analysis. A variance of 2 may be large for a variable that only takes values from 1 to 5, but small for income measured in millions of currency units.
Use the following principles when making a decision instead of setting one cutoff for every data set.
| Situation to check | How to read the variance | Basis or source to acknowledge |
|---|---|---|
| Variance equals 0 | All observations have the same value, so the variable has no variation | The rule for calculating variance from the statistical formula |
| Very small variance | Respondents are concentrated around one or a few levels, so check frequencies and the distribution | Descriptive statistical interpretation, with no fixed item-removal cutoff |
| Large variance | The data are widely dispersed, so check outliers and the measurement range | Descriptive statistical interpretation, with no fixed item-removal cutoff |
| Variance used in a regression model | Also check assumptions, residuals, and the influence of outliers, rather than drawing a conclusion from Variance alone | (Field, 2013) |
| Dispersion among items in a scale | Combine it with Cronbach's Alpha, Corrected Item-Total Correlation, and EFA | (Nunnally and Bernstein, 1994) and (Hair et al., 2010) |
A large number does not automatically mean that the data are good. It may reflect real differences between respondents, but it may also result from an incorrect unit, a few unusual values, or unclear item wording. Conversely, a small variance does not automatically mean that the scale fails. Place it in the context of the full variable and model.
In the article on variance, you can also see how population variance, sample variance, and common reporting formats differ in descriptive statistics.
How to read variance in SPSS output
To view the variance of one or more variables, open SPSS and select Analyze > Descriptive Statistics > Descriptives. Move the variables you want to check into Variable(s), click Options, select Mean, Std. deviation, Variance, Minimum, and Maximum, then click Continue > OK.
SPSS creates a Descriptive Statistics table. The Variance column contains variance, while Std. Deviation contains standard deviation. You should not take the square root of an unrelated number and compare it with Variance. The correct relationship is Std. Deviation² = Variance, apart from rounding in the displayed output.
Illustrative output
The table below is illustrative output, not the result of a real study. It follows the format of the Descriptive Statistics table that SPSS commonly produces.
| N | Minimum | Maximum | Mean | Std. Deviation | Variance |
|---|---|---|---|---|---|
| 120 | 1.00 | 5.00 | 3.68 | 0.742 | 0.551 |
For this illustrative row, Variance = 0.551 and Std. Deviation = 0.742. Calculating 0.742² gives approximately 0.551. Mean, at 3.68, shows the average level, while Variance shows the dispersion around that average. You cannot say that the variable passes or fails simply because Variance is 0.551.
If you use Analyze > Descriptive Statistics > Frequencies, you can also view a frequency table and charts. This is worth doing when the variance appears unusually small or large. A frequency table shows whether respondents are concentrated at one level, distributed relatively evenly, or include a few values far from the rest.
For a scale with multiple items, do not look only at the Variance of each item and then conclude that the scale is reliable. Run Analyze > Scale > Reliability Analysis, move items measuring the same construct into Items, click Statistics, select Scale if item deleted, and then review Reliability Statistics and Item-Total Statistics.
What to do when variance appears unacceptable
First, identify what “unacceptable” means in your case. If your supervisor asks you to check whether a variable has variance equal to 0, you handle it differently from a variable with a large variance caused by an outlier. Variance itself has no universal removal cutoff.
Check coding and value ranges
Open Variable View and inspect Values, Measure, Missing, and the variable labels. If the scale runs from 1 to 5 but Data View contains 55 or 0, the data may have been entered incorrectly or missing values may not have been coded properly. Correct the source data and record the change. Do not silently delete large numbers of values.
Check reverse-coded items
If the questionnaire contains reverse-coded items, recode them before calculating the scale score. For a 1 to 5 Likert scale, the new value is commonly created using the relationship 6 − old value. Check the questionnaire codebook before using this formula, because not every scale has the same number of response levels.
Check outliers
Use Analyze > Descriptive Statistics > Explore, move the variable into Dependent List, open Plots, and select the appropriate charts. Review the boxplot, minimum, maximum, and list of unusual observations. You can read more about outliers to distinguish an extreme value caused by an entry error from a valid observation.
If an outlier is an entry error, correct it using the original data. If it is a valid response, do not delete it simply to make the variance look better. Run the analysis with and without that observation, explain your treatment criterion, and report the decision transparently to your supervisor.
Rerun the statistics after making corrections
After each change, save a new version of the data file and rerun Descriptives. If a variable has variance equal to 0 because all responses are identical, the software cannot create variation through a technical adjustment. Review the questionnaire design or sample quality instead of trying to manipulate the number in SPSS.
Distinguishing variance from standard deviation and similar statistics
Variance and standard deviation both describe dispersion. Variance is the mean of the squared deviations, while standard deviation is the square root of variance. Because standard deviation keeps the original measurement unit, it is usually easier to interpret in descriptive statistics.
For example, if a variable has variance of 0.551, its standard deviation is approximately 0.742. If you report both, the numbers must be consistent. Do not label Variance as standard deviation in the results table, because they are two different columns in Descriptive Statistics.
The mean shows the central position, while variance and standard deviation show dispersion. The median is the middle value after the data are ordered, so it may differ from the Mean when the distribution is skewed or contains outliers. The article on the median is more suitable when you need to compare measures of central position.
You should also avoid confusing variance with R² in regression or extracted variance in EFA. R² refers to the proportion of variation in the dependent variable explained by the model. Extracted variance refers to the amount of variation in the observed variables retained by the factors. These terms both contain “variance” in English, but they answer different questions.
Common errors when calculating variance
The first error is dividing by n in every situation. When calculating population variance, dividing by N is appropriate; when estimating from a sample, the formula usually divides by n − 1. If you calculate variance manually to compare it with SPSS, check whether you are comparing sample or population variance.
The second error is calculating the mean incorrectly because of missing data. If blank cells were entered as 0, Mean and Variance will be pulled downward. In SPSS, check the missing-value coding and the number of valid observations in the N column.
The third error is concluding that a variable is good simply because its variance is large. Strong dispersion may occur together with outliers, entry errors, or respondents interpreting the item in different ways. Check the histogram, boxplot, frequencies, and item wording as well.
The fourth error is using variance as a substitute for Cronbach's Alpha. A scale whose items are dispersed does not necessarily have items measuring the same construct. Reliability must be assessed using an appropriate procedure, including Reliability Statistics, Item-Total Statistics, and EFA or CFA steps depending on the research design.
Finally, do not round too early during a manual calculation. Keep more decimal places during intermediate steps and round only the final result. When comparing your calculation with SPSS, a very small difference may simply result from how many decimal places the software displays.
Frequently asked questions
How do you calculate variance with a calculator?
You can enter the numbers into Excel or a statistical calculator. For sample variance in Excel, the usual function is VAR.S; for population variance, it is VAR.P. Identify the type of data before selecting the function, because the two functions use different denominators.
What does a variance of 0 mean?
A variance of 0 means that every observation has the same value. The variable has no variation in the current data set. Check the coding, data entry, and filters before deciding whether to remove the variable.
Is a larger variance always better?
No. A larger variance only means that the values are farther from the mean. It may reflect real differences, but it may also result from an outlier or a data error. Read it together with the Mean, standard deviation, frequencies, and scale context.
How are variance and standard deviation different?
Standard deviation is the square root of variance. Variance uses squared units, which makes it harder to interpret directly, while standard deviation keeps the variable's original unit. The two statistics must match according to the relationship Variance = Std. Deviation².
Which formula does SPSS use for variance?
When you run descriptive statistics for sample data, SPSS usually reports sample variance, meaning the sum of squared deviations divided by n − 1. If your manual result does not match, check the number of valid observations, missing values, filters, and whether you used n or n − 1.
Open Analyze > Descriptive Statistics > Descriptives in your file now, check N, Mean, Std. Deviation, and Variance, and then inspect unusual values before placing the table in Chapter 4. If you want to run statistics on your own .sav or .csv file and learn how to explain the results, see DoThesis M4 data analysis.