How to Run Pearson Correlation in SPSS and Interpret the Results

SPSS··8 min read

Pearson correlation is often the output table that appears before linear regression in a quantitative study. You can use it to check whether two variables have a linear relationship, whether the relationship is positive or negative, and how strong that relationship is. A correlation coefficient does not, by itself, prove that one variable causes the other.

This guide follows the SPSS file you have open: prepare the data, run the command, read the Correlations table, check Sig., and write the result in Chapter 4. The numbers in the table are identified as illustrative output, not as the results of a specific study. The procedure also shows where the sample size appears, which matters when missing data are handled differently across variable pairs.

When to use Pearson correlation

Use Pearson correlation when your research question is phrased as, “Is variable X related to variable Y?” It is also appropriate when your hypothesis predicts a linear relationship between two quantitative variables. For example, your study might examine the relationship between service quality and satisfaction, perceived usefulness and usage intention, or academic pressure and academic performance.

In a research model, a hypothesis is often written as: “Service quality has a positive relationship with customer satisfaction.” Pearson correlation then helps you see whether the data support this direction at the bivariate correlation level. If the hypothesis states that X increases Y, the main test will usually require regression or a structural model. Pearson describes the strength of the relationship between X and Y, but it does not estimate an effect while controlling for other variables.

You can run Pearson correlation on composite or mean scores for constructs after checking the scale. For example, after running Cronbach's Alpha, calculate mean scores for variables such as PU, SAT and INT, then enter these variables into the correlation analysis. If you enter every item separately, the table becomes long and it becomes easier to answer the wrong research question. The output may then describe relationships among individual indicators rather than relationships among the constructs in your hypotheses.

Pearson is suitable for numeric variables when the relationship you want to examine is linear. If the data contain many outliers, are severely skewed, or show a nonlinear relationship, inspect the scatterplot and consider Spearman correlation. Do not choose Pearson only because SPSS provides a button for it. The choice should follow the measurement level, data pattern and analysis plan stated in your thesis.

Prepare the data before running the analysis

Open Data View and Variable View to check the variables you will use. Each variable should have the Numeric type, each row should represent one respondent, and each column should represent one variable. If you are using Likert data from 1 to 5, make sure the values have been coded consistently. A variable containing text such as “agree” or “disagree” will not be handled as an ordinary numeric variable.

Before calculating a representative score for a construct, deal with reverse-coded items. For example, one item in a scale may be worded negatively while the remaining items are worded positively. You need to reverse-code that item before calculating the mean. If you skip this step, the Pearson coefficient may reflect a coding error rather than the real relationship in the data. Check the value labels and the recoding rule before you combine the items into a score.

Check missing values with Frequencies or Descriptives. Review the number of missing values for each variable, values outside the Likert range, and special codes such as 99 or 999. If 99 means “no response” but is still treated as a score, the correlation result will be distorted. Declare or replace this value according to the data-cleaning rule stated in your research methods. Keep a record of how many cases remain after each cleaning decision.

Next, check outliers and linearity with a graph. Go to Graphs > Chart Builder, select Scatter/Dot, and create a chart with one variable on the X axis and the other on the Y axis. The points should show a reasonably straight trend. Any points that sit far away from the rest should be checked against the original response file. Do not delete them simply because they make the coefficient look better. A valid unusual response is still part of the data unless your stated rule gives a defensible reason to exclude it.

If you are analysing several constructs, record the number of valid observations for each pair of variables. SPSS can exclude rows with missing data pairwise when it runs the correlation. As a result, the sample size for two pairs in the same matrix may differ when the missing-data patterns differ. Report the relevant N with the coefficient instead of assuming that every cell in the matrix uses the full sample.

You can also review the SPSS materials to compare the data-preparation steps before running the analysis. If your study is still checking the structure of the scale, EFA is usually performed before calculating scores and testing the model.

Steps to run Pearson correlation in SPSS

Step 1: Open the Bivariate Correlations dialog box

In SPSS, select Analyze > Correlate > Bivariate. The Bivariate Correlations dialog box will open. The list of variables in your file appears on the left, while the Variables box on the right is where you place the variables to be tested.

Select the variables representing your research constructs and click the arrow to move them to the Variables box. For example, you can enter SERV, SAT and INT in one run to obtain the correlation matrix for every pair. If you are testing only one hypothesis, you can enter the two relevant variables. Use the names of the calculated construct scores, not the raw item names, when your hypothesis concerns latent constructs.

Step 2: Select Pearson and a two-tailed test

Under Correlation Coefficients, keep Pearson selected. This option produces the r coefficient that you need to report. If your hypothesis predicts a relationship without specifying its direction in advance, select Two-tailed under Test of Significance.

Select One-tailed only when your proposal specified a one-tailed test and you have a clear methodological basis for it. Do not choose a one-tailed test simply to make the p-value smaller. The choice must be consistent with the hypothesis written before you viewed the result.

You may clear Flag significant correlations if you want to read the Sig. column yourself. If you keep it selected, SPSS marks statistically significant correlations with asterisks. The asterisks make the table easier to scan, but they do not replace reporting the r coefficient, sample size and p-value. Read the exact Sig. value and state whether the test was one-tailed or two-tailed.

Step 3: Choose how to handle missing values

Click Options to review how SPSS will handle missing data. Pairwise uses cases with complete data for each pair of variables, while Listwise keeps only rows with complete data for every variable in the matrix. With a table containing several variables, the two options can produce different numbers of observations.

If you choose Pairwise, pay attention to the sample size in each cell or in the output note. If you choose Listwise, the sample size is usually more consistent, but it can decrease substantially when the questionnaire contains many unanswered items. Select the option that matches the data-cleaning process described in Chapter 3. Do not change this setting after seeing the result without documenting the reason.

Step 4: Run the command and save the output

Click Continue, then click OK to run the analysis. In the Output Viewer, SPSS usually creates a Correlations table. Save the output file with the .sav file, using a name that records the run date and data version. If you remove an item or change the way scores are calculated, save a new version instead of overwriting the previous output.

A controlled workflow records the variable names, valid number of observations, choice of Pearson or Spearman, one-tailed or two-tailed testing, and the reason for removing any variable. If your supervisor asks why the Chapter 4 result differs from an earlier run, this log lets you trace the change. It also helps you reproduce the table when you revise the thesis after feedback.

Read the output

The main table to read is Correlations. For each pair of variables, SPSS usually displays three rows: Pearson Correlation, Sig. (2-tailed) and N. The first row is the r coefficient, the second row is the p-value, and the third row is the number of observations used for that pair.

Variable pairPearson CorrelationSig. (2-tailed)NIllustrative interpretation
SERV and SAT0.4680.000214Positive relationship, statistically significant
SERV and INT0.3210.000214Positive relationship, moderate in this example
SAT and INT0.5870.000214Positive relationship, stronger than the two pairs above

Illustrative output, simulating the structure of SPSS Correlations table. The value 0.000 in SPSS does not mean that p is exactly 0. When reporting it, you can write p < 0.001 if the table displays Sig. as .000.

The r coefficient ranges from -1 to +1. A positive sign means that the two variables tend to increase together. A negative sign means that as one variable increases, the other tends to decrease. A value closer to 0 indicates a weaker linear relationship, while an absolute value closer to 1 indicates a stronger linear relationship.

In the example above, r = 0.468 between SERV and SAT indicates a positive relationship. Sig. = 0.000 is smaller than 0.05, so the relationship is statistically significant at the selected threshold. You still should not write that SERV “causes” SAT based only on a Pearson table. The result describes the observed association in the sample.

If r is positive but Sig. is greater than 0.05, do not conclude that there is a statistically significant relationship. If Sig. is smaller than 0.05 but r is very small, report the relationship at its actual magnitude rather than calling it a strong effect. Statistical significance and practical significance are different issues.

When the diagonal of the matrix equals 1.000, it represents a variable's correlation with itself. You do not need to interpret the diagonal cells. The matrix also repeats the same variable pair in two symmetrical positions, so you can retain either the lower or upper triangle when presenting a more compact thesis table.

Evaluation criteria

There is no single threshold that fits every study when you describe a coefficient as strong or weak. Separate two questions: is the coefficient statistically different from 0, and is the relationship meaningful for the research? The table below focuses on decision rules commonly used in quantitative reporting.

Point to checkCriterionWhere to read it in SPSSSource
Correlation coefficientr ranges from -1 to +1Read the Pearson Correlation row(Field, 2013)
Statistical significanceSig. is smaller than 0.05Read the Sig. (2-tailed) row(Field, 2013)
Sample size for the pairN must be reported with the resultRead the N row(Field, 2013)
Check of linearityPoints show a trend close to a straight lineInspect the Scatterplot before running the analysis(Field, 2013)
Correlation between independent variablesVery high correlation requires further review before regressionRead the Correlations matrix, then check VIF in regression(Hair et al., 2019)

In the table above, Sig. 0.05 is a commonly used decision rule, but you should state the selected significance level in your methods section. The sign of r answers the question about direction. The absolute size of r answers the question about strength.

If the independent variables have very high correlations, a problem may appear when you run the regression. Pearson is an initial screening step. It does not replace VIF or the diagnostic checks for the regression model. Reading one correlation matrix and declaring the entire model acceptable is not enough.

How to write this in your thesis

In Chapter 4, state the method, number of observations, r coefficient, Sig. and direction of the relationship. Then connect the result to the hypothesis, using language that matches the type of analysis. You could write: “The Pearson correlation analysis showed that SERV had a positive correlation with SAT, with r = 0.468 and p < 0.001, based on 214 observations. This result supports the hypothesis of a positive relationship between the two variables in the study sample.”

This is illustrative output for you to replace with the result in your own file. Do not copy these numbers into your thesis if they are not from your output.

Use the template below by replacing the text in brackets:

The Pearson correlation analysis showed that [variable X] had a [positive/negative] correlation with [variable Y], with r = [r value], Sig. (2-tailed) = [p-value] and N = [sample size]. Because p was [smaller than/greater than] 0.05, the correlation [was/was not] statistically significant at the [selected level] significance level. Therefore, hypothesis [H number] was [supported/not supported] at the correlation analysis stage.

If the output shows Sig. = .000, write p < 0.001 instead of p = 0.000. If p = 0.127, write p = 0.127 and conclude that there is insufficient statistical evidence at 0.05. You should also report r with a consistent number of decimal places throughout the chapter. Keep the wording about correlation separate from wording about causation unless a later model tests a causal claim.

Common mistakes when running Pearson correlation

Entering individual items instead of representative scores in the matrix. When a scale contains many items, the correlation table becomes dense and difficult to connect to the hypotheses. Decide first whether you are testing relationships between constructs or between individual questionnaire items.

Reading Sig. as the strength of the correlation. Sig. indicates statistical evidence at the selected threshold. The strength and direction of the relationship are shown by r. Two tables can have the same p-value and very different r values.

Inferring causality from Pearson correlation. The statement “X affects Y” requires regression or another suitable model. With Pearson correlation, safer wording is “X is correlated with Y” or “X and Y have a linear relationship.”

Ignoring missing data and outliers. A few unusual rows can increase or decrease r substantially. Check the original record, the reason for the unusual data, and the approved treatment rule. Do not delete a row only because the result looks more consistent with the hypothesis after deletion.

Running Pearson for a nonlinear relationship. Two variables can be related along a curve while their Pearson coefficient remains low. A scatterplot helps you identify this situation before choosing the analysis.

Confusing Pearson with EFA or AMOS. Pearson is a bivariate correlation analysis. EFA explores factor structure, while AMOS is suitable for covariance-based SEM. Review the rotated matrix if you are reading EFA output, or the AMOS materials if your model uses that software.

Frequently asked questions

What is Pearson correlation?

Pearson correlation measures the strength and direction of a linear relationship between two numeric variables. SPSS returns the r coefficient, p-value and N in the Correlations table.

A positive coefficient indicates that the two variables tend to increase together, while a negative coefficient indicates an opposite direction. The coefficient does not, by itself, prove a causal relationship.

What Pearson correlation value is acceptable?

Do not use one r value to decide that every result is acceptable. Review the sign of r, its magnitude, the p-value, the sample size and its consistency with the hypothesis.

If Sig. is smaller than 0.05, you can conclude that the relationship is statistically significant at the 5% level, provided that the data and assumptions of the analysis are appropriate. Then report the magnitude of r instead of writing only “acceptable” or “not acceptable.”

What Sig. value is statistically significant in Pearson correlation?

In common reporting, Sig. smaller than 0.05 is considered statistically significant at the 5% level. If Sig. is greater than or equal to 0.05, you do not have enough evidence to conclude that the relationship differs from 0 at this level.

Check that you are reading the correct Sig. (2-tailed) row and the correct variable pair. Do not use a diagonal value or the Sig. value for another pair to make a conclusion about the current hypothesis.

Can Pearson correlation be used for a Likert scale?

In many theses, the mean or composite score of multiple Likert items is treated as a numeric variable for correlation analysis. You need to explain how the score was calculated and check the quality of the scale before running the analysis.

If you have only one ordinal variable, many outliers, or a nonlinear relationship, consider Spearman correlation. The decision should be based on the characteristics of the data and the analysis plan, not only on the habit of using SPSS.

What should I do if Pearson correlation is not significant?

First, recheck the coding, reverse-coded items, missing values, score calculation and outliers. Then inspect the scatterplot to make sure the relationship you are testing is linear. Do not remove a variable or switch tests simply to obtain a smaller p-value.

If the hypothesis is not supported, report that result honestly and explain it within the limits of the sample, scale and model. The result may differ when you move to regression because several variables enter the model at the same time, but you still need to retain the output and the procedure you ran.

Open the .sav file again, check the names and calculation of the representative variables, then rerun Analyze > Correlate > Bivariate with Pearson, Two-tailed and the missing-data treatment recorded in your methods section. If you need to run the analysis on your own .sav or .csv file, M4 SPSS and SmartPLS analysis can help you check the output and turn the results into a reporting table without replacing your real data.