
What Is Quantitative Data? How to Identify and Analyze It
What Is Quantitative Data
Quantitative data is data represented by numbers that can be measured, coded, summarized, and analyzed using statistical methods. When you ask 250 people to rate their satisfaction on a Likert scale from 1 to 5, the score selected by each person is quantitative data. Age, income, number of purchases, and grade point average are also quantitative data.
The key point is whether you can place the data into a clear measurement structure. Each row in an SPSS or CSV file usually represents one respondent, while each column represents a variable, such as gender, age, SAT1, SAT2, or purchase intention. From these columns, you can calculate frequencies, percentages, means, standard deviations, correlations, regression results, Cronbach's Alpha, EFA, or a PLS-SEM model.
For example, a study of the factors affecting online cosmetics purchase intention might produce the following observed variables:
PU1toPU4: agreement with statements about perceived usefulness.TR1toTR3: agreement with statements about trust.INT1toINT3: agreement with statements about purchase intention.AGE: the respondent's age.GENDER: gender coded numerically.
Quantitative data is not limited to directly measured figures such as money or height. A questionnaire response also becomes quantitative data when you define a consistent coding scheme, such as 1 = strongly disagree and 5 = strongly agree. The Likert scale has been used to measure attitudes in this way since the foundation proposed by Likert (Likert, 1932).
If you are asking broader questions about the research process, see what is scientific research and the Scientific Research topic to place the data in the right thesis context. The article What is scientific research also helps you distinguish between a topic, research question, method, and data.
Why Quantitative Data Matters in Quantitative Research
Quantitative data lets you turn a research question into variables that can be tested. For example, the question “does service quality affect satisfaction” must be translated into observed variables for service quality and satisfaction. After collecting responses, you can test the relationship using correlation or regression, depending on your model and hypotheses.
Your data also determines which analysis method you should choose. If you have a score-based outcome variable and several explanatory variables, linear regression may be appropriate. If your model contains latent variables and multiple observed variables, you may consider using SPSS together with SmartPLS. With PLS-SEM, the usual process assesses the measurement model first and then the structural model, following the two-step approach presented by (Anderson and Gerbing, 1988).
In a real data file, quantitative data helps you answer at least four groups of questions:
- What are the characteristics of the survey sample, such as gender, age, and occupation?
- Are the observed variables reliable enough to combine into scales?
- Do the scales demonstrate convergent validity and discriminant validity?
- Does the data support the hypotheses about relationships between variables?
Quantitative data therefore does not mean simply running a table of means. You need to move from questionnaire design and coding through data cleaning, scale assessment, and model testing. A number only has meaning when you know which concept it represents, how it was measured, and how many observations produced it.
How Much Quantitative Data Is Enough
There is no single number that applies to every dataset. “Enough” may refer to sample size, scale quality, whether the data is suitable for EFA, or the indicators used in PLS-SEM. Compare each criterion with the purpose of your analysis instead of reducing everything to one general threshold.
The table below gives reference points commonly used in quantitative theses. Each row includes its source, so you know what to cite when your supervisor asks for the basis of a decision.
| Check | Reference point | Source |
|---|---|---|
| Observations per observed variable | 5 to 10 observations per variable | (Hair et al., 2010) |
| Cronbach's Alpha | 0.7 or above | (Nunnally, 1978) |
| Alpha for a new or exploratory scale | May be 0.6 or above | (Hair et al., 2010) |
| Corrected Item-Total Correlation | 0.3 or above | (Nunnally and Bernstein, 1994) |
| KMO when running EFA | 0.5 or above | (Kaiser, 1974) |
| Bartlett's Test | Sig. below 0.05 | (Kaiser, 1974) |
| Factor loading | 0.5 or above | (Hair et al., 2010) |
| Total Variance Explained | 50% or above | (Hair et al., 2010) |
| Composite Reliability | 0.7 or above | (Fornell and Larcker, 1981) |
| AVE | 0.5 or above | (Fornell and Larcker, 1981) |
| HTMT | Below 0.85, or below 0.90 for closely related concepts | (Henseler et al., 2015) |
| VIF in PLS-SEM | Below 5 | (Hair et al., 2019) |
These reference points do not replace an examination of your research design. A dataset with 300 rows can still have problems if many respondents selected the same answer, left large sections blank, or answered inconsistently. Conversely, a new scale with an Alpha below the classic threshold may need to be explained in its exploratory context rather than having items deleted automatically until the number looks better.
How to Read Quantitative Data in the Output
You can start with a simple descriptive output table in SPSS. Go to Analyze > Descriptive Statistics > Descriptives, move the variables you want to inspect into Variable(s), and then select Options to display Mean, Std. Deviation, Minimum, and Maximum. The table below is illustrative output, not the result of a specific study.
| Descriptive Statistics | N | Minimum | Maximum | Mean | Std. Deviation |
|---|---|---|---|---|---|
| PU1 | 248 | 1 | 5 | 3.81 | 0.84 |
| PU2 | 248 | 1 | 5 | 3.74 | 0.91 |
| TR1 | 248 | 1 | 5 | 3.56 | 0.97 |
| INT1 | 248 | 1 | 5 | 3.92 | 0.79 |
| Valid N (listwise) | 248 |
The N column shows the number of observations used in the calculation for each variable. If one variable has a substantially lower N than the others, check for missing data. Minimum and Maximum show the actual range of values. With a Likert scale from 1 to 5, a Maximum of 7 may indicate that you entered an incorrect code or that the questionnaire used a different scale from the one you expected.
Mean is the average value. A Mean of 3.81 for PU1 only describes the response tendency in the illustrative sample. It does not show that PU1 affects the outcome variable. Std. Deviation shows the spread around the mean. A large or small standard deviation must be interpreted together with the scale and context. You should not impose one general acceptable threshold on every topic.
If you move on to scale assessment, the Reliability Statistics output may look like this:
| Reliability Statistics | Cronbach's Alpha | N of Items |
|---|---|---|
| PU | 0.842 | 4 |
In this illustrative output, the Alpha for the PU scale is 0.842 across four observed variables. You also need to inspect Item-Total Statistics, especially Corrected Item-Total Correlation and Cronbach's Alpha if Item Deleted. Do not delete an item simply because deleting it increases Alpha. Check whether the item fits the content, coding, and theoretical model.
What to Do When Quantitative Data Does Not Meet the Criteria
Work from data errors toward model problems. This order is easier to explain than deleting many rows or variables at the beginning.
Check Coding and Variable Type
Open Variable View in SPSS. Check Type, Values, Missing, and the variable labels. A Likert variable entered as a string, such as “Agree,” will not work in many numerical analyses. Establish the codebook before converting the values to numbers. Do not replace them arbitrarily.
Check Missing Data and Unusual Responses
Use Analyze > Descriptive Statistics > Frequencies to inspect frequencies and unusual values. Check rows with too many blanks, rows that use almost the same response for nearly every variable, or rows with an unusual completion time if the survey file stores that information. Define your exclusion criteria before removing rows and record how many rows were removed.
Check Reverse-Coded Variables
If the questionnaire contains negatively worded variables, recode them before running Alpha or EFA. Go to Transform > Recode into Different Variables, create a new variable, and record the coding rule clearly. For a scale from 1 to 5, the usual rule is 1 becomes 5, 2 becomes 4, and 3 stays unchanged. Apply this only when the item content was actually designed in the reverse direction.
Review the Scale and Model
If Alpha, factor loading, or AVE does not meet the criterion, reread the item wording and the scale source. An item may be measuring a different concept, may have been translated incorrectly, or may have been placed in the wrong group. Deleting an item based on both content and statistical evidence is easier to defend than repeatedly rerunning the analysis in search of a favorable result.
Collect More Data When the Problem Is the Sample
If the data lacks adequate representation, contains too few observations in one group, or includes many invalid rows, software cannot repair the problem for you. Discuss the criteria for additional data collection with your supervisor. Do not duplicate old rows to increase the sample size. That changes the data structure and creates a high risk of difficult questions at the defence.
Quantitative Data Versus Qualitative Data
Quantitative data answers questions such as how many, to what extent, whether a relationship exists, and how strong an effect is. Qualitative data usually focuses on meaning, experience, or how participants interpret an issue. You can read what is qualitative research to distinguish the two approaches at the methodological level.
| Criterion | Quantitative data | Qualitative data |
|---|---|---|
| Data format | Numbers, codes, measured scores | Text, speech, open descriptions |
| Common instrument | Closed-ended questionnaire, Likert scale | Open-ended questions or text documents |
| Purpose | Measurement, comparison, hypothesis testing | Understanding meaning and context |
| Typical results | Means, percentages, p-values, regression coefficients | Themes, descriptions, and interpretations |
| Common software | SPSS, SmartPLS, AMOS, JASP, R | Text-processing or coding software |
The distinction depends on how the data is designed and analyzed, not only on whether you can see numbers. An open response coded into topic groups can produce a frequency table, but that does not turn the entire research design into quantitative research. Conversely, gender coded as 1 and 2 is only a numerical representation of a categorical variable.
For a quantitative thesis, define these points in your proposal: which variables are independent, dependent, mediator, or moderator variables; how many observed variables measure each construct; and whether the data will be analyzed with SPSS or SmartPLS. You can also see research subjects to avoid confusing the object of study with the people who provide the data. If you need to define the scope of your topic, the research fields list can help you place the data in the right context.
Common Errors
The first error is treating every number in the file as high-quality quantitative data. A respondent ID, province code, or group code is only an identifier if it has no ordinal or numerical distance meaning. You should not calculate the mean of province codes or major codes as though they were measured quantities.
The second error is using the mean to confirm a hypothesis. The Mean describes the sample tendency, while a hypothesis about an effect usually requires a correlation, regression, or path coefficient test. Two scales can both have high Means without one scale affecting the other.
The third error is deleting items according to a mechanical threshold. An increase in Alpha after deleting an item is not enough to show that the scale is better. You need to examine Corrected Item-Total Correlation, item content, factor loading, theoretical structure, and the number of items remaining.
The fourth error is mixing criteria between SPSS and SmartPLS. CFI, TLI, and RMSEA belong to the group of indices commonly used in CB-SEM. They should not be inserted into a PLS-SEM data interpretation as though both approaches used the same criteria. In SmartPLS, focus on the relevant indicators, such as outer loading, CR, AVE, HTMT, VIF, R², and Q².
Frequently asked questions
What is quantitative data in a thesis?
It consists of numerical values collected or coded according to defined rules and used to describe the sample and test research questions. Examples include Likert scores, age, income, number of purchases, and grade point average.
Is quantitative data always an actual measurement?
No. The response “agree” can be coded as 4 on a Likert scale, but 4 only represents the response level defined by the coding scheme. Interpret the variable using its label and coding method, not only the number displayed in the file.
Which software can analyze quantitative data?
SPSS is suitable for descriptive statistics, reliability assessment, EFA, correlation, regression, and many hypothesis tests. SmartPLS is suitable for PLS-SEM when the model contains latent variables and requires assessment of both the measurement model and the structural model. JASP and R can reduce software costs, but you will usually need to learn more about the interface and syntax yourself.
What should I do if quantitative data is missing?
First, identify which variables contain missing data and what pattern the missingness follows. Then check the questionnaire, coding, and row-exclusion criteria. Replacing missing values should only be done when there is a clear methodological reason and the decision is recorded in the methods section. Do not fill values arbitrarily just to increase the number of observations.
How large should a quantitative data sample be?
Sample size depends on the number of observed variables, analysis method, number of groups to compare, and model complexity. A commonly cited reference point is 5 to 10 observations per observed variable (Hair et al., 2010), but you still need to compare it with your model and the requirements of your field.
Open your .sav or .csv file now and check the variable names, data types, missing values, and coding scale before running any analysis. If you need to run the analysis on your own SPSS or SmartPLS data, see M4 data analysis.