
Sample Size Formulas for Quantitative Research
What is sample size in quantitative research
Sample size is the number of valid responses used in a quantitative study. If you distribute 350 questionnaires but only 312 responses pass your checks, your analysis sample size is 312. Keep these two numbers separate in the methods chapter because the number of responses collected and the number retained for analysis are not always the same.
Your sample needs to be large enough for the data to represent the target population reasonably well and to support the planned analysis. A simple descriptive survey, linear regression, EFA, and PLS-SEM may require different justifications. Identify your model, number of items, and analysis method before deciding how many responses to collect.
You can review the background explanation in what is sample size, then return to the formula that fits your study. A formula is only the starting point. Sample quality also depends on how you select participants, your screening criteria, and the number of valid responses.
How to determine sample size
Use the Yamane formula when the population size is known
When you have a reasonably clear finite population size, you can use the Yamane formula:
n = N / (1 + N × e²)
Here, n is the minimum sample size, N is the population size, and e is the allowable error. For a 5% error, replace e with 0,05. For example, if the population list contains 10.000 people, the minimum sample size from the formula is approximately 385. This is the minimum under the formula's assumptions. It does not include invalid responses or nonresponses. The finite-population formula is presented in (Yamane, 1967).
If your population is the customers of an app with 50.000 accounts, explain whether 50.000 means the number of people within the research scope or only the number of registered accounts. One person with multiple accounts can change how you define the population. When you do not have a reliable population list, applying the Yamane formula mechanically will be difficult to defend.
Use the Cochran formula when the population is unknown
For a proportion survey in a large population or one whose size is unknown, the Cochran formula is commonly written as:
n₀ = Z² × p × (1 - p) / e²
Z is related to the confidence level, p is the estimated proportion of the characteristic under study, and e is the allowable error. When you do not have a reliable estimate for p, state the assumption used in your proposal and apply it consistently in the calculation. The basis for calculating sample size for a proportion is presented in (Cochran, 1977).
Do not put a number in your thesis without stating the population size, error, confidence level, or related assumptions. If the committee asks, “Why did you choose 300 responses?”, your answer should lead back to the formula and the target population, rather than simply saying that 300 is commonly used.
Use the number of items for EFA and quantitative models
If your study includes EFA, regression, or a model with several scales, the number of items gives you a practical basis. A commonly cited rule is 5 to 10 observations per item (Hair et al., 2010). For example, a questionnaire with 30 items would have a reference range of 150 to 300 valid responses.
This rule does not replace a review of the model. If you have several factors, many independent variables, or expect to remove some items during EFA, set a target above the minimum. For regression, the reference formula is n ≥ 50 + 8m, where m is the number of independent variables (Tabachnick and Fidell, 2013). This supports regression planning, not every type of quantitative study.
Quick reference by situation
The table below is illustrative output.
| Situation | How to determine it | Basis to report in the thesis |
|---|---|---|
| Finite population size is known | n = N / (1 + N × e²) | Formula from (Yamane, 1967) |
| Population size is unclear | Cochran formula for a proportion | (Cochran, 1977) |
| Many items and EFA is planned | 5 to 10 observations per item | (Hair et al., 2010) |
Regression with m independent variables | n ≥ 50 + 8m | (Tabachnick and Fidell, 2013) |
| Factor analysis requires adequacy checks | Also review sample quality and factor structure | (Comrey and Lee, 1992) |
For example, you have 24 items and 5 independent variables. Using the 5 to 10 rule, the reference sample size is 120 to 240. Using the regression formula, the minimum is 50 + 8 × 5 = 90. When the two bases produce different numbers, choose the higher level you can realistically implement and explain both checks. If you expect invalid responses, distribute 10% to 20% more than your target analysis sample, but describe this as a reserve rather than a new minimum sample size.
Designing the questionnaire for this objective
Sample size only matters when the questionnaire measures the concepts in your study correctly. Before distributing it, make a list of latent variables, items, scale sources, and coding rules. A scale with four items should use clear codes such as DV1, DV2, DV3, and DV4, rather than long and inconsistent column names.
Choose the question types
A quantitative questionnaire commonly includes screening questions, Likert-scale measurement items, and demographic questions. Screening questions establish whether a respondent belongs to the target group, such as having used a service within the past six months. Demographic questions describe the sample or support grouping. Do not put every demographic variable into the model without a research reason.
State the scores and the meaning of both endpoints of the Likert scale. A 5-point scale can run from strongly disagree to strongly agree. The foundation of the Likert attitude scale is presented in (Likert, 1932). Keep the interpretation in the same direction across items, and mark any items that require reverse coding before importing the file into SPSS.
Estimate how many questionnaires to distribute
Separate three numbers: the minimum analysis sample, the target number of valid responses, and the number of questionnaires to distribute. For example, you need 250 valid responses and expect 10% of responses to be excluded, so the number distributed must be higher than 250. If the response rate is low, also account for people who are invited but do not complete the questionnaire.
You can review how to organize what is a questionnaire and the structure of an online survey form before creating the final version. Open the questionnaire on a phone, check required items and page logic, and verify the coding of responses before sending the link widely.
Check the questionnaire before official data collection
At the checking stage, focus on display errors, unclear wording, missing response options, and completion time. If an item can be understood in two different ways, the resulting data will be difficult to interpret even when the sample is large.
Also check whether exported columns have the correct variable names, order, and numeric format. Code “Not applicable” separately rather than automatically entering it as a middle value on the Likert scale. Record these rules in the data-coding guide.
Conducting data collection
State clearly where the sample comes from, when data collection takes place, and which criteria apply. If your study concerns users of an e-commerce platform, define who counts as a user instead of writing only “customer survey”. Convenience sampling may fit the limits of a thesis, but you need to report its limits for generalization.
You can create an online survey questionnaire with a suitable tool. fillform.info is DoThesis's own form product, built to create questionnaires and collect responses in Vietnamese. Google Forms is a common alternative if you want to create the form yourself within the Google ecosystem. Whichever tool you use, download the data and inspect every column before importing it into SPSS.
During collection, save versions of the questionnaire, the start date, the end date, and the number of responses per day. If you change an item halfway through collection, the data before and after the change may no longer be fully comparable. If a change is necessary, record when it occurred and what was changed so you can explain it in the methods chapter.
Do not stop immediately after collecting exactly the minimum sample size. Some responses will be removed because of missing data, unusually fast completion, or signs that the respondent selected the same option across an entire scale. Decide the reserve based on access to participants, time, and budget. Do not present it as a mandatory percentage for every study.
Cleaning data before importing it into SPSS
First, check rows with substantial missing data, responses that fail the screening questions, and duplicate records. If the questionnaire requires every item to be answered but a respondent leaves many items blank, set the exclusion criterion before analysis. Apply it consistently to the full dataset.
Next, check straight-lining, where a respondent selects the same level across many items without evidence of reading and distinguishing their content. This is a signal to investigate, not by itself sufficient evidence for exclusion. Combine completion time, screening responses, and unusual response patterns before deciding.
Then check variable coding. Columns used for Cronbach's Alpha and EFA must be numeric variables, not text strings. Reverse-code items before running the reliability test. Columns such as submission time, email, access code, or internal notes should not enter the model unless they are research variables.
Finally, save two separate files. One should be the untouched raw data, and the other should be the cleaned data used for analysis. Version names such as data_raw, data_clean_v1, and data_clean_v2 make it easier to investigate when an advisor asks why the number of responses differs between tables.
How to write this in your thesis
The sample-size section in Chapter 3 should answer five questions: who is the population, how was the sample accessed, which formula was used, what was the minimum sample size, and how many valid responses were actually obtained. Also state the collection period and exclusion criteria if the data went through screening.
You can write it using this template: “The study determined a minimum sample size of [sample size] based on [formula or rule], where [state the key variables]. After data collection and cleaning, the study obtained [number of valid responses] responses for [EFA/regression/PLS-SEM].” Replace every bracketed placeholder with the actual numbers in your file.
When you use more than one basis, state which one determined the final number. For example, you can report that the item-based requirement was 200 and the regression requirement was 90, so the target was set at a minimum of 200 valid responses to support EFA. This is easier to check than listing several formulas without connecting them to the final decision.
When reporting the sample, include a table showing the number distributed, the number returned, the number excluded, and the number of valid responses. Do not put illustrative numbers into Chapter 3. The table in this article explains the reasoning, while your thesis must use counts from the actual data.
Common mistakes
Choosing a formula because the number looks familiar
Many students choose 200 or 300 because they see those numbers in other theses. The number in another study may reflect a different population, model, or analysis method. Start with your research objective, then check at least one population-based basis and one model-based basis.
Confusing distributed questionnaires with valid sample size
If you distribute 300 questionnaires and exclude 28, you cannot report an analysis sample size of 300. Record each step separately so the figures in Chapters 3 and 4 and the appendix do not conflict.
Collecting from the wrong population
A complete response still has no value if the respondent does not belong to the study population. Put screening questions before the scale items, and define the exclusion criterion before collection begins.
Deleting data only to reach the target sample size
Do not remove responses arbitrarily to make Cronbach's Alpha or EFA look better. Each excluded response needs an observable reason and a record. If the data do not meet the requirements, you may need to review the questionnaire, sampling approach, or analysis plan rather than deleting more rows.
Using Yamane when N is unknown
The Yamane formula requires a population size N. If you have no basis for determining N, describe the information limit and choose a more suitable basis. Inventing a large N to produce a plausible-looking number will invite questions during the defense.
Frequently asked questions
What sample size is suitable for a quantitative study?
There is no single number that fits every study. Compare the population size, number of items, number of independent variables, and analysis method. For EFA, 5 to 10 observations per item is a commonly used basis (Hair et al., 2010), while regression can also be checked using n ≥ 50 + 8m (Tabachnick and Fidell, 2013).
Which sample size formula should I use when the population is unknown?
Consider the Cochran formula for estimating a proportion (Cochran, 1977). In the thesis, state the assumptions about confidence level, error, and estimated proportion. If your study focuses on EFA or a model with several scales, also check the requirements of that analysis rather than relying only on a proportion-survey formula.
Should I collect more responses than the minimum sample size?
A reserve is useful because some responses may contain missing data, come from the wrong population, or show straight-lining. The reserve depends on the collection channel and the expected response quality. In your report, separate the number required for analysis from the number distributed and the number excluded.
What should I do if the valid sample is smaller than expected?
First, review the exclusion criteria and check whether any responses were removed because of data-entry errors. If the valid number is still low, discuss with your advisor whether to continue collecting data, adjust the population scope, or limit the analyses. Do not lower the standard only after seeing the test results.
Can I use the same sample for SPSS and SmartPLS?
A dataset can be used in both SPSS and SmartPLS if the research design, measures, and data quality support both plans. However, the sample-size justification must be connected to the specific model and analysis procedure. Do not take the threshold from one method and claim that every other analysis is automatically supported.
Open your questionnaire file and data file again. Write down the number of items, number of independent variables, and population size if known, then calculate the sample size using the most appropriate basis before continuing collection or cleaning. If you need to run the formula and check real data in .sav or .csv, you can use DoThesis M4 analysis.