
What Is the Research Population? How to Define and Sample It Correctly
What Is a Research Population
You can understand the research population as the group of people, organizations, businesses, households, or units from which you approach participants and collect data for your study. More specifically, the research population answers the question: “Who or which units does the research data come from?”
For example, a study analyzing the factors affecting university students’ intention to use e-wallets in Ho Chi Minh City has a research population consisting of students currently studying at universities and colleges in Ho Chi Minh City. If you distribute a questionnaire to 250 students and enter their responses into SPSS, those 250 people are the observation units in your sample, while the group of students targeted by the study is the research population.
You need to describe the population using verifiable conditions, such as age, occupation, place of residence, product-use status, or relationship with an organization. A useful description usually has four parts: who, where, during what period, and what participation conditions apply.
You can read what scientific research is to place this concept within the full research process. The population determines whom you distribute the questionnaire to, which screening criteria you use, and which group you can generalize your results to.
Why the Research Population Matters in Quantitative Research
In quantitative research, the population defines the scope of inference for your results. When you run Cronbach's Alpha, EFA, linear regression, or PLS-SEM, the software processes only the data from people who submitted responses. Your results therefore have meaning only within the population you defined and sampled.
Suppose you survey people who purchased cosmetics online within the past six months. If you distribute the questionnaire mainly to students who have never purchased cosmetics online, the data no longer match the original population. The output tables may still produce numbers, but your conclusions about the behavior of people who have made such purchases will lack a sound basis.
Defining the population leads to three practical decisions. The first is the screening criteria, such as whether respondents have used the product, made a transaction, or belong to the required age group. The second is the data collection channel, such as a Facebook group, classroom, business, user community, email, or online questionnaire. The third is the sampling method and sample size, followed by an explanation of why the number of observations is appropriate for the model.
You also need to distinguish the research field, geographical area, time period, and research population. You can refer to research field when defining the topic boundary, then use the population to turn that boundary into a specific data group.
How Large Should the Research Population Be
There is no fixed population size that applies to every study. Here, “adequate” means that your description is clear enough for another person to know whom you surveyed, check whether respondents belong to the group, and determine whether the actual sample relates to the research question.
| Checking criterion | Question to answer | How to present it | Basis |
|---|---|---|---|
| Correct research group | Are the respondents directly related to the topic? | State the required behavior, role, or experience | Research objective and research question |
| Screening conditions available | Which question verifies eligibility to participate? | State the screening question at the beginning of the questionnaire | Questionnaire design |
| Geographical area or context available | Where was the data collected? | State the province, city, school, business, or platform | Research scope |
| Time reference available | Which period does the experience refer to? | State the month, quarter, or survey period | Data collection plan |
| Appropriate sample size | Are there enough observations for the planned method? | Explain it based on the number of items or independent variables | (Hair et al., 2010); (Tabachnick and Fidell, 2013) |
| Accessible participants | Can you actually obtain valid responses? | State the access source and questionnaire distribution method | Data implementation feasibility |
If your scale has many items, a commonly used rule is 5 to 10 observations for each item (Hair et al., 2010). For regression, you can refer to a minimum sample size of n at or above 50 + 8m, where m is the number of independent variables (Tabachnick and Fidell, 2013). This provides a basis for explaining sample size. It is not a reason to change the population simply to reach a target number of respondents.
For example, suppose your study has 25 items and you plan to run EFA. The range from 125 to 250 observations is a reference point under the rule above. You still need to review the number of items, number of factors, sampling method, and your supervisor’s requirements before finalizing the number.
How to Check the Research Population in the Output
The population does not appear as a separate statistic in the Coefficients, KMO and Bartlett's Test, or Reliability Statistics tables. You check it through the sample-description variables and screening questions in your data file. In SPSS, the Frequencies and Descriptive Statistics tables show whether the actual sample matches the group you described.
Suppose you are studying office workers’ intention to use digital banking applications in Hanoi. The questionnaire includes the SCREEN_USE variable to identify respondents who have used the application within the past six months, and the AGE_GROUP variable to describe age. The table below is illustrative output, not the result of an actual study.
| Frequencies | Frequency | Percent | Valid Percent | Cumulative Percent |
|---|---|---|---|---|
| SCREEN_USE = Yes | 214 | 89.2 | 89.2 | 89.2 |
| SCREEN_USE = No | 26 | 10.8 | 10.8 | 100.0 |
| Total | 240 | 100.0 | 100.0 |
If the participation criterion is having used the application within the past six months, the 26 records with the value No need to be considered for removal from the analysis sample. You also need to check how missing values are coded. A blank cell, code 99, or the response “Do not remember” should not be placed in the group of users if the questionnaire did not define it that way.
In SPSS, open Analyze > Descriptive Statistics > Frequencies, move the screening and demographic variables into the Variables box, and then request the frequency table. To check the mean age, standard deviation, or minimum range, open Analyze > Descriptive Statistics > Descriptives. These tables describe the sample you obtained. They do not prove that the sample fully represents the population.
When reading the output, compare three places: the population criteria in the methods chapter, the screening question in the questionnaire, and the actual values in the .sav or .csv file. If these three parts do not match, clean the data before running the main analysis.
What to Do When the Research Population Does Not Meet the Criteria
First, identify whether the problem is in the description or in the data. If the description is still too general, rewrite the criteria. If many respondents belong to the wrong group, check the screening question and the record-exclusion rule. If the number of valid observations is low, then consider collecting more data.
Review the Participation Criteria
Write each criterion as a condition that can be answered with Yes or No. For example, “Purchased from e-commerce platform X at least once within the past three months” is clearer than “is an online consumer.” The more specific the criterion, the easier it is to explain why a record was retained or removed.
Check Coding and Missing Data
Open Variable View to check Values, Missing, and the data type of the screening variable. Then use Analyze > Descriptive Statistics > Frequencies to detect unexpected codes. If code 1 means “Yes” but some rows contain code 3, trace the questionnaire or coding convention before filtering the data.
Remove Ineligible Records With a Clear Record
You can use Data > Select Cases to select cases that meet the criteria. Save an original data file, create a separate cleaned version, and record the number of removed records and the reason for each removal. Data removal must be based on criteria defined in advance. Do not remove a record merely because it lowers Cronbach's Alpha or R².
Collect More Data if the Valid Sample Is Too Small
If the valid number of observations is below your plan after cleaning, return to the correct population. Do not expand to another group just to reach the required number. If you need to change the geographical area, age range, or usage condition, update the research scope and discuss the change with your supervisor.
If you use an online questionnaire, you can deploy it through Google Forms or fillform.info, DoThesis’s Vietnamese form tool for collecting survey responses. Whichever tool you use, responsibility for checking the population still rests on the screening questions, the exported data, and the cleaning rules.
Distinguishing the Research Population From the Research Subject
These two terms are often mixed up because both appear in the opening sections of a study. An easy distinction is this: the research subject is the issue or relationship you want to analyze, while the research population is the group that provides the data for analyzing that issue.
For example, consider the study “Factors affecting customer satisfaction with food delivery services in Hanoi”:
- Research subject: the factors affecting customer satisfaction with food delivery services.
- Research population: customers who used food delivery services in Hanoi during the specified period.
- Observation unit: each customer who provided a valid response in the sample.
- Research scope: Hanoi, the selected service type, and the survey period.
You can also read research subject to compare how these two sections are written. A study can have the same population but a different research subject. For example, you could survey the same student group to study purchase intention, satisfaction, or usage behavior.
You also need to distinguish the population from the target population of the study. The target population is the entire group you want to address, while the sample is the part selected to answer the questionnaire. The population describes the target group, while the sample shows who actually appears in the data file.
If you need to place this concept within the broader content system, you can read the Scientific Research topic and scientific research.
Common Mistakes
Defining the population too broadly. “Vietnamese consumers” may be beyond the data collection capacity of a student thesis. If you survey only one city or one age group, state that actual scope.
Confusing the population with a product or phenomenon. “Satisfaction with digital banking” is the content to be measured, not the group of respondents. The population must be users or an organizational group with relevant experience.
Having no screening question. You describe the population as people who have purchased a product, but the questionnaire allows anyone to respond. In that case, the data file cannot demonstrate that respondents belong to the research group.
Changing the population after collecting data. You initially survey office workers, then add students while keeping the original sample description. This creates inconsistency between the proposal, questionnaire, and results chapter.
Generalizing beyond the scope. A convenience sample from one university is not enough to state that the results represent all university students in Vietnam. State the limits of the geographical area, sampling method, and respondent group.
Reporting only the number of responses without the eligibility conditions. Having 300 rows of data does not mean that you have 300 suitable observations. If your study process includes these steps, report the number of questionnaires distributed, questionnaires returned, questionnaires excluded, and valid questionnaires.
Frequently asked questions
What is a research population, and must it be stated in a thesis?
The research population is the group of people or units that provides data for the study. In a quantitative thesis, you should state the population in the research subject, scope, or methods section so that readers know which group the results are based on.
How are the research population and research subject different?
The research subject is the issue, phenomenon, or relationship to be analyzed. The research population is the group surveyed or providing the data. For example, “online cosmetics purchase intention” is the research subject, while “consumers aged 18 and above who have purchased cosmetics online” is the research population.
Can a study have more than one research population?
It can, if the model and research objectives genuinely require data from multiple groups. You then need to state each group, the sampling method, and how you will compare or combine the data. If you add a group only to reach a target number, the research design will be difficult to explain.
Is the research population the same as the sample size?
No. The population is the target group, while sample size is the number of selected units with valid data. For example, the population may be students using e-wallets in Ho Chi Minh City, while the sample size may be 250 valid questionnaire responses.
Do you need to rerun SPSS when the research population is wrong?
You need to check and clean the data first. If the problem is limited to an incorrect code or a few records that do not meet the criteria, correct the data according to the documented rule and rerun the affected analyses. If you must change the survey group, review the questionnaire, sample size, and methods presentation before rerunning the full process.
How should you write the research population in a thesis?
How to write this in your thesis
You can write: “The research population consists of [people or units] in [location], who meet [participation conditions] during [time period]. Data were collected through [method]. After removing questionnaires that did not meet [screening criteria], [number] valid observations remained for analysis.” Replace every bracketed section with the actual information in your data file.
Open your questionnaire file and data file now, compare the population criteria with the screening variable, and record the number of valid questionnaires before running the next analysis. If you need help running the analysis on your .sav or .csv file, see M4 analysis.