What Is a Population Regression Model and How to Apply It

Research models··14 min read

What Is a Population Regression Model

You can understand a population regression model as a model that describes the relationship between a dependent variable and one or more independent variables across the entire population under study. In a thesis, this concept usually appears when you want to test how strongly factors X affect outcome Y, such as the effects of service quality, perceived value, and satisfaction on repurchase intention.

The general form of a multiple linear regression model is usually written as follows:

Y = β0 + β1X1 + β2X2 + ... + βkXk + ε

Here, Y is the dependent variable, X1 through Xk are the independent variables, β0 is the constant, β1 through βk are the regression coefficients, and ε is the error term. The word “population” refers to the true relationship in the research population. Your survey data are only a sample used to estimate that relationship.

You should distinguish the population regression model from the sample regression equation. The population model is the theoretical framework you propose before analyzing the data. After running the data in SPSS, you obtain an estimated equation from the sample, such as Y = 0.42 + 0.31X1 + 0.18X2. These numbers do not automatically prove that the relationship exists in the population. You still need to examine the p-value, confidence intervals, model fit, and regression assumptions.

If you are building the theoretical framework, you can also read research model and what is theory to separate three parts clearly: the theoretical foundation, the proposed model, and the hypotheses to be tested.

Components of the Model

A population regression model in a thesis usually contains the following components. Identify each one before entering the data into SPSS or creating a model in SmartPLS.

ComponentCommon notationMeaning in the study
Dependent variableYThe outcome that the study aims to explain or predict
Independent variableX1, X2, X3Factors assumed to affect Y
Constantβ0The expected value of Y when the X variables equal 0
Regression coefficientβ1, β2, β3The direction and amount of change in Y when one X changes, while the other variables are held constant
ErrorεThe part of the variation in Y that is not explained by the X variables in the model
Research samplenThe number of observations used to estimate the population model

Dependent Variable

The dependent variable is the outcome you want to explain. For example, if your study examines factors affecting online purchase intention, “online purchase intention” can be Y. In a questionnaire, Y is usually measured with multiple items and then combined into a construct before you run the regression.

Do not choose Y simply because it is easy to survey. Y must connect to your research question and theoretical model. If Y is a latent construct, check the reliability of the scale before entering its mean score or representative factor score into the regression.

Independent Variables

Independent variables are the factors placed on the explanatory side of the model. For example, in a repurchase intention model, X1 might be product quality, X2 perceived price, X3 trust, and X4 satisfaction.

Each independent variable needs a theoretical reason for appearing in the model. Adding many variables only to search for small p-values can make the model difficult to explain, create multicollinearity, and lead the committee to ask why each variable was selected.

Coefficients and Error

The coefficient β shows the direction of the effect when the other conditions remain unchanged. A positive coefficient indicates that X and Y vary in the same direction in the model. A negative coefficient indicates an inverse relationship. The sign of the coefficient is only part of the conclusion. You must also check the p-value or confidence interval to assess the statistical evidence.

The error term ε reminds you that the model does not explain all of Y's behavior. A social science study usually includes many factors outside the model. Read a high or low R² in the context of the topic, the number of variables, and the research objective. Do not treat one number alone as the rule for deciding whether a model is good or bad.

The Original Model and Common Extensions

A population regression model is an analytical framework, while each specific model should come from an appropriate theory. For technology acceptance research, the Technology Acceptance Model proposed by Davis (1989) commonly focuses on perceived usefulness and perceived ease of use. For intended behavior, the Theory of Planned Behavior proposed by Ajzen (1991) includes attitude, subjective norm, and perceived behavioral control.

If your study examines technology acceptance in an organizational setting, the UTAUT model of Venkatesh et al. (2003) is a commonly used foundation. UTAUT2 by Venkatesh et al. (2012) extends the context to consumers through components that better fit individual technology-use behavior.

The original model is a starting point rather than a list of variables that must be copied exactly. When applying it to your topic, you can retain core variables, remove variables that do not fit, or add variables supported by the literature. Explain every change before you formulate the hypotheses.

For example, the original model may propose that perceived usefulness affects intention to use. In your study, an electronic payment context may support adding trust and perceived risk. The extended model then needs to explain why these two new variables fit the context, where the scales came from, and what direction of relationship you expect.

You should also distinguish “hypothesis” from “assumption” when writing Chapter 2. The article what is a hypothesis explains how to turn relationships between concepts into testable statements. If you are unsure which term to use, see assumption or hypothesis.

Common Scales for Each Construct

A regression model is only as credible as the measurement of the constructs it contains. Do not use one isolated question to represent a complex construct when the foundational literature uses multiple items.

ConstructScale or reference sourceUse in the model
Perceived usefulnessDavis (1989)Can be an independent variable explaining intention to use
Attitude, subjective norm, perceived behavioral controlAjzen (1991)Independent variables in an intended behavior model
Technology acceptanceVenkatesh et al. (2003)Use UTAUT components according to the research context
Consumer technology acceptanceVenkatesh et al. (2012)Can extend the model to a consumer context
Service qualityParasuraman et al. (1988)Can measure service quality dimensions and their effects on satisfaction

These sources are starting points for building a questionnaire, not confirmed Vietnamese translations for every context. Check the wording, survey population, and usage context. Two studies can both use the construct “trust” and still require different items if one examines digital banking and the other examines an e-commerce platform.

After collecting the data, you can run Cronbach's Alpha to check reliability. Cronbach's Alpha of 0.7 or above is commonly accepted according to (Nunnally, 1978), while Corrected Item-Total Correlation of 0.3 or above is discussed in (Nunnally and Bernstein, 1994). For a new scale or exploratory research, Alpha from 0.6 may be considered according to (Hair et al., 2010), but explain the context instead of using this threshold to keep every item.

If you run EFA, KMO of 0.5 or above and a statistically significant Bartlett's test are commonly used criteria according to (Kaiser, 1974). Factor loading of 0.5 or above and total variance explained of 50% or above are referenced in (Hair et al., 2010). When an item loads strongly on several factors, examine the theoretical content, loading difference, and effect of deleting the item before deciding.

Applying the Model to Your Study

Suppose you are studying factors affecting customers' repurchase intention on an e-commerce platform. You select four independent variables: information quality, trust, perceived value, and satisfaction. The dependent variable is repurchase intention.

The conceptual model can be presented with four arrows pointing toward Y:

  • X1: Information quality → Y: Repurchase intention.
  • X2: Trust → Y: Repurchase intention.
  • X3: Perceived value → Y: Repurchase intention.
  • X4: Satisfaction → Y: Repurchase intention.

From this model, write the hypotheses as complete statements:

  • H1: Information quality has a positive effect on customers' repurchase intention.
  • H2: Trust has a positive effect on customers' repurchase intention.
  • H3: Perceived value has a positive effect on customers' repurchase intention.
  • H4: Satisfaction has a positive effect on customers' repurchase intention.

After testing the measurement scales, you can estimate a sample equation in the following form: Y = β0 + β1X1 + β2X2 + β3X3 + β4X4. If β2 is positive and statistically significant, the data provide evidence supporting H2 within the scope of your sample and research design. This wording is more careful than saying that “trust definitely causes repurchase intention,” because survey regression usually reflects a statistical relationship and does not automatically prove an absolute causal relationship.

How to write this in your thesis

You can adapt the following paragraph for Chapter 2 or Chapter 4: “Based on [theory name] and previous studies, this study proposes [number] independent variables, including [X1], [X2], [X3], and [X4], to explain [Y]. The regression results show that [X] has a coefficient of [β] = [value], p-value = [value]. Therefore, hypothesis [H] is [accepted/not accepted] at the [significance level] level.”

In the discussion section, connect the results to the model rather than merely ranking β values. Explain which variable has the stronger effect when the other variables are present, whether the result has the same direction as the hypothesis, and whether it differs from the foundational study. If the data do not support a hypothesis, report the result accurately instead of modifying the model until every hypothesis is accepted.

Testing the Model with SPSS or SmartPLS

SPSS is suitable when you have combined the items into representative scores and want to run linear regression, ANOVA, a t-test, or PROCESS. For multiple linear regression in SPSS, the usual menu path is Analyze > Regression > Linear. Move Y into the Dependent box and X1 through X4 into Independent(s), then open Statistics and select Estimates, Model fit, Collinearity diagnostics, and Confidence intervals.

In the Model Summary table, examine R and R Square. In the ANOVA table, examine the significance of the F test to assess whether the overall model fits the data under the test you selected. In the Coefficients table, read the B column if you want to write the equation in the original units. Read Standardized Coefficients Beta if you want to compare the relative strength of the variables. Use t and Sig. to test each coefficient.

You also need to check VIF and Tolerance in Coefficients. VIF below 5 is a reference threshold in PLS-SEM reporting according to (Hair et al., 2019), but in SPSS you still need to consider the context and correlations between variables. For the residuals, you can inspect the Histogram, Normal P-P Plot, and Scatterplot. Durbin-Watson between 1 and 3 is a criterion commonly mentioned according to (Field, 2013), but do not use this number alone to conclude that all regression assumptions have been met.

SmartPLS is suitable when your model contains latent variables measured by multiple indicators, mediation or moderation relationships, or when you need to assess the measurement model and structural model together. The PLS-SEM procedure is usually divided into two parts: assess the measurement model first, then assess the structural model according to (Anderson and Gerbing, 1988) and the procedure presented in (Hair et al., 2022).

In SmartPLS, check Outer Loadings, CR, AVE, and HTMT before reading Path Coefficients. Outer loading of 0.7 or above is a reference threshold according to (Chin, 1998) or (Hair et al., 2022). CR of 0.7 and AVE of 0.5 or above are referenced in (Fornell and Larcker, 1981). HTMT below 0.85, or below 0.90 for closely related constructs, is discussed in (Henseler et al., 2015).

When assessing the structural model, examine Path Coefficients, Standard Deviation, T Statistics, P Values, and Confidence Intervals from bootstrapping. The current PLS-SEM procedure commonly uses 5,000 bootstrap subsamples according to (Hair et al., 2022). For R², the levels 0.75, 0.50, and 0.25 are described as substantial, moderate, and weak, respectively, in a PLS-SEM guide by (Hair et al., 2011). These are interpretation points, not rules for rejecting a model simply because its R² is low.

If your model is a simple observed-variable regression, SPSS is usually easier to explain at the defense. If the model has latent variables and several relationships between constructs, SmartPLS may be more suitable. AMOS follows the CB-SEM approach and provides fit indices such as CFI, TLI, RMSEA, and SRMR according to (Hu and Bentler, 1999). These indices should not be included in an article about SmartPLS-specific evaluation criteria.

Common Mistakes When Using This Model

Calling Every Equation a Population Regression Model

The equation you obtain from a .sav or .csv file is an estimate from a sample. In your thesis, state clearly that it is an estimated regression equation, and use the test results to make careful inferences about the population.

Adding Variables Without a Basis

A variable's correlation with Y in preliminary data is not enough to place it in the theoretical model. Explain the scale source, expected relationship, and reason it fits the research context.

Looking Only at the p-value

The p-value shows statistical evidence under the selected test, but it does not describe the full quality of the model. Read it together with the coefficient, confidence interval, R², VIF, and assumption checks. A coefficient with a small p-value still needs attention if VIF is high.

Deleting Variables Until the Model Looks Good

EFA or regression may not look good on the first run. You can make adjustments when you have a methodological basis, but record each run, the reason for deleting an item, and the effect on the model. Deleting items only because it changes the p-value will be difficult to explain at the defense.

Confusing Regression with Causality

Cross-sectional survey data can show the strength of an association and explanatory ability, but they are not enough to claim that one variable definitely causes another. Your wording should follow “has an effect under the tested model” or “has a statistically significant relationship,” depending on the research design.

Frequently asked questions

Is a population regression model the same as a regression equation?

The two concepts are related, but they are not the same thing. A population regression model is the theoretical relationship between Y, the X variables, and the error term in the population. The regression equation calculated from survey data is an estimate of that relationship in the sample.

Does the population regression model need to appear in Chapter 2?

Present the conceptual model and general equation in Chapter 2 or the methodology chapter, depending on your department's requirements. Chapter 4 presents the estimated equation and the Model Summary, ANOVA, and Coefficients tables from the actual data.

How many independent variables are appropriate in a model?

There is no fixed number for every study. The number should be based on theory, the research objective, the number of items, and the sample size. For regression, a commonly cited guideline is a minimum sample size of 50 + 8m, where m is the number of independent variables, according to (Tabachnick and Fidell, 2013). You still need to consider the questionnaire design and your supervisor's requirements.

Should I run the population regression model in SPSS or SmartPLS?

Choose SPSS when you already have composite variables and your main objective is linear regression, model fit testing, and testing the effects of variables. Choose SmartPLS when the model includes latent variables, multiple indicators, mediation, or moderation. The software does not replace the theoretical basis or measurement procedure.

What should I do if a hypothesis is not accepted?

Keep the result, then check the coding, reverse-coded items, missing data, scale, and model assumptions. If the procedure contains no error, report that the data did not provide sufficient evidence supporting the hypothesis in this context and sample, then discuss possible reasons.

Open your data file again and list Y, X1 through Xk, the scale source, and the corresponding hypothesis before running the model. Then match each output table to your research question. If you need to complete the analysis using your own .sav or .csv file, see DoThesis M4 data analysis.