What Is a Research Hypothesis? How to Build and Write One

Research models··12 min read

What Is a Research Hypothesis?

A research hypothesis is a statement about a relationship between two or more variables that can be tested with data. In a quantitative thesis, a hypothesis usually predicts that an independent variable affects a dependent variable, or explains how a mediator or moderator changes that relationship.

For example, the statement “Service quality affects customer satisfaction” only describes a research relationship. When you write it as a hypothesis, you need to state the direction of the effect clearly: “H1: Service quality has a positive effect on customer satisfaction.” You can then test this relationship with questionnaire data and an appropriate model.

You should distinguish a hypothesis from a research question. A question states what you need to find out, while a hypothesis is a proposed answer that the data may support or reject. If you are confusing “assumption” and “hypothesis,” see Assumption or hypothesis. This confusion often leads students to use the wrong term from Chapter 1 onward.

A good quantitative hypothesis usually contains four elements: the name of the influencing variable, the name of the affected variable, the direction of the relationship, and the research context. For example: “Perceived usefulness has a positive effect on students’ intention to use digital banking applications in Ho Chi Minh City.”

Components of a Hypothesis Model

Before writing H1, H2, or H3, you need to identify the variables in your research model. A model is not just a diagram with arrows. Each concept needs a role, a definition, and an appropriate measurement approach.

ComponentMeaning when building a hypothesisExample
Independent variableA factor predicted to be a cause or influenceService quality
Dependent variableThe outcome that needs to be explained or predictedCustomer satisfaction
MediatorExplains the mechanism through which an independent variable affects a dependent variablePerceived value
ModeratorChanges the strength or direction of a relationshipUsage experience
Control variableIncluded to reduce the influence of background factorsAge, gender, income

The independent variable usually appears on the left side of the diagram, with an arrow pointing toward the dependent variable. A variable can be an independent variable in one hypothesis and a dependent variable in another relationship. For example, perceived value may be affected by service quality and then affect repurchase intention.

A mediator answers the question, “Through what mechanism does the effect occur?” If service quality increases satisfaction and satisfaction increases intention to return, satisfaction may be proposed as a mediator. The logic for distinguishing a mediator from a moderator is presented in (Baron and Kenny, 1986), while bootstrapping the indirect effect can be examined using (Preacher and Hayes, 2008).

You should read What Is a Research Model before drawing the diagram. If you have not identified which variable is the cause, outcome, or mediating mechanism, assigning labels from H1 to H4 is only arranging sentences. It is not yet model building.

Original Models and Common Extensions

A hypothesis should come from a theory, an original model, or previous research related to your topic. You do not need to copy an existing model exactly. Your task is to identify the model that explains the behavior or outcome you are studying, then adjust the variables to fit your context.

For example, a technology acceptance topic may begin with TAM. In TAM, perceived usefulness and perceived ease of use are used to explain acceptance or behavioral intention. The original foundation of this model is commonly attributed to Davis (1989). If your study examines technology use in a broader context, you can consider the TPB of (Ajzen, 1991), UTAUT of (Venkatesh et al., 2003), or UTAUT2 of (Venkatesh et al., 2012).

An original model helps you answer three questions. First, why could this variable affect that variable? Second, which theory explains the relationship? Third, which scale can be adapted for your questionnaire? Every variable you add needs a reason. Do not add variables simply because the diagram has too few arrows.

A reasonable extension usually takes one of three forms. You may add a new variable to reflect the context of Vietnam, add a mediator to explain the mechanism, or add a moderator to test differences between groups. For example, a TAM model may be extended with trust, perceived risk, or social influence if these variables fit the research question.

When consulting What Is a Theory, record the author, publication year, main variables, and proposed relationships. Do not write “according to many previous studies” and leave the source blank. The committee may immediately ask which studies you mean, which context they examined, and which part you are carrying forward.

Scales Commonly Used for Each Concept

A hypothesis tells you how the variables are related, while a scale tells you which items measure those variables. These two parts must match. If your hypothesis concerns service quality but your questionnaire measures only price, you will not be able to defend the proposed relationship.

ConceptReference scale or modelHow to use it in the study
Service qualitySERVQUAL by (Parasuraman et al., 1988)Adapt the components to the specific type of service
Perceived usefulnessTAM by (Davis, 1989)Measure the extent to which users feel that the technology helps them achieve their goals
Behavioral intentionUTAUT by (Venkatesh et al., 2003)Adapt it to use or continued use behavior
Purchase intentionCan be developed from the theory of planned behavior by (Ajzen, 1991)Examine intention in a product or service context
AttitudeA scale associated with TPB by (Ajzen, 1991)Measure positive or negative evaluations of the behavior

These sources are starting points, not the final wording for your questionnaire. Read the original definition, check the number of items, examine the sample context, and review how the authors measured the concept. Then translate and adjust the wording so that respondents understand it correctly.

If a variable has several possible measurement sources, create a comparison table with the item code, original wording, proposed translation, and reason for retaining or removing it. Do not rename items simply to make Cronbach's Alpha look better. Item removal should be based on both substantive meaning and test results.

In quantitative research, scales are often checked in this order: Cronbach's Alpha, EFA if the model uses SPSS, and then regression or a structural model. With PLS-SEM, you will examine outer loading, CR, AVE, and HTMT. These criteria belong to measurement model assessment, while hypotheses belong to testing relationships between variables.

Applying the Model to Your Topic

Suppose you are studying “students’ intention to continue using online learning applications.” After reviewing the literature, you select four independent variables: perceived usefulness, perceived ease of use, content quality, and social influence. The dependent variable is intention to continue using the application.

The conceptual model can be written as follows: perceived usefulness, perceived ease of use, content quality, and social influence all affect intention to continue using the application. If you want to test the mechanism, satisfaction can be placed between the independent variables and continued-use intention. Your questionnaire, model, and analysis will then be longer.

Illustrative hypotheses:

  • H1: Perceived usefulness has a positive effect on intention to continue using the online learning application.
  • H2: Perceived ease of use has a positive effect on intention to continue using the online learning application.
  • H3: Content quality has a positive effect on intention to continue using the online learning application.
  • H4: Social influence has a positive effect on intention to continue using the online learning application.

This is illustrative output, not the result of a real study:

HypothesisPath coefficientt-valuep-valueIllustrative conclusion
H10.2843.9120.000Accept H1
H20.1171.8460.065Insufficient evidence to accept H2
H30.3564.7280.000Accept H3
H40.2012.6790.008Accept H4

This table shows how to connect a hypothesis to the output. It does not predict your numbers. When you run the real data, use the exact coefficient, t-value, p-value, or bootstrap confidence interval from your own file. A rejected hypothesis is still a valid result if you explain it correctly and do not alter the data to preserve the hypothesis.

How to write this in your thesis

You can adapt the following sentence to your topic: “Based on [theory name] and previous studies, this study proposes H1: [independent variable] has a [positive/negative] effect on [dependent variable] in the context of [research population and location].” In the results section, write: “The hypothesis test shows that [independent variable] has a [positive/negative] effect on [dependent variable], with a coefficient of [β = ...] and p-value [= ...]. Therefore, H[ ] is [accepted/not supported].”

Avoid writing “H1 is correct” simply because the coefficient is positive. You need to examine statistical significance, the size of the coefficient, and whether the direction matches the hypothesis. If you use PLS-SEM, you can also report the bootstrap confidence interval and the R² of the dependent variable.

Testing the Model with SPSS or SmartPLS

SPSS is suitable when your model uses linear regression, a t-test, ANOVA, or PROCESS. Import the questionnaire data into a .sav file, check for missing data, recode reverse-worded items, calculate representative scale scores, and then run the analysis. For regression, the usual menu path is Analyze > Regression > Linear. Put the dependent variable in the Dependent box and the independent variables in Independent(s), then inspect the Coefficients table.

In the Coefficients table, B or Beta indicates the direction and size of the relationship, while Sig. is the p-value as displayed by SPSS. If your hypothesis predicts a positive effect, the coefficient must have a positive sign. If the p-value does not meet the level specified in your methods section, you should not conclude that the hypothesis is supported merely because the coefficient has the expected sign.

SmartPLS is suitable when the model contains latent variables, multiple indicators, mediation or moderation relationships, and you choose PLS-SEM. In SmartPLS 4, create a project, import the .csv file, drag the indicators into constructs, draw arrows between the constructs, and run PLS-SEM Algorithm. Then use Calculate > Bootstrapping to test the significance of the path coefficient. Assessing the measurement model before the structural model follows the two-step approach described in (Anderson and Gerbing, 1988).

With SmartPLS, inspect the Outer Loadings, Construct Reliability and Validity, Discriminant Validity, Path Coefficients, R-Square, f-Square, and Collinearity Statistics tables. Common criteria include outer loading from 0.7, CR from 0.7, AVE from 0.5, HTMT below 0.85 or 0.90 depending on how close the concepts are, and VIF below 5. These thresholds are supported respectively by (Chin, 1998), (Fornell and Larcker, 1981), (Henseler et al., 2015), and (Hair et al., 2019).

When running bootstrapping, the number of subsamples is commonly set to 5,000 following the recommendation of (Hair et al., 2022). R² indicates how much of the dependent variable is explained by the variables in the model, while f² indicates the contribution of each independent variable. R² is not the proportion of hypotheses that are accepted, so do not use R² alone to conclude anything about H1, H2, or H3.

Common Errors When Building Hypotheses

The first error is writing a hypothesis without a direction. The statement “Service quality has a relationship with satisfaction” does not tell the reader whether the relationship is positive or negative. If the literature does not provide enough basis for predicting a direction, explain why and choose an appropriate test instead of adding the word “positive” without support.

The second error is deciding on hypotheses first and then searching for a theory to justify them. A more reliable sequence is to identify the problem, read the relevant models, select the concepts, define the relationships, and then write the hypotheses. A diagram with ten hypotheses but no theoretical foundation is harder to defend than a short model with clear reasoning.

The third error is mixing levels of analysis. For example, the independent variable may measure individual perceptions while the dependent variable is the revenue of an entire organization. The two variables may be related conceptually, but individual-level data are not enough to draw a conclusion about an organizational-level outcome.

The fourth error is confusing a hypothesis with a research objective. “Measure customer satisfaction” is an objective. “Satisfaction has a positive effect on repurchase intention” is a testable hypothesis.

The final error is changing the model after looking at the p-value. If H2 is not supported, record the result and discuss it. Deleting H2, reversing an arrow, or removing a variable simply because the output does not look good damages the transparency of the study.

Frequently asked questions

What is a hypothesis in scientific research?

It is a statement about a relationship between variables that can be tested with data. In a quantitative thesis, a hypothesis usually describes the direction of an effect between an independent variable, dependent variable, mediator, or moderator.

How many hypotheses should a study have?

There is no fixed number that applies to every study. The number depends on how many relationships need to be tested and how broad the model is. Prioritize relationships with a theoretical basis and measurable data instead of adding hypotheses merely to make the model look more complex.

How should you write hypotheses H1, H2, and H3?

Each hypothesis should include the influencing variable, the affected variable, and the direction of the relationship. The basic template is: “H1: [Variable X] has a positive or negative effect on [Variable Y] in the context of [research population].”

What should you do when a hypothesis is not supported?

Keep the result, report the coefficient and p-value, and discuss possible explanations such as the context, research sample, measurement approach, or characteristics of the respondents. Do not alter the data or delete the hypothesis simply to make every relationship statistically significant.

Can you build a hypothesis using only SPSS?

SPSS processes data and tests relationships, but it does not create the theoretical basis for a hypothesis. You need to build the model from the research question and literature first, then use SPSS to run regression or other appropriate tests. SmartPLS is another option when the model contains latent variables and several structural relationships.

Open your proposal and data file, then create a table with the independent variable, dependent variable, theoretical source, predicted direction, and H1 to H4 code before running the tests. If you need to run the model on your own .sav or .csv file and turn the output into thesis results, see M4 data analysis.