Customer Satisfaction Survey: How to Conduct and Process the Data

Surveys··14 min read

What a customer satisfaction survey means in quantitative research

A customer satisfaction survey is a structured way to collect data about customers' perceptions after they use a product, receive a service, or interact with a business. You use a standardized questionnaire that usually includes items measuring service quality, perceived value, responsiveness, experience, and intention to return or recommend.

In a quantitative thesis, the survey produces numerical data. Each customer selects a response level for each item, and you then code those choices as numbers for entry into SPSS or SmartPLS. For example, SAT1 might measure the statement “I am satisfied with this service,” while SAT2 measures “The service meets my expectations.”

You should distinguish the concept being measured from the questionnaire used to measure it. The concept is customer satisfaction, while the questionnaire is the instrument containing screening questions, descriptive information, and measurement items. If you want to understand the structure of this instrument, you can also read what is a questionnaire. Defining the concept before writing the items helps you avoid having each item measure a different subject.

A good quantitative questionnaire must connect three parts: the research objective, the model, and the data you need to analyze. If your hypothesis tests the effect of service quality on satisfaction, the questionnaire needs items measuring service quality, items measuring satisfaction, and control variables if your study uses them.

How to determine the customer satisfaction survey sample size

Sample size depends on your model, the number of items, the analysis method, and your access to customers. Do not choose a number simply because another study used it. State the basis for your calculation in Chapter 3, then collect data from the population you have described.

If the customer population is finite and you know its size, you can use the Yamane formula:

n = N / (1 + N.e²)

Here, N is the population size, n is the minimum sample size, and e is the accepted error. You should cite this formula as (Yamane, 1967). If you do not have a complete customer list, state clearly that the accessible population is difficult to determine instead of entering an unsupported value for N.

For a questionnaire with many items, a commonly used rule is 5 to 10 observations for each item (Hair et al., 2010). For example, a questionnaire with 25 items would produce a planning range of 125 to 250 valid responses under this rule. This supports your planning, but it does not guarantee that the data will pass every test.

If you run a regression with m independent variables, another guideline sets the sample size at 50 + 8m or above (Tabachnick and Fidell, 2013). When the two approaches produce different results, use the higher figure and allow for invalid responses. For example, if you need 200 valid questionnaires and expect 15% to be excluded, you need about 236 responses, because 200 / 0,85 is approximately 236.

Sample-size situationCalculation or reference levelSource to report in the thesis
Finite population, N knownn = N / (1 + N.e²)(Yamane, 1967)
Questionnaire with items5 to 10 observations for each item(Hair et al., 2010)
Regression with m independent variablesn ≥ 50 + 8m(Tabachnick and Fidell, 2013)
Factor analysisConsider sample size and data suitability(Comrey and Lee, 1992)

You should also anticipate exclusions caused by extensive missing responses, straight-lining, or failure to pass a screening question. If you collect only the minimum number of questionnaires, a few invalid cases can leave your final sample below the planned size.

Designing a questionnaire for this objective

Start with a specification table. Each construct in the model needs a definition, a scale source, a variable code, and an expected number of items. You can consult survey form template to understand the layout, but adapt the content to the product, service, and customer group in your study.

The introduction should briefly state the academic purpose, target respondents, completion time, and confidentiality principle. Do not request identifying information unless it serves the analysis. A clear introduction tells respondents which experience they should evaluate, such as their most recent transaction or their use of the service during the previous three months.

Place the screening section before the measurement items. For example, you can ask whether the respondent used the service during the study period, whether they completed a transaction at a particular branch or platform, and whether they belong to the target customer group. Respondents who do not meet the criteria should exit the questionnaire or be directed to the appropriate section.

Descriptive information may include age, gender, occupation, usage frequency, length of service use, and transaction channel. Include only variables that serve the research objective. The more unnecessary items you add, the higher the dropout rate and the longer data cleaning will take.

Customer satisfaction items commonly use a multi-point Likert scale. The Likert scale was introduced in the original work of (Likert, 1932). For a student questionnaire, a 5-point scale is often easy to explain, ranging from 1 for strongly disagree to 5 for strongly agree. You can read more about what a Likert scale is before deciding how many points to use.

Each item should express one main idea and identify what the respondent is evaluating. “The consulting staff are enthusiastic and process requests quickly” combines two subjects. A respondent may agree with one part but not the other. Separate these into two items if both subjects matter to the study.

Avoid leading wording such as “Do you agree that our service is very good?” This wording places a positive judgment inside the item. Use a neutral statement instead, such as “The service meets my expectations,” and keep the coding direction consistent wherever possible.

Before distribution, test the questionnaire with a small group that fits the research population. The purpose is to identify unclear wording, missing response options, overlapping items, or excessive completion time. Do not automatically merge pilot data with the main dataset if you changed the questionnaire afterward.

Implementing data collection

Define who will be invited, where they will be reached, and the collection period. If your study examines a store's customers, you can ask customers after a transaction or send a link to customers who agreed to receive the survey. If you study an application's customers, state the condition that they must have used the application during a specific period.

Avoid distributing a public link without controlling the target population when your research requires actual customers. A response from someone who has never used the product can distort the satisfaction variable. Record the start date, end date, collection channel, and number of responses from each channel so that you can report the process transparently.

DoThesis has fillform.info, our form tool for creating and collecting Vietnamese questionnaire responses. You can use it when you need a centralized form for your study. Google Forms is a familiar alternative, especially when your research team already has an account and a sharing process.

Do not describe the collection period as “collecting enough responses in a short time” if you have no specific plan. Set an expected period, monitor the number of valid responses each day, and check the sample composition. If 90% of responses come from one age group while your model needs several groups, adjust the access channel instead of continuing to distribute the same link.

Do not give gifts or ask respondents to choose answers that benefit the business. If you offer a gift, describe the conditions clearly and separate the contact information for receiving it from the analysis data where possible. This reduces the risk of directly linking identifying information to responses.

Cleaning data before importing it into SPSS

First, export the data as .csv or .xlsx, then create a read-only original file and a working file. Do not edit the only copy. Check variable names, value coding, and column order before opening the file in SPSS. Each row should represent one respondent, and each column should represent one variable.

Remove technical columns that are not used in the analysis, such as session IDs, system timestamps, or addresses that do not serve the research objective. Keep a response ID if you need to trace a case, but do not include it in regression, EFA, or Cronbach's Alpha. Identifying questions should not be treated as measurement variables either.

Check incomplete responses. A questionnaire that leaves most measurement items blank should generally not be treated as a valid observation. For a few isolated blank cells, define the treatment rule in advance and record the number of affected cases. Do not fill every blank with the mean simply to preserve the number of rows.

Check for straight-lining. If a respondent selects the same level throughout the questionnaire, this is a sign that requires review, but it is not enough to exclude the case automatically. Also check completion time, reverse-coded items if any, and the screening question. A short questionnaire can be completed quickly and still be valid.

Check out-of-range values. For a 1 to 5 scale, a value of 0, 6, or text outside the defined codes should be corrected in the original data if you can confirm that it is an entry error. If you cannot determine the correct value, code it as missing and record the rule. Do not replace extreme values with the mean without a reason.

If an item is reverse-coded, reverse the code before running reliability tests or creating a composite variable. For a 1 to 5 scale, the usual transformation is new value = 6 - old value, but apply it only when the scale design requires it. A mistake here can lower correlations between items and cause you to remove the wrong item.

After cleaning, save a log containing the response ID, exclusion reason, processing time, and person responsible. This process helps you answer the supervisor's question about why the initial number of responses differs from the number of observations in SPSS. You can also read Google survey if your data came through Google Forms and you need to check how to export the dataset.

How to write this in your thesis

In Chapter 3, describe the survey population, sampling method, collection period and location, measurement instrument, scale used, and criteria for valid data. Your wording must reflect what you actually did. Do not report probability sampling if you really distributed a convenience link through your personal network.

You can present the sample size using this structure: the questionnaire contained [number of items] items, and the study applied the rule of [5 to 10] observations for each item according to (Hair et al., 2010), producing a minimum sample size of [minimum sample size]. After allowing for [rate]% invalid responses, the study distributed [number distributed] questionnaires and received [number collected] responses.

You should separate the number distributed, the number collected, the number excluded, and the number of valid responses. If you used several collection channels, create a table by channel so readers can see where the data came from. Do not call every form visit the sample size, because only responses meeting the criteria enter the analysis.

You can write the description as follows: “The study used a structured questionnaire consisting of [number of sections] sections. The respondents were [customer description], who had [product or service use condition] during [time period]. After excluding [number of questionnaires] invalid responses, [number of valid questionnaires] observations were used for the subsequent analyses.” Replace every bracketed section with the actual values from your file.

In the appendix, include the questionnaire and variable codes. For example, SAT1 through SAT4 can represent the satisfaction construct, while SER1 through SER5 can represent service quality. Consistent coding helps you compare the questionnaire with the Reliability Statistics, KMO and Bartlett's Test, Rotated Component Matrix, and Coefficients tables in the SPSS output.

Common mistakes

The first mistake is using one general questionnaire for every industry without defining the specific customer experience. Bank customers, e-commerce customers, and clinic customers encounter different touchpoints. The items must describe the service you are studying, or the satisfaction score will be difficult to interpret.

The second mistake is asking people who do not belong to the target population. A large number of responses cannot compensate for an incorrectly defined sample. Set the screening criteria before the main items and check the form's page-logic rules.

The third mistake is changing the items after collecting part of the data and then merging all responses. If you change the wording, add an item, or alter the scale, mark the change date and consider separating the two data waves. Record the decision so you can explain it later.

The fourth mistake is removing data simply because one item has a low value or because you want a better result. Exclusion must follow criteria stated in advance, such as failure to pass screening, excessive missing data, or evidence of invalid responding. Do not delete an observation merely because it lowers the mean satisfaction score.

The final mistake is entering variable names with diacritics, spaces, or excessive length. SPSS and SmartPLS apply different rules to variable names. Use short codes without diacritics, such as SAT1, SAT2, and AGE, and keep a codebook mapping each variable to its Vietnamese content.

Frequently asked questions

How many questions should a customer satisfaction survey include?

There is no single number for every topic. The number depends on the constructs in your model and the number of items in each scale. You need enough items to represent each construct, but an overly long questionnaire increases dropout and straight-lining.

Use the scale source, research objective, and pilot-test results as your basis. Every item should have a reason to exist, be linked to a construct, and have a code for analysis.

What sample size is sufficient for a customer survey?

You need to state the basis instead of reporting only one number. You can use 5 to 10 observations for each item (Hair et al., 2010), or the 50 + 8m formula for a regression with m independent variables (Tabachnick and Fidell, 2013).

The final sample size is the number of valid responses after cleaning. If you expect some questionnaires to be excluded, collect additional responses and report the numbers distributed, collected, excluded, and entered into SPSS separately.

Should you use a 5-point Likert scale to measure satisfaction?

A 5-point scale is easy to understand and implement in many questionnaires. State the meaning of each level, keep the order from low to high, and use the same coding direction across variables.

If an item is reverse-coded, mark it in the codebook and process it before running the reliability test. Do not add a “don't know” option to the Likert scale unless you have a separate coding plan for that option.

What should you do when many survey responses are identical?

First, check completion time, missing-response rates, attention-check items, and screening conditions. Identical responses do not automatically prove that a respondent answered incorrectly, especially when the items are similar or the questionnaire is short.

Set the exclusion criteria before processing and apply them consistently. Save the list of excluded responses and the reasons, then report the final figures in Chapter 3.

Can you import Google Forms data directly into SPSS?

Export the data as .xlsx or .csv, clean the column names, and check the coding before opening it in SPSS. Do not use timestamps, email addresses, or response IDs as analysis variables unless they belong to the model.

Before running the analysis, compare several rows in the exported file with the original form. A common problem is that response options are exported as text while measurement variables need numerical codes from 1 to 5.

Should you conduct the survey in person or online?

The collection method depends on where you can reach customers and on the requirements of your study. Online collection makes it easier to monitor response counts and export data, while in-person collection can reach customers at a transaction point but requires more control over data entry.

Whichever method you choose, describe the population, period, channel, and validity criteria. The collection format does not replace sample design and data cleaning.

Open your current data file, create the variable codebook, check the number of valid responses, and record the exclusion rules before running SPSS. If you need help running the analysis on your own .sav or .csv file, you can use the M4 data analysis module.