What Is DEA? Model, Input and Output Variables, and Testing

Research models··15 min read

What is a DEA model?

DEA stands for Data Envelopment Analysis. It is a quantitative method for evaluating the relative efficiency of multiple decision-making units, or DMUs, when each unit uses one or more inputs to produce one or more outputs.

For example, you can treat each bank as a DMU, with the number of employees, operating costs, and capital as inputs, and revenue or profit as outputs. In an education study, each university can be a DMU, with the number of lecturers, budget, and campus area as inputs, and the number of graduates or research results as outputs.

DEA does not simply ask which unit has the highest result. The model builds an efficiency frontier from the best-performing DMUs in the dataset, then compares the remaining units with that frontier. A DEA score therefore usually represents relative efficiency within the sample, not an absolute ratio that applies to every industry.

When you search for “what is DEA”, you will also find DEA analysis in Stata. This is a suitable option when your data are panel data or cross-sectional data covering multiple DMUs. DEA differs from what is ARDL, what is ARIMA, and other time-series forecasting models because DEA focuses on measuring relative efficiency across units.

Each DMU receives an efficiency score in DEA. Under the common interpretation, a score of 1 means that the DMU lies on the sample's efficiency frontier. A score below 1 indicates that the DMU may reduce inputs or increase outputs according to the model's orientation. A score of 1 does not mean that the unit operates optimally under every condition, and it does not automatically establish a causal relationship.

Components of a DEA model

Identify the components of your DEA model before opening Stata. If you select the wrong DMUs or include too many variables, most units may receive an efficiency score of 1. The model then has little ability to distinguish performance across units.

ComponentMeaning when building a DEA modelIllustrative example
DMUThe unit whose efficiency is evaluatedBank, hospital, school, province
InputResources usedLabor, capital, costs, number of beds
OutputResults producedRevenue, profit, treated patients
OrientationThe model's optimization directionReduce inputs or increase outputs
Returns to scaleAssumption about operating scaleCRS or VRS
Efficiency scoreRelative efficiency scoreUsually between 0 and 1 for input-oriented DEA

DMUs must be comparable

DMUs should perform similar functions and use resources that have comparable meanings. Comparing commercial banks with insurance companies in the same model can create problems involving technology, processes, and operating objectives. State the criteria for including a unit in your sample instead of selecting only the units whose data are easiest to download.

The number of DMUs must also be adequate for the number of inputs and outputs. When a model contains too few DMUs and too many variables, the DEA frontier becomes broad and many units may be classified as efficient. This is why you should consider combining variables with closely related content or retain only variables that are genuinely connected to your research question.

Inputs and outputs must reflect the efficiency mechanism

An input is a resource consumed or used by the DMU. An output is a result produced by the DMU. Revenue is often treated as an output, while labor costs are usually treated as an input. The classification still depends on the theory of the industry and the objective of your study.

Do not select a variable simply because it appears in a financial report. Each variable needs an explanation: why it is a resource, why it is a result, and why it can be compared across DMUs. If a variable is used as both an input and an output in the same model, explain the logic clearly and check the sensitivity of the results.

Orientation and returns to scale

Input-oriented DEA answers this question: how much can a DMU reduce its inputs while maintaining its current output level? Output-oriented DEA asks how much the DMU can increase its outputs with its current resources. If managers have more control over costs than revenue, an input-oriented model is often easier to explain.

CRS, or Constant Returns to Scale, assumes that a proportional increase in inputs produces a corresponding proportional increase in outputs. VRS, or Variable Returns to Scale, allows efficiency to change with operating scale. For banks or hospitals with very different sizes, consider VRS so that disadvantages caused by scale are not treated as poor management efficiency.

The basic model and common extensions

The basic DEA model is commonly presented in two forms: CCR and BCC. CCR is associated with the CRS assumption, while BCC is associated with the VRS assumption. When you run the study, do not select a model simply because the software uses one option as its default. Start with the research question and the operating characteristics of the DMUs.

The CCR model gives an overall efficiency score under constant returns to scale. The BCC model separates pure technical efficiency from scale effects. By comparing the two results, a study can examine scale efficiency using the CRS and VRS scores, but the formula and reporting approach must remain consistent with the software you use.

Another extension concerns the choice of orientation. If your study focuses on cost reduction, you can use an input-oriented model. If the objective is to assess the ability to expand output using existing resources, an output-oriented model may be more appropriate.

DEA can also be extended over time with the Malmquist index to examine productivity changes across periods. This is different from ordinary cross-sectional DEA. You need the same variable system and the same measurement logic across years. If you have data for only one year, do not describe the result as productivity growth over time.

You can also review FEM and REM for panel data when your dataset contains multiple DMUs across multiple years. FEM and REM estimate relationships between variables, while DEA builds an efficiency frontier. The two groups of methods can address different questions, but you should not substitute one for the other simply because both can use panel data.

Common measures for each concept

DEA does not use Likert scales in the way that Cronbach's Alpha, EFA, or CFA does. Input and output variables are usually observed quantitative variables taken from financial reports, industry databases, or administrative statistics. The table below is therefore a map of commonly used variables to help you start designing the model, not a validated scale for every topic.

Concept to measureExample input or output variableSuitable data sourceNote when using it
Labor scaleNumber of employees, hours workedAnnual report, human resources reportUse consistent units across DMUs
Capital or assetsTotal assets, owners' equityFinancial statementsConsider a log transformation if the gap is very large
Operating costsOperating costs, personnel costsFinancial statementsSpecify exactly which costs are included
Service outputNumber of transactions, customers, or patientsOperating reportAvoid using an indicator that duplicates an input
Financial resultRevenue, profit, incomeFinancial statementsCheck negative values and outliers
Output qualityGraduation rate, service level, completion rateIndustry reportExplain how units are compared

If your data concern a latent concept such as satisfaction or perceived quality, DEA is not a direct scale-measurement tool. You can use a questionnaire to create a composite index, but you need to describe how the index is calculated and why it is included as an input or output. For a PLS-SEM model, you can also review what is GARCH to distinguish DEA from a model for handling volatility in financial data.

Check the measurement units before running the analysis. If one variable is measured in million currency units and another in billion currency units, the results may be misunderstood during interpretation. DEA is relatively insensitive to some proportional transformations, but consistent units are still necessary for the descriptive table and for reproducing the analysis.

Applying DEA to your study

Suppose your study evaluates the operating efficiency of 20 bank branches in the same year. Each branch is a DMU. You select three inputs, namely the number of employees, operating costs, and total assets, together with two outputs, namely service revenue and profit before tax.

The conceptual model can be written as follows: branch resources consist of labor, costs, and assets, and these resources are transformed into financial results. DEA builds an efficiency frontier from the branches with the best input and output combinations in the sample, then calculates a relative score for the remaining branches.

You could formulate the following management-oriented research hypotheses:

  • H1: The technical efficiency of bank branches differs across DMUs.
  • H2: Efficiency under the VRS assumption is greater than or equal to efficiency under the CRS assumption for each DMU.
  • H3: Scale efficiency differs between groups of branches with different sizes.
  • H4: Operating-environment factors are associated with DEA efficiency scores in a subsequent analysis.

H1 to H3 fit the description and comparison of DEA scores. H4 requires more caution because, if you use the DEA score as a dependent variable in regression, the score is usually bounded and may contain many values equal to 1. You need to choose an appropriate regression model and treatment method instead of running linear regression and immediately stating a conclusion.

If your study evaluates provinces, a province or city can serve as a DMU. Inputs may include investment capital, labor, and budget expenditure. Outputs may include GRDP, the number of active businesses, or a social indicator. Check that the provinces have the same measurement scope and period, especially when the figures come from multiple sources.

How to write this in your thesis

You can adapt the following paragraph: “The study uses DEA to evaluate the relative efficiency of [number] DMUs. The inputs include [list of inputs], and the outputs include [list of outputs]. The [CRS or VRS] model and [input or output] orientation are selected because [reason connected to the research question]. The DEA results are used for [descriptive, comparative, or subsequent analytical purpose].”

In the methods chapter, present a table defining the variables, measurement units, data sources, and reasons for classifying each variable as an input or output. In the results chapter, report the number of efficient DMUs, the distribution of efficiency scores, the reference DMUs, and unusual results. Do not present only a score table and call it a complete DEA analysis.

How to test the model using SPSS or SmartPLS

DEA is usually run in Stata, R, Python, or specialized software rather than SPSS and SmartPLS. SPSS can support data cleaning, descriptive statistics, and correlation checks. SmartPLS is suitable for measurement and structural models in PLS-SEM, not as the primary tool for building a DEA frontier.

Preparing the data file

Each row should represent one DMU in cross-sectional data, or one DMU at one point in time when you have panel data. Each column should have a short code without diacritics and without special characters. Keep separate columns for the DMU identifier, year, inputs, and outputs.

Check missing values, negative values, zero values, and outliers. Variables such as profit may be negative, but the treatment of negative values in DEA depends on the model and software. Do not automatically add a constant to every negative value without explaining the effect of that transformation.

Running DEA analysis in Stata

In Stata, commands and user-written packages may differ according to the version and installation method. Check the help file for the command you are using before the official run. The practical process usually includes importing the file, describing the data, identifying input and output variables, choosing CRS or VRS, choosing input-oriented or output-oriented analysis, and saving the efficiency scores.

After running the model, save the do-file, Stata version, command name, model options, and run date. At a minimum, the results table should contain the DMU code, CRS efficiency score, VRS efficiency score, and explanatory variables if you conduct a subsequent regression step.

The table below is illustrative output, not the result of a real study:

DMUCRS efficiencyVRS efficiencyScale efficiencyIllustrative interpretation
CN011.0001.0001.000Lies on the frontier under both assumptions
CN020.8200.9100.901May have room to improve inputs
CN030.7600.8000.950Technical and scale efficiency both require review
CN041.0001.0001.000Relatively efficient within the sample

Do not conclude that CN01 is the best branch in the market simply because its score is 1. The more accurate conclusion is that CN01 lies on the efficiency frontier created by the sample DMUs and selected variables. For CN02, a score of 0.820 in an input-oriented model may be interpreted as potential for reducing inputs while holding outputs constant, but the specific reduction depends on the slack results and model configuration.

Comparing results with SPSS and SmartPLS

SPSS is useful when you need to inspect Mean, Standard Deviation, Minimum, and Maximum, identify unusual data, or create a descriptive table before DEA. SmartPLS is appropriate only when the study also includes a PLS-SEM model, such as testing the effect of management capability on DEA efficiency. In that case, clearly present DEA and PLS-SEM as two separate analytical stages.

CFI, TLI, RMSEA, and SRMR are fit indices for evaluating CB-SEM, not substitute measures for a DEA score. You should also not use Cronbach's Alpha to assess the quality of a single financial variable. A scale and an efficiency frontier answer two different methodological questions.

Common errors when using this model

The first error is calling every efficiency indicator DEA. If you simply divide revenue by cost, that is a simple efficiency ratio, not DEA. DEA requires multiple DMUs, a set of inputs and outputs, and a specified orientation and returns-to-scale assumption.

The second error is including too many variables in a small sample. As the number of inputs and outputs increases, the model may classify too many DMUs as efficient. Build a base model and then check sensitivity using several reasonable configurations instead of running dozens of models and selecting only the most attractive table.

The third error is combining non-comparable DMUs. A central hospital and a small clinic should not be compared without an argument about their functions and scale. If the DMUs cannot be made comparable, narrow the sample scope.

The fourth error is ignoring outliers. A DMU with extremely high revenue may create the efficiency frontier and pull down the scores of other units. Check descriptive statistics, charts, and data sources before deciding whether to retain or remove an observation. Record every data exclusion.

The final error is writing a causal conclusion from a DEA score. DEA reports relative efficiency under the selected configuration. If you want to explain factors associated with efficiency, you need an additional analytical design, a clear theoretical basis, and checks for the statistical issues created by a bounded dependent variable.

Read also: what is ARDL, what is ARIMA, what is contingent valuation method CVM, what are FEM and REM for panel data, what is GARCH.

Frequently asked questions

What is DEA and what is it used for?

DEA is a method for evaluating the relative efficiency of multiple DMUs with multiple inputs and outputs. It is suitable when you want to compare banks, hospitals, schools, branches, businesses, or provinces with similar functions.

A DEA score of 1 only indicates that the DMU lies on the sample's efficiency frontier. You still need to review the inputs, outputs, model assumptions, and sensitivity before drawing a management conclusion.

Is DEA analysis in Stata difficult?

The main difficulty is usually data design, not typing the command. You need to identify the correct DMUs, inputs, outputs, orientation, and CRS or VRS assumption, then check missing data, negative values, and outliers.

Stata can perform DEA when you use a command or user-written package compatible with your installed version. Save the do-file and the command documentation so that another person can reproduce the result.

What DEA score is acceptable?

DEA does not have a universal acceptable threshold like Cronbach's Alpha or a p-value. A score of 1 generally indicates relative efficiency on the sample frontier, while a score below 1 indicates a distance from the frontier under the selected orientation.

Report the score distribution, the number of efficient DMUs, and the interpretation used by the model. Do not apply an arbitrary threshold and divide DMUs into good and poor performers without a methodological basis.

Can SPSS be used to run DEA?

SPSS can support data import, descriptive statistics, and preliminary checks, but it is not the primary choice for running DEA. Stata, R, Python, or specialized software is usually more suitable for building the efficiency frontier.

SmartPLS also does not replace DEA. Use SmartPLS only when the study includes a latent-variable or structural model, and clearly separate the results from the two methods.

Can DEA be used with panel data?

DEA can be applied to multiple DMUs across multiple years, but you must distinguish DEA run separately by year from an analysis of productivity change over time. The data need consistent measurement across periods and stable DMU codes.

If your main objective is to estimate the effect of one variable on another in panel data, consider separate panel-data regression models. DEA can serve as an efficiency-measurement step, but it does not automatically replace FEM, REM, or dynamic models.

Open your data file, create a table containing the DMU code, inputs, outputs, and year, then record why you selected the orientation and CRS or VRS assumption before the first run. When you need to check the model structure and analysis process using your own .sav or .csv file, you can use DoThesis M4 data analysis.