Double Materiality 8 min

Why 1–5 Surveys Fail Materiality Assessments (And What to Use Instead)

ExecutESG Editorial Team 01 Sep 2026
ExecutESG

Why 1–5 Surveys Fail Materiality Assessments (And What to Use Instead)

For years, sustainability professionals have relied on a familiar tool to gauge stakeholder sentiment for Double Materiality assessments: the trusty 1-to-5 Likert scale survey. It's ubiquitous, easy to deploy, and broadly understood.

You send out an email asking stakeholders to rate a list of 30 Environmental, Social, and Governance (ESG) topics from "Not Important" (1) to "Very Important" (5). A few weeks later, you aggregate the results, average the scores, and plot them on a scatter graph to create your materiality matrix.

However, as the regulatory landscape tightens with directives like the CSRD and ESRS, sustainability teams are discovering a hard truth: traditional Likert surveys are fundamentally flawed when applied to ESG.

If you are using a 1-to-5 scale to determine your material impacts and risks, you are likely suffering from deep methodological errors that render your data mathematically invalid and strategically useless. In this article, we explore why Likert scale ESG surveys fail and introduce the robust alternative: forced-choice pairwise comparison.


1. The Ubiquity of Likert Surveys in ESG Materiality

Why does everyone use Likert surveys? The answer is simple: convenience.

Platforms like SurveyMonkey and Google Forms have democratized survey creation. When sustainability teams are tasked with engaging hundreds of stakeholders—investors, employees, community members, and suppliers—a digital survey seems like the path of least resistance. It requires minimal technical knowledge to set up and provides the illusion of quantitative rigor. After all, if you have a spreadsheet full of numbers, you have data, right?

Unfortunately, taking the average of flawed subjective opinions does not yield objective truth. The illusion of rigor evaporates when you examine the behavioral psychology of how respondents actually answer ESG survey methodologies.

For a broader look at errors in this process, read our guide on Common DMA Mistakes to Avoid.


2. The Four Fatal Flaws of ESG Surveys

Relying on a 1-to-5 scale introduces four distinct behavioral and statistical flaws that corrupt materiality assessment data.

A. Central Tendency Bias

When faced with a scale of 1 to 5, human beings naturally gravitate toward the high end of the spectrum when evaluating serious topics. In an ESG context, stakeholders rarely rate any issue as a 1 or 2. Even topics that are tangentially related to a company's core operations receive a 3 ("Neutral") or a 4 ("Important").

Because stakeholders are reluctant to declare any sustainability issue as "unimportant," the entire dataset shifts toward the 4 and 5 range. This central tendency bias ensures that everything scores highly, making differentiation impossible.

B. Social Desirability Bias

Social desirability bias is the tendency of survey respondents to answer questions in a manner that will be viewed favorably by others. This is particularly toxic in ESG.

Imagine a stakeholder rating the topic of "Child Labor in the Supply Chain." Even if your company is a domestic software provider with zero risk of child labor, a respondent will feel morally uncomfortable giving it a low score. They will rate it a 5 because opposing child labor is socially desirable.

The survey stops measuring what is material to the company and starts measuring what is generally considered bad in society.

C. Survey Fatigue

A comprehensive materiality assessment often involves evaluating upwards of 30 to 40 distinct ESG topics across both impact and financial dimensions. If you ask a stakeholder to rate 40 topics on a 1-5 scale, cognitive fatigue sets in by question 15.

Respondents stop thinking critically and start "straight-lining"—clicking 4 for every single answer just to finish the survey. When 50 questions meet 30 stakeholders, the result is garbage data driven by exhaustion rather than insight.

D. Scale Subjectivity

What exactly is a "4"? To a cautious risk manager, a 4 might represent an existential threat to the business. To an optimistic marketing executive, a 4 might mean "a good idea we should look into."

Because the scale is entirely subjective, averaging a risk manager's 4 with a marketing executive's 4 is mathematically unsound. You are adding apples and oranges, yet treating the resulting average as a precise metric for resource allocation.


3. Real-World Example: The Flatline Chart

To understand the impact of these biases, look at the typical output of a Likert-based materiality survey.

Imagine you survey 100 stakeholders on 30 ESG topics. When the results come back, you calculate the average score for each topic.

  • Your lowest-scoring topic (e.g., "Noise Pollution") gets an average score of 3.8.
  • Your highest-scoring topic (e.g., "Climate Change Mitigation") gets an average score of 4.6.

We call this the "Flatline Effect." All 30 topics are clustered tightly within a 0.8 point spread.

What does this tell management? It tells them that everything is material. When every topic scores between a 3.8 and a 4.6, it is impossible to draw a line in the sand and say, "We will focus our budget on the top 5 issues." The variance is simply too small to justify strategic trade-offs.

If an auditor asks why a topic scoring 4.1 was excluded while a topic scoring 4.2 was included, you have no defensible answer. The margin of error is larger than the difference between the scores.


4. The Alternative: Forced-Choice Pairwise Comparison

If 1-to-5 surveys are broken, how do we fix them? By changing the paradigm of how we ask questions. We must abandon subjective scoring and embrace forced-choice relative prioritization.

The most mathematically robust way to achieve this is through Pairwise Comparison.

How It Works

Instead of presenting a stakeholder with a list of 30 topics and asking them to rate each one in isolation, pairwise comparison presents two topics side-by-side.

The stakeholder is asked a simple binary question: "Which of these two issues poses a greater financial risk to our company over the next 5 years: Supply Chain Disruption or Data Privacy Breaches?"

The stakeholder must choose one. They cannot say both are highly important; they are forced to make a trade-off.

Forcing Genuine Trade-Offs

By forcing a choice, you completely eliminate social desirability bias and central tendency bias. A stakeholder is no longer declaring that an issue is "unimportant"; they are simply stating that in the context of this specific business, Topic A is a higher priority than Topic B.

This mirrors how real-world business decisions are made. A CEO does not evaluate investments in a vacuum; they evaluate them relative to competing priorities.

The Mathematical Advantage: 40:1 Variance vs 1.2:1 Variance

When you aggregate pairwise comparisons using a mathematical framework like the MLE Consensus Model for ESG Materiality, the results are staggering.

Because stakeholders are forced to differentiate, the data spreads out. Instead of the tight 1.2:1 variance seen in Likert surveys (where the top topic is 4.6 and the bottom is 3.8), a pairwise consensus model generates massive variance.

It is common to see a 40:1 variance, where the mathematical consensus proves that the top priority topic is 40 times more significant to the stakeholder group than the bottom priority topic. This provides executive teams with crystal clear, undeniable strategic direction.

For more on the mechanics of this process, read our article on Collective Decision Making via Pairwise Prioritization.


5. How to Transition: Practical Steps for Sustainability Teams

Transitioning away from flawed surveys to a mathematically robust methodology doesn't have to be painful. Here is how your sustainability team can make the shift:

  1. Acknowledge the Flaws: Start by showing management the "flatline" results of past surveys. Explain that central tendency bias is preventing the company from making focused, strategic decisions.
  2. Define the Context: Ensure that your stakeholders understand how they are comparing topics. Are they voting based on outward impact severity, or inward financial risk? Context is critical for valid pairwise comparisons.
  3. Adopt a Mathematical Engine: Do not try to calculate pairwise matrices manually in Excel. Adopt a platform that automates the presentation of pairs and the calculation of the final consensus scores.
  4. Communicate the Value: Let stakeholders know that their time is valuable. Explain that a short series of A/B choices provides far more strategic value than clicking "4" fifty times on a Likert scale.

(Note: For smaller enterprises looking to comply with VS (VSME) standards, moving to a forced-choice model simplifies the reporting burden by clearly isolating the few metrics that actually matter).


6. Upgrade Your Methodology with ExecutESG

The regulatory requirements of the CSRD mean that "good enough" methodologies are no longer acceptable. You need an auditable, mathematically rigorous approach to stakeholder engagement.

ExecutESG's AHP Pairwise Consensus Engine replaces flawed surveys with an intuitive, mobile-friendly interface. Stakeholders make simple A/B choices, and our proprietary algorithms calculate the true mathematical consensus, generating defensible, high-variance data for your double materiality matrix.

Stop relying on garbage data. Start making strategic decisions based on mathematical truth.


Free 2-Minute Diagnostic

Check Your VS (VSME) & CSRD Readiness Score

Answer 6 quick questions under the EU Voluntary Standard VS (VSME) to discover your compliance gap score, estimated time savings, and generate your free starter report.

Take the Free VS (VSME) Quiz →

🍪 Your Privacy Options

We use strictly necessary cookies to keep you signed in and protect your session. With your explicit consent, we also use analytics cookies (Google Analytics GA4) to improve our service. You can choose to accept all cookies or only allow essential ones. Read our Privacy Policy.