Methodology

Durability Score (Ds)

A composite score at the brand × category × market level. Five indicators, public weighting, validated against Stiftung Warentest and Consumer Reports.

Ds v3 (evidence-based) · published 2026-08

Core principles

Independence

Ratings are not based on payments by manufacturers or retailers. No financial relationship may influence a result.

Transparency

Data sources, weighting, and calculation method are published. Anyone may challenge the methodology.

Reproducibility

Each release carries a methodology version number. Code and data are public on GitHub and Zenodo.

Currency

Ratings are reviewed annually; a weighting change causing a shift of 2+ positions triggers a 30-day public consultation.

Five Durability Score indicators

The final Ds is a weighted average of five indicators. Each rating carries an evidence grade A–D.

D0.35Tested durabilityF0.25Field failureX0.15Warranty returnsC0.15General qualityO0.10Repairability→ also published as QIweights_for(present)a missing component is dropped, never zeroed — the rest rescale to sum to 1.0Σ wᵢ · valueᵢ = ds_totalPublish-tier gateneeds ≥1 durability-specific component — D, F, X or OTier A≥2 componentsTier Blimited evidenceTier Cnot published as DsEE — Consumer Indexnever enters WEIGHTSSecondhand Deal Findermarket data collected, no indicator yetDsDurability ScoreTiers A–B · 742QIQI — Quality Index= component C alone3,352
Five signals are renormalized over whichever are actually present for a brand×category — a missing one is dropped, not treated as zero — then summed and passed through a tier gate before a Durability Score is published at all. Component C is scored into Ds and published standalone as the Quality Index. Component E never joins the pipeline: crowd ratings correlate with Ds at ρ = 0.28, but the link is an attention-and-expectation effect, not durability evidence — folding it in would import the reputation bias the score exists to avoid. Secondhand Deal Finder market data is collected separately for the same reason — as raw data so far, not yet a finished indicator.
D

Tested durability

35 %

Laboratory sub-ratings for build quality, wear and robustness. The only signal that is both measured in a laboratory and directly about how long something lasts, which is why it carries the most weight.

Sources: Stiftung Warentest, dTest, Which?, OCU, Altroconsumo

F

Field failure & lifespan

25 %

Share of owners reporting a fault, time to first fault, and reliability rating. Three views of the same survey, so they form one component rather than three.

Sources: Which? Reliability Survey (owner panel, 7-year window)

X

Warranty returns

15 %

Return rate across the full two-year statutory warranty. An independent instrument and population from F — different country, different observation window.

Sources: Alza.cz (two-year statutory warranty)

C

General quality (supporting)

15 %

Laboratory ratings across the remaining dimensions, weighted toward longevity. Never sufficient on its own: without at least one signal directly about durability, no score is published.

Sources: Stiftung Warentest, dTest, Which?, OCU, Altroconsumo

O

Repairability

10 %

The French Indice de Réparabilité (note_c1–c5). Lowest weight because it is declared by the manufacturer rather than measured, and its values sit near the top of the scale (mean 89.8), so it separates brands weakly.

Sources: Indice de Réparabilité (FR), data.gouv.fr

Cross-lab calibration: additive fixed-effects model, validated by leave-one-brand-out. Reduces cross-lab disagreement by 29% (D) and 40% (C).

Why every number carries a range

Two accredited laboratories testing the same brand in the same category still differ by roughly 9.8 points after their scales are reconciled — measured leave-one-brand-out, not assumed. About 90 % of brands are measured by a single independent source, where no disagreement can be observed at all because there is nothing to compare against. Publishing the bare figure would make such a brand look exactly as certain as one confirmed by five laboratories. The interval therefore carries that residual: a single-source brand reads about ±19 points, a corroborated one about ±10.

A

Confirmed by several sources

At least two independent components directly about durability. Published as a Durability Score.

B

Limited evidence

A single durability component. Published, but always with a visible caveat.

C

Not published

General quality only, or a single laboratory with no way to calibrate. The number exists; we do not present it as durability.

General quality (C) is never enough on its own. The database holds 1,373 brand × category combinations with laboratory quality ratings but no signal about durability at all. If C could stand alone, those would be published as durability under a quality number — precisely the error this version of the methodology removes.

Open the Brand Durability Monitor →

Consumer Index (E) — published separately

Customer ratings from e-shops are held separately in the database (signal_type = crowd_rating) and do not enter Ds — an earlier version of this methodology gave them 30 %, but analysis showed the correlation with actual longevity is weak and easily confused with expert verdicts in the data. The signal itself is real and measured, though, so since August 2026 it has been published on its own as the Consumer Index at /en/evaluation and in the Brand Durability Monitor — it never changes the value of Ds.

Agreement between independent signals — a supplementary view, not a sixth component

If a product sells well, that is weak — but real — evidence of its quality, even when price or brand reputation confound it. On /en/evaluation we therefore show, below the main table, whether Ds (lab evidence) and the Consumer Index (market signal) agree on a given brand. This is not a new score and never changes the value of Ds or its weights — it is a purely descriptive overview.

Why it is not a sixth weighted component: Ds and the Consumer Index correlate at only rho ≈ 0.28 (n = 162, p = 0.0003) — well below the agreement the methodology requires before blending any signals into one weighted component (Which?'s three field-failure signals inside component F agree at 0.75–0.92). Weighting in a signal that correlates this weakly would produce a number that looks more precise than the evidence supports — and would quietly re-admit exactly the reputation effect Ds deliberately excludes E from its weighting to avoid.

Finding from real data: agreement was tested with a permutation test (signal_agreement_validation.py). At the same 95% confidence interval shown on the Ds and E numbers themselves, only 15 of 170 cells with both signals present were conclusive — too few to test agreement at all (p = 0.60, indistinguishable from noise). At a wider ±1 standard deviation band (~68% confidence), 41 cells are conclusive and agreement runs 82.9% versus 51.0% expected by chance — permutation test p = 0.0016. The table on /en/evaluation therefore uses this wider band and discloses it visibly — it is not the same confidence level as the headline Ds and Consumer Index numbers.

Price level — descriptive metadata, not a Ds component

For eleven categories we publish where a brand sits on price: budget, mid or premium. It exists to answer the one question the monitor previously could not — which cheap brands actually last. Price never enters Ds or any other number here; folding it in would be a new methodology version, not an addition.

A level describes the brand, not the product. It is a median across that brand’s products in that category. Bosch sells both a €400 and a €1,200 washing machine — “premium” says where the brand typically sits, not what any one model costs.

Prices arrive in seven currencies and our exchange-rate table is empty, so we do not convert them. Converting would be the wrong move even with rates: an exchange rate does not make a Czech price comparable to a Swedish one, because the price level of the market itself differs through tax, purchasing power and competition. “Premium or budget” is inherently a within-market question, so each product is ranked against others in the same (category, currency, retailer) group and a brand’s position is the mean of those dimensionless ranks. Tertile boundaries are cut across every priced brand in the category, not only those with a Ds, so they reflect the real price distribution.

Where we withhold it, and why: a category appears only if at least ten of its brands have both a price and a Durability Score. Tertiles over four brands are not a segmentation. Eleven of roughly forty categories meet that bar. The euro figure shown comes exclusively from prices recorded by the testing laboratories — we found that three e-commerce scrapers write incorrect units, so retail prices are used for ranking (a percentile is scale-invariant) but never as an absolute amount. That is why many brands show a level with no price beside it.

Planned indicators

These signals are designed but do not yet feed any published number. They are listed separately so they cannot be mistaken for part of the score.

Quality stability over time

Slope of a brand's monthly ratings. Needs a longitudinal series that is still being collected.

Observed in-use lifespan

Presence on the secondhand market years after launch. Data is being gathered through the Secondhand Deal Finder.

Price-tier separation

A durable budget brand and a durable premium brand currently sit on the same axis. Separating price tiers is in preparation.

Derived metrics

QCR

Quality Coverage Ratio

Share of a brand's portfolio that meets the minimum quality threshold. Distinguishes premium brands from those with just a few good models.

BCI

Brand Consistency Index

Intra-class correlation of Ds across markets. High BCI = the brand pursues a consistent quality strategy. Low BCI = markedly different quality by market.

ORS

Obsolescence Risk Score

Derived metric 0–100 computed by the ODA algorithm (PELT change-point detection) from longitudinal review trajectories. Signals whether managed obsolescence is occurring.

Evidence grades A–D

Every rating carries a visible evidence grade so readers cannot mistake a weak signal for a verdict. Ratings of C or D are published with a warning and may not be used in headlines or comparison tables.

AStrong

Multiple independent laboratory tests + repairability data + warranty/complaint statistics + consumer survey, all consistent.

BModerate

Multiple independent sources in agreement, but limited scope or sample size.

CLimited

One independent source or anecdotal data. The result is preliminary.

DInsufficient

Data unavailable or inconsistent. Ratings are not published in headlines.

Data sources and partners

The diagram below shows only what actually feeds Ds's five weighted components — the full list, including calibration-only and planned sources, is in the table beneath it.

Lab test panelsWarentest · dTest · Which?Euroconsumers (OCU + Altroconsumo)Which? reliability surveylongitudinal fault tracking —its own instrument, separate from the tests aboveAlza warranty ledger2-year return rateexponential decay transform, R₀ = 1.0French repairability registryIndice de Réparabilité (data.gouv.fr)+ UFC-Que Choisir disclosureE-shop crowd ratingsstar ratings + recommend ratesacross e-shops · ≥5 reviews/productDdurabilityCqualityFfield failureXreturnsOrepairability
Four of the five scored components come from independent expert instruments — three general-purpose lab panels, one longitudinal reliability survey, one e-commerce warranty ledger, one government repairability registry. Consumer Index E is the only signal sourced from shoppers themselves, and is deliberately drawn outside the main flow here — for the same reason it never enters the weights.

Stiftung Warentest (DE)

Independent foundation, ~200 comparative tests per year, endurance tests, reference for calibrating R and O

Which? (UK)

Consumer association, ICRT member; longitudinal reliability surveys, calibration of T

Consumer Reports (US)

Non-profit ICRT member; annual survey with hundreds of thousands of households; primary benchmark for US-active brands

dTest (CZ)

Czech consumer association, ICRT member; laboratory tests and consumer context for the Czech market

iFixit.com

Repair community; repairability scores, teardown documentation, parts availability — direct input to R

Indice de Réparabilité (FR)

Mandatory government index since 2021; largest source of structured repairability scores in the EU — direct input to R

Indice de Durabilité (FR)

Successor to the repairability index for TVs (since 1/2025) and washing machines (since 4/2025); combines repairability and reliability. Mandatory open data on data.gouv.fr

Belgian Repairability Index

Mandatory federal index since 5/2025 (dishwashers, vacuums, laptops, etc.); the only one that also includes spare-part prices. Source: repairdatabase.be (FPS Public Health)

Fnac Darty Baromètre du SAV

Retailer service report; primary calibration benchmark for European brands

EU EPREL database

EU energy-label registry; official repairability class A–E and index for smartphones/tablets (0–5, Reg. 2023/1669) and tumble dryers (0–10, Reg. 2023/2534). For other appliances the registry publishes no repairability data — we use the category-level ecodesign spare-parts obligations, flagged as estimates

QualityDB (internal)

Proprietary longitudinal architecture (DCT/FMT layers, PELT ODA); basis for W and T indicators

Secondhand Deal Finder (internal)

Secondary-market data from Bazoš, Sbazar, Vinted; lifespan proxy for the O indicator

How the French Indice de Durabilité is actually measured

France's statutory index (decrees of 5 April 2024, under the AGEC law) is one of the few sources where the measurement methodology is publicly prescribed and auditable (by the DGCCRF). It currently covers TVs (since 1/2025) and washing machines (since 4/2025). The 0–10 score is split into two equally-weighted blocks:

R

Repairability

50% of score
  • Documentation — availability of service documentation for professionals and consumers.
  • Disassembly — number of steps and fastener types needed to replace a part.
  • Spare-parts availability — tracked separately across 4 distribution channels: repairer, manufacturer, distributor, end consumer (delivery delay + years of availability).
  • Spare-parts price.
F

Reliability (fiabilité)

50% of score
  • Resistance to wear and stress — the only criterion with a prescribed test (see below).
  • Ease of maintenance and cleaning — e.g. accessibility of a usage counter.
  • Commercial warranty and quality process — extended warranty beyond the statutory 2 years.

How reliability is physically tested

TVs

Tested on a minimum of 5 units of the same chassis/platform, following the manufacturer's own documented methodology. The manufacturer must keep proof of a successful pass on file for DGCCRF audits — there is no single government test protocol, it is an auditable self-declaration.

Washing machines

Measured as the number of completed wash cycles under the CEN EN 50731 standard (or an equivalent method). Points are awarded across a range of 1,400–3,400 cycles.

Finding from real data: across every manufacturer we checked (Sony, Bosch/Siemens, Haier/Candy, Samsung, LG, Miele), the "commercial warranty and quality process" score was consistently near zero (0.25–0.5 out of 10) — none of them offer an extended warranty or a declared quality process beyond the legal minimum. In practice this criterion is a market-wide "free point left on the table" rather than something that differentiates brands.

Sources: Légifrance (decrees of 5 April 2024 for TVs and washing machines), schema.data.gouv.fr (etalab/schema-indice-durabilite), economie.gouv.fr, ADEME.

Product matching and match confidence

Manufacturers sell slightly different variants of the same model in different countries (e.g. Bosch WGG14402BY for Czechia vs. WGG14402FF for France — same platform, different market suffix). Every score in our database and API therefore carries a match-type label, so it is always clear how confidently a score belongs to a specific product. We never present estimates as official values.

Match typeWhat it meansConfidence
exactIdentical EAN/GTIN code — the same specific product as registered by the manufacturer.highest
modelIdentical manufacturer model code.very high
familySame base model, different regional variant (market suffix). Hardware is generally identical; the match is labelled transparently.high
estimatedThe model has no direct score; the value is estimated from scores of the same brand in the same category (at least 3 rated models).indicative

Scores labelled official come from model values in official registries (FR/BE index, EU EPREL) or expert teardowns; scores labelled estimated are our own derived estimates and are always marked as such — in the database, in the API, and in partner widgets.

Existing research base

N = 408 respondents

Czech consumer survey

15 appliance brands, 6 quality dimensions, five-item planned-obsolescence belief scale (α = 0.896), expected lifespan of washing machine, refrigerator, dishwasher and dryer.

2,106,167 listings

Amazon marketplace dataset

Analysis of rating distributions, review concentration (Gini, Lorenz) and Bayesian-shrunk quality estimates; submitted to Journal of Big Data.

DCT/FMT layers, PELT ODA

QualityDB longitudinal architecture

Obsolescence-detection algorithm based on change-point detection in time series; basis for the ORS metric and MSK 2025 grant applications.

7 agencies · 30,681 tests · 326,797 sub-ratings

Cross-agency testing database

Harmonised results from European consumer organisations (Stiftung Warentest, dTest, Which?, OCU, Altroconsumo, UFC-Que Choisir). 986 products tested by two or more agencies — 612 once shared laboratories are accounted for. The basis for components D and C.

637 identical measurements

Shared laboratories: OCU and Altroconsumo

OCU (ES) and Altroconsumo (IT) both belong to Euroconsumers and republish part of the same test. Modelled explicitly (lab_group, shared_result_group) so one test cannot appear as independent confirmation by two sources. Underpins the revision of a Quality & Quantity paper.

ρ 0.481 → 0.17–0.21 after correction

Signal provenance: reviews vs. expert tests

The finding that one rating column carried customer stars for some sources and a rescaled expert verdict for others. Led to mandatory signal-type provenance in the schema and to the correction of two manuscripts. The method transfers to any database merging multiple sources.

Interval-censored Weibull MLE

Manuscript: appliance lifespan

Estimation of real lifespan for washing machines, refrigerators, dishwashers and dryers; benchmarked against EU Ecodesign reference values.

Waste Management Bulletin, R1

Washing-machine e-waste: collection gaps

Systematic analysis of WEEE flows for large appliances and gaps in collection systems; quantifies the difference between what is placed on the market and what is documented as collected.

Rating updates

Every rating is reviewed annually. A weighting change causing a shift of 2+ positions triggers a 30-day public consultation. The review date is visible on each record.

Challenge a rating

Manufacturers and directly affected parties may submit a rebuttal within 30 days of publication (max. 800 words). Rebuttals are published alongside the report — they never confer veto rights.

Report an error or submit a rebuttal

This is the abbreviated web version of the Ds v3 methodology. The full specification — formulae, code, validation results and data schema — is published on GitHub and Zenodo with each rating release.

Transparency and funding