We compared 16,978 test results from four European consumer testing organisations. Four of our studies: how strongly agencies agree (ρ = 0.84–0.92), whether higher prices mean higher quality, why the market rewards brands more than tested performance — and when the crowd agrees with the experts. With links to the full texts.
Why we compare agencies
Independent testing organisations (dTest in Czechia, Stiftung Warentest in Germany, and others) exist to reduce the information asymmetry between manufacturers and consumers. But that role only works if different agencies reach comparable conclusions. If every agency produced a different quality ranking, which one should you trust? Four of our studies examine this question empirically — on the largest harmonised dataset of consumer tests assembled in our region. Below we summarise the main results and link to the full manuscripts.
1. Cross-agency quality convergence across four European organisations
The core of our research: the harmonised QualityTest dataset with 16,978 test results and 125,978 sub-criterion ratings from four European organisations — dTest (CZ), Stiftung Warentest (DE), Which? (UK), and UFC-Que Choisir (FR). Agreement is measured on 282 products tested by multiple agencies. The result: strong rank-order agreement across all agency pairs (Spearman ρ = 0.84–0.92, p < 0.001), with no systematic proportional bias (Bland–Altman diagnostics). The overall price–quality correlation, by contrast, is only moderate (ρ = 0.33) and varies substantially by category — from 0.62 for smartwatches down to 0.19 for laptops. Sub-criterion analysis (Random Forest, XGBoost) further shows that sub-scores discriminate quality tiers (+33% F₁ over a majority baseline) despite low linear predictability (R² < 0.05) — agencies weight criteria non-linearly, "editorially".
Why it matters: if independent agencies converge on quality, their tests are a credible public standard — and aggregating them into a single score, like ours, is methodologically defensible. Our data shows the convergence is measurable and documented. And because price predicts quality only weakly, expert tests add the most value precisely in categories with strong branding and complex specifications.
Full text: Cross-Agency Product Quality Convergence (Google Drive)
2. Do you pay for quality, or for the brand?
This study links independent test results with market prices and firm-level data in the European home appliance industry — unlike brand-only studies, it directly tests the relationship between objectively measured quality and price. The finding: price premiums often attach to strategic brand position, not objectively measured quality, and persist even where differences in tested quality are modest. The study also connects these micro-level pricing patterns to broader strategic polarisation in the industry.
Why it matters: "more expensive = better" is the most widespread purchasing heuristic. Our data shows it works only weakly for appliances — and that consumers without access to test results systematically overpay for brands.
Full text: Objective Quality and Brand-Position Pricing (Google Drive)
3. A bifurcated market: quality mismatch in the global appliance industry
A hedonic pricing analysis linked to independent quality scores, framed by vertical differentiation theory and the rise of Chinese manufacturers. Tested quality translates weakly into prices, and the market is splitting in two: the premium segment sells brand architecture, challengers compete on price — and objective quality sits somewhere in between. The study shows that market bifurcation can coexist with only partial willingness to pay for tested quality.
Why it matters: market bifurcation means consumers need an objective quality signal more than ever. The price tag won't provide it.
Full text: Quality Mismatch and Market Bifurcation (Google Drive)
4. Do the stars agree with the lab? Expert tests versus crowd ratings
The first large-scale dataset linking standardised laboratory scores from four European testing agencies to consumer signals from seven retail platforms in five countries: 1,551 matched product–platform observations covering 1,332 unique laboratory-tested products across eight categories. The overall correlation between star ratings and lab scores is moderate (Spearman ρ = 0.481; 95% CI [0.43, 0.53]) — but it hides dramatic category heterogeneity: televisions ρ = 0.86, smartphones and laptops ρ ≈ 0.65, yet headphones ρ = 0.05 and smartwatches ρ = 0.09. The pattern matches the search/experience/credence-goods taxonomy: stars work where quality can be assessed objectively and fail where personal experience and fit dominate. Practically: a star-ranked top-20% shortlist recovers 80% of the lab-defined top quintile (a 4× lift over random) — stars are a usable coarse screener, not a measure of quality.
Why it matters: the answer to "can you trust the stars?" is "it depends on the category" — and our data quantifies, for the first time, in which categories you can and in which you cannot. Exactly where stars fail, independent tests and aggregated scores add the most value.
Full text: Expert vs Crowd Agreement (Google Drive)
The takeaway
Independent testing works — agencies agree enough that their results can be aggregated into a single signal. But the market doesn't yet reward that signal: people pay for brands, not for measured quality, and crowd ratings substitute for expert tests only in some categories. That is exactly why we display repairability and quality scores directly on products — so objective data is visible at the point of decision.
A note on the linked texts: the links lead to full-text author manuscripts (preprints) shared via Google Drive. Several are currently under peer review — where a published version exists, please cite that version.
Další čtení
Jak dlouho doopravdy vydrží spotřebiče? Shrnutí našeho výzkumu životnosti
Šest našich studií o životnosti spotřebičů: co očekávají čeští spotřebitelé, jak se vyvíjí skutečná životnost praček, proč víra v plánované zastarávání nemění nákupní chování — a kolik vysloužilých spotřebičů mizí mimo oficiální evidenci. Včetně odkazů na plné texty.
Recenze, reklamovanost a informační hodnota veřejných signálů v e-commerce: shrnutí našeho výzkumu
Analyzovali jsme 2,1 milionu produktů na Amazonu a 31 000 produktů největšího českého e-shopu. Šest našich studií: inflace hodnocení, strop predikce reklamovanosti, opravitelnost jako signál, selhání reputace, engagement signály — a rotace katalogu. Včetně odkazů na plné texty.