energy · construction · family: it worked in the lab
measuring heat lossin a lived-in house
A Home's real heat loss cannot yet be measured reliably while anyone lives in it — smart-meter methods missed the reference by up to 50% in a blind test
Problem statement
The single number that says how well a house's fabric keeps heat in — the heat transfer coefficient (HTC, in W/K) — can only be measured reliably by a co-heating test: the house is emptied, heated to a constant temperature for two to three weeks in winter, and the input power is regressed against the indoor–outdoor temperature difference. That is impossible for the hundreds of millions of occupied homes that retrofit programmes need to assess before and after work. The IEA EBC Annex 71 project (2016–2021, nine countries) set out to extract HTC instead from data occupied homes already produce — smart-meter gas and electricity, a few room temperature sensors, local weather — and found that the methods work in principle but not yet as a quality-assurance tool: static and dynamic statistical estimates on five monitored case-study houses "deviated often significantly (20% or more) from the target value," and in a blind test on five inhabited UK dwellings different participants' estimates "for some of the buildings were in close agreement with the target values (co-heating test results), while for other buildings deviations up to almost 50% were found," with "a significant underestimation of the HTC" across all five houses.
Why this matters
Retrofit finance, energy-performance certificates, heat-pump sizing, and pay-for-performance schemes all depend on knowing a building's actual fabric performance rather than the value a designer or assessor assumed — and the two routinely diverge (the "performance gap"). Without an in-use measurement, a retrofit programme cannot verify that installed insulation delivered the heat-loss reduction it was paid for, cannot rank a housing stock by real need, and cannot detect the leaky outliers that most deserve attention. Annex 71's own conclusion is that the methods "show promise" but that "a further in-depth analysis on more case studies is advisable to turn the methods into reliable tools to be used in actual performance assessment," with "specific attention" needed on "uncertainty and repeatability before moving to large scale applications." The UK's SMETER programme, which tested eight smart-meter-based technologies against co-heating baselines in 30 occupied houses, found the concept "effective" but the field still lacks an accepted accuracy standard for the occupied case.
What’s been tried and why it hasn’t worked
Annex 71 followed IEA EBC Annex 58, which had established that HTC can be identified from dynamic tests in unoccupied buildings; the step to occupied buildings is where the methods lose robustness. The Subtask 3 team derived the full heat-balance equation and tested every simplification against five well-characterised houses: static approaches (averaging, single and multiple linear regression) and dynamic approaches (ARX and state-space models). Both families gave internally consistent results but often missed the co-heating reference by 20% or more, and the analysis showed that "assumptions on almost all parameters (measurement time and period, internal heat gains, temperature averaging,…) significantly impact the outcome" — for example, using one room's temperature instead of a multi-sensor average changed the computed HTC by up to 15%, and splitting metered gas between space heating and domestic hot water, or estimating occupant and appliance gains, were among the most influential guesses. In the blind test on five occupied UK homes, static methods gave consistent answers across participants but dynamic methods — which give the analyst freedom over data frequency and model selection — diverged between users. The binding constraint is therefore not computing power or algorithms but disentangling three confounded heat sources (fabric, systems, occupants) from a handful of low-resolution signals, in a building whose occupants open windows, run showers, and switch heating on and off unpredictably.
What would unlock progress
Two things would move the field from "promising" to "usable": a validated, standardised uncertainty budget for in-use HTC (so a result comes with a defensible confidence interval rather than a point estimate that may be 50% off), and cheap ways to observe the confounders directly — a domestic-hot-water disaggregation signal, an occupancy or window-state proxy, a solar-aperture estimate — so that they stop being free parameters. Adjacent fields have solved structurally similar problems: non-intrusive load monitoring disaggregates appliance signatures from a single electricity meter, and system-identification practice in process control routinely delivers parameter estimates with credible intervals from noisy operational data. Annex 71's own framing points at the cheapest data source: the building's "on-board monitoring systems" — the controls and meters of its own heating services — rather than dedicated test instrumentation.
Entry points for student teams
A team could take one instrumented occupied house (or a public dataset such as the UK SMETER release cited in the Annex 71 report) and build a Monte Carlo uncertainty analysis over the assumption set Annex 71 identified — measurement window, gain estimates, temperature averaging, DHW split — to produce HTC estimates with intervals and to rank which assumption most inflates the error. A sensing-focused team could prototype a low-cost add-on (a single hot-water pipe temperature clip, a window-contact sensor, or a boiler-bus reader) and measure how much each shrinks the estimate's uncertainty. A data-science team could test whether a fixed, non-discretionary dynamic-model recipe removes the between-analyst spread seen in the blind test. Relevant skills: building physics, statistical system identification, time-series analysis, low-cost sensing.
Genome — every gene is a door
Structural cousins — same reason stuck, other fields
Sources
"Building energy performance assessment based on in-situ measurements — Physical Parameter Identification (Subtask 3 report)," IEA EBC Annex 71 (Operating Agent: Staf Roels, KU Leuven), accessed 2026-08-17; "Factsheet — Building Energy Performance Assessment Based on In-situ Measurements, EBC Annex 71," IEA EBC, accessed 2026-08-17; "New report finds smarter way to make homes more energy efficient," Loughborough University (SMETER programme summary), accessed 2026-08-17 go to source 1 ↗ go to source 2 ↗ go to source 3 ↗
verification notes (working record)
The collection team’s own sourcing notes for this brief, kept verbatim:
IEA EBC (Energy in Buildings and Communities Technology Collaboration Programme) annexes are multi-year international research projects; the Subtask 3 report is the primary technical deliverable and all quantitative claims above (20%+ deviations, up to ~50% in the blind test, up to 15% from temperature averaging, systematic underestimation) are quoted from its Chapter 8 conclusions and Chapter 9. Nine participating countries and 2016–2021 duration are from the Annex 71 factsheet. The SMETER details (eight technologies, 30 occupied houses, co-heating baseline) are from the Loughborough University summary; the SMETER technical evaluation report itself was not read for this brief and its per-technology accuracy figures should be pulled before quoting further. `failure:lab-to-field-gap` chosen because the methods are validated in unoccupied/dedicated-test conditions (Annex 58) and degrade specifically in occupied deployment; `constraint:coordination` and `failure:not-attempted` were considered and rejected — the problem is heavily attempted (two IEA annexes and a national programme) and the barrier is signal separation, not actor alignment. Related collection briefs: `energy-building-performance-prediction-gap` (models overestimate retrofit savings — a modelling problem; this brief is the measurement problem that would let those models be checked) and `energy-building-retrofit-digital-gap` (buildings lack digital models). Check the IEA EBC annex list for successor projects on in-situ performance measurement before treating the Annex 71 conclusions as the current state of the art.
Verifier note (2026-08-17): Annex 71 Subtask 3 report (114 pp, KU Leuven, August 2021) fetched and read; all quoted phrases located verbatim in §8.6 and Ch. 9 ("significant underestimation" refers to the initial guideline-based blind estimates for all five houses); nine participating countries and 2016–2021 confirmed from the factsheet; SMETER details (eight technologies, 30 homes, co-heating blind trial) confirmed on the Loughborough page. Title hedged from "cannot be measured" to "cannot yet be measured reliably" to match Annex 71's own "show promise" framing.
Verified at intake 2026-08-17: gate (net) + adversarial source check + contested-tag second coding.