resilienciahabitat · Mallorca

Evidence protocol

  • Statusdraft
  • OwnersNico & Laure
  • Last reviewed2026-09-14

How we get from “this seems like a good idea” to “this produced a measured result.” This protocol applies to every intervention we design, implement or recommend.

The six stages

  1. BASELINE      2. NEED         3. PREDICTION    4. INTERVENTION   5. VERIFICATION   6. PUBLICATION
  measure the      show the        state the        implement         re-measure        write it up,
  starting  ──▶    starting   ──▶  expected    ──▶  the change   ──▶  the same     ──▶  including
  point            point is a      result,                            way                misses
                   problem         before acting

The order matters, and the expensive mistake is starting at stage 4.

1. Baseline — measure the starting point

Before any intervention, establish what the system currently does, in numbers.

  • Define the metric before collecting it (see the metric definition template): quantity, unit, instrument, method, frequency, expected accuracy.
  • Record the conditions alongside the values: weather, occupancy, season, how the system was operated. Without these, a later comparison is meaningless.
  • Prefer a full annual cycle where the phenomenon is seasonal (pool, heating, cooling, irrigation all are). Where a year is not available, cover at least one full characteristic season and say so explicitly.
  • Sub-metering beats whole-site metering. “The house used 14,000 kWh” supports no decision; “the pool pump used 2,400 kWh of it” supports several.

Minimum honest baseline: if we genuinely cannot measure, we say the baseline is estimated, record the estimation method, and carry that uncertainty through every subsequent claim.

2. Need — establish that there is a problem worth solving

A baseline on its own is not a case for action. Compare it against a reference: a benchmark, a physical lower bound, a comparable site, a regulatory standard, or the system’s own theoretical performance. State the size of the gap and what it costs per year in money, resource units, comfort or risk exposure.

If the gap is small, the correct output is “leave it alone”, and that is a real result.

3. Prediction — state the expected result before acting

Written down, quantified, and with an uncertainty range, before implementation:

  • Expected change in each affected metric, with a range (not a single number).
  • The mechanism: why we expect it, physically.
  • Time to effect, and how long verification will need to run.
  • What would falsify it — what result would mean the intervention did not work.
  • Known confounders and how we will handle them.

Predicting in advance is the single discipline that separates this from marketing. A prediction written after the fact always fits.

4. Intervention — implement, and record what actually happened

Record the as-built reality, not the plan: what was installed, model numbers, settings, who did it, what changed during the work, what was compromised, actual cost. Log the date precisely — it is the boundary of the before/after comparison.

Change one thing at a time where possible. Where several changes are bundled (usually for cost reasons), record that the effects will not be separable, and say so in the result.

5. Verification — re-measure, the same way

  • Same metric, same instrument, same method as the baseline. A change in method is indistinguishable from a change in performance.
  • Normalise for conditions. A mild winter reduces heating consumption on its own. Adjust for degree-days, rainfall, occupancy and any other driver that moved between the two periods, and show the adjustment.
  • Compare against the stage 3 prediction, and report the difference in both directions: under-performance and over-performance are both signals that our model of the system was wrong.
  • Where the result is ambiguous, say it is ambiguous.

This is the standard practice of measurement and verification in the energy sector — the IPMVP framework is the established reference and is worth adopting rather than reinventing. (To do: review IPMVP option A/B/C/D and record which we use per domain.)

6. Publication — write it up

Every completed cycle produces a case study containing: baseline, need, prediction, what was done, verified result, cost, and what we would do differently. Published whether or not it worked, subject to the site owner’s consent and anonymity preferences.

Guardrails against fooling ourselves

  • Pre-register the prediction. Commit it to the repository before implementation, so it is timestamped in version control.
  • Beware the seasonal illusion. Comparing August to October will show a dramatic improvement in cooling energy that has nothing to do with the intervention.
  • Beware the attention effect. Sites behave better simply because someone is watching them. If behaviour changed as well as hardware, attribute honestly.
  • Beware the vendor’s number. Manufacturer performance figures are measured under standard conditions that your site does not have.
  • Keep a control where you can. An unmodified comparable circuit, zone or site, measured over the same period, is worth a great deal.
  • Report cost with the same rigour as benefit. Including maintenance and expected replacement.

Applying this when the client will not pay for a baseline

A common, real constraint: someone wants the pool fixed, not a twelve-month study. Practical fallbacks, in descending order of quality:

  1. Short intensive baseline — days or weeks of high-frequency sub-metering rather than a year of low-frequency data.
  2. Retrospective baseline from existing records — utility bills, water bills, meter photographs, service history.
  3. Reference baseline from our own accumulated site data, adjusted for the differences.
  4. Physics-based estimate from first principles, with assumptions listed.

Any of these is acceptable if the tier is labelled and the uncertainty is carried forward. What is not acceptable is presenting fallback 4 as if it were stage 1.