Measurement and verification

How Energy Agent verifies savings

A number promised before a fix is an estimate. A verified saving is what the meter actually did afterwards, compared with what the building would have used if nothing had changed. This page explains how that comparison is built, how its quality is judged, and what the numbers on a verification report mean.

Savings are verified to IPMVP Option C (whole-building meter regression, or Option B on an isolating submeter), with the ASHRAE Guideline 14 fitness metrics CV(RMSE) and NMBE reported on every baseline, so a claim can stand up to a utility incentive review.

That is the sentence on the Energy Agent page. Below, every term in it, then every step behind it.

The sentence, unpacked

IPMVP International Performance Measurement and Verification Protocol
A published protocol, maintained by the Efficiency Valuation Organization (EVO), that sets out how to measure the energy saved by a change to a building in a way that two parties can agree on. Utilities, energy service companies and lenders use it as the shared rulebook: if a saving was measured “to IPMVP”, the other side knows what was measured, what was assumed, and where the uncertainty lies. It does not prescribe one method; it defines four options (A to D) and the discipline around each.
M&V Measurement and verification
The general term for proving a saving after the fact rather than predicting it before. Measurement is the metering; verification is the comparison against a baseline and the statement of how confident the result is. An estimate says “this should save about X”. M&V says “the meter shows it saved Y, give or take Z”.
Option C IPMVP Option C: whole-facility measurement
Use the building’s main utility meter, fit a statistical baseline model to the period before the change, then compare what the meter recorded after the change with what the model says it would have recorded. Option C is the right choice when the change is large enough to show up at the main meter, which is true of most schedule, setpoint and plant fixes. It is the method Energy Agent runs by default.
Option B IPMVP Option B: retrofit isolation with full measurement
The same idea, applied to a submeter that isolates the affected system, for example a chiller plant meter or a lighting panel, when the change is too small to be seen reliably at the main meter. In Energy Agent, Option B runs the same baseline model and the same fitness checks as Option C, on the isolating channel instead of the main meter.
Option A IPMVP Option A: retrofit isolation with key-parameter measurement
Measure one parameter, typically the kilowatt reduction, and stipulate the rest, typically the annual operating hours. A lighting retrofit is the classic case: the new fixtures draw a measured 4 kW less, the lights run an agreed 3,000 hours a year, so the saving is 12,000 kWh. Energy Agent supports it as a calculator. It has no baseline model, no confidence interval, and is never marked incentive-grade, because half the answer was agreed rather than measured.
ASHRAE American Society of Heating, Refrigerating and Air-Conditioning Engineers
The professional society whose standards and guidelines define good practice for building mechanical systems. Its Guideline 36 describes how air handlers should be controlled; its Guideline 14 describes how savings should be measured.
Guideline 14 ASHRAE Guideline 14: Measurement of Energy, Demand and Water Savings
The document that turns IPMVP’s principles into pass or fail numbers. Its most-used part says how well a baseline model has to fit the data before you are allowed to use it as a counterfactual, using two statistics, CV(RMSE) and NMBE, with limits that depend on whether the model works at hourly, daily or monthly resolution.
CV(RMSE) Coefficient of variation of the root-mean-square error
A measure of scatter. Take every hour in the baseline, find how far the model missed the meter, square those misses, average them, take the square root: that is the root-mean-square error, the typical size of a miss in kWh. Divide it by the average hourly use to express it as a percentage: that is the coefficient of variation. A CV(RMSE) of 5% means the model is typically off by 5% of an average hour. Guideline 14 allows up to 30% for hourly models.
NMBE Normalized mean bias error
A measure of lean. Add up every miss with its sign, so misses above and below cancel, and divide by the total use. A model can have small scatter and still run systematically high or low; NMBE catches that. It matters more than CV(RMSE) for savings, because a biased baseline turns straight into a biased saving. Guideline 14 allows within ±10% for hourly models; Energy Agent typically reports well under 1%.
Utility incentive review What a program evaluator checks before paying
Many utilities pay incentives for measured savings and employ an evaluator to check the claim. The evaluator asks the questions this page answers: what baseline, how long, how good a fit, what happened to weather, how wide is the uncertainty, and was the change actually on the date claimed. A report that carries those answers survives the review; one that carries a spreadsheet estimate does not.

What does “verified” mean?

You cannot measure a saving directly. Once the fix is in, the building only ever shows you what it uses now; the “before” no longer exists in the same weather with the same occupancy. So verification builds a stand-in for it: a model of how the building behaved before the change, fed with the weather and calendar of the period after. The model’s answer is called the counterfactual: what the meter would have read if nothing had been done.

The verified saving is simply the counterfactual minus the meter, added up over the period after the change. Everything else on this page is about making that counterfactual trustworthy and saying honestly how far it could be wrong.

Worked example 1: a known saving, recovered. A synthetic office with a 20 kWh per hour cut during occupied hours from October 27, 2024, so the true answer is known.

How to read it. Each point is one day after the change. The navy line is what the meter recorded. The teal line is the baseline model’s answer to “what would this building have used on this day, with this weather, if nothing had changed?” The shaded gap below the teal line is energy the building did not use: the verified saving. Over the 100 days shown the gap adds up to 15,604 kWh.

A known saving, recovered: metered use against the counterfactual5007501,0001,2501,500Oct 27Nov 16Dec 6Dec 26Jan 15Feb 3kWh/day
metered (actual)baseline model (counterfactual)verified saving
Computed by Energy Agent’s TOWT engine on a synthetic building (Synthetic office (the TOWT test building)), not a customer result.

Step 1: build a baseline the building would have followed

Energy Agent’s baseline model is a TOWT model: time-of-week and temperature. It is the standard approach for interval-data verification and works at hourly resolution. Time-of-week means the model carries one term for every hour of the week, 168 in all, so it learns that Tuesday at 2 pm looks different from Sunday at 2 am without being told the schedule. Temperature means it also learns how the load responds to outdoor air temperature, in segments rather than one straight line, because a building that is flat below 60 °F and climbs steeply above 75 °F should not be forced onto a single slope. The segments are placed where the building’s own temperatures actually fall, and the model learns separate temperature responses for occupied and unoccupied hours.

The baseline period is the stretch before the change that the model learns from. A full year is the target, because it lets the model see every season the performance period will bring. The engine will fit on as little as three weeks, but it will not call the result incentive-grade with less than a year.

How to read it. Two weeks of hourly energy from inside the baseline period, starting April 1, 2024. Navy is the meter. Teal is the model, which has learned the building’s weekly rhythm (168 hour-of-week terms) and its response to outdoor temperature. Where the two lines sit on top of each other the model has captured the building; where they part, the difference is the noise the fitness metrics measure.

A known saving, recovered: baseline fit over two weeks0255075100MonTueWedThuFriSatSunMonTueWedThuFriSatSunkWh/h
metered (actual)baseline model
Computed by Energy Agent’s TOWT engine on a synthetic building (Synthetic office (the TOWT test building)), not a customer result.

A second, simpler picture underlies the temperature part. Take each baseline day, plot its energy against its mean outdoor temperature, and a characteristic shape appears: flat on mild days, rising once the outdoor temperature crosses a threshold. That threshold is the balance point, the outdoor temperature at which the building starts needing cooling (or, on the other side, heating). ASHRAE’s change-point models describe this shape with two to five parameters, and Energy Agent searches for the balance points rather than assuming the textbook 65 °F.

The same idea gives heating degree days and cooling degree days (HDD and CDD): for each day, how many degrees the mean outdoor temperature sat below the heating balance point or above the cooling balance point. Utility engineers have used degree days to weather-normalise bills for decades; the change-point model is the same thing with the balance point measured instead of assumed. Energy Agent uses this daily model for weather normalisation and reporting, and the hourly TOWT model for verification, because the hourly model also captures the schedule.

How to read it. One dot per baseline day: the day’s mean outdoor temperature across, the day’s energy up. The teal line is the change-point model Energy Agent fitted, here a three-parameter cooling model. The dashed line marks the balance point it found, 65 °F: below it, outdoor temperature does not move the load; above it, every extra degree costs energy. The dots fall in two bands, weekdays above and weekends below, and a daily temperature model cannot tell them apart: that spread is the schedule, which the hourly TOWT model captures and this one cannot. It is why verification uses the hourly model and this daily model is kept for weather normalisation.

A known saving, recovered: daily energy against outdoor temperature5007501,0001,2501,5001,7502,00035°F45°F55°F65°F75°F85°FkWh/daybalance point 65 °F
one baseline daychange-point modelbalance point
Computed by Energy Agent on a synthetic building (Synthetic office (the TOWT test building)), not a customer result. Daily fit: CV(RMSE) 22.5%, NMBE 0%, R² 0.366.

Step 2: prove the baseline is good enough

A counterfactual is only as good as the model behind it, so before any saving is reported the model is scored on how well it reproduced the baseline it was trained on. Guideline 14 names the two scores.

Both are computed on the residuals, the hour-by-hour differences between what the meter recorded and what the model says, and both divide by the number of hours minus the number of things the model had to learn (its degrees of freedom), so that a model with many terms cannot look better simply by having more knobs.

CV(RMSE)
square root of the average squared miss, divided by the average use, as a percentageCV(RMSE) = 100 × √( Σ(actual − model)² ÷ (n − p) ) ÷ mean(actual)
NMBE
the sum of the signed misses, divided by the total the model should have matched, as a percentageNMBE = 100 × Σ(actual − model) ÷ ( (n − p) × mean(actual) )

n is the number of hours in the baseline; p is the number of parameters the model learned. R², also reported, is the share of the hour-to-hour variation the model explains; Guideline 14 does not gate on it, so neither does Energy Agent.

The bars depend on the model’s resolution, because an hourly model has more noise to explain than a monthly one. Energy Agent applies:

Model resolutionCV(RMSE) at mostNMBE withinWhere the bar comes from
Hourly (the TOWT verification model)30%±10%Guideline 14’s hourly limits
Daily (the change-point weather model)25%±5%Stricter than the guideline’s 30% hourly, looser than its 15% monthly
Monthly (utility bills)15%±5%Guideline 14’s monthly limits, quoted for reference

A baseline that fails either bar is not used to claim a saving. The engine says why and what would fix it, usually more data or weather.

Worked example 1, scored. The office baseline runs 299 days; the model fits with CV(RMSE) 5.0% and NMBE 0.04%.

What the engine reported for “A known saving, recovered”
MeasureValueBarResult
Baseline length299 days365 for incentive gradeshort of a year
CV(RMSE), hourly5.0%at most 30%pass
NMBE, hourly0.04%within ±10%pass
R² of the baseline0.992reported, not gated
Weather in the modelyesrequiredpass
Hours outside the trained temperature range0.1%at most 5%pass
Post-change period100 daysat least 14pass
Metered after the change87,282 kWh
Counterfactual (model)102,886 kWh
Verified saving+15,604 kWh (15.2%), $1,872 at $0.12/kWh
95% confidence interval14,627 to 16,581 kWh (± 977)
Fractional savings uncertainty (90%)5.3%lower is better
Day-to-day autocorrelation of residuals0.69widens the band
The saving that was actually seeded15,620 kWhrecovered within 0.1%
Annualisednot annualisedneeds 300+ post daysseasonal window
Incentive-gradeno: the baseline is 299 days, not a full yearno

Step 3: measure the gap and say how sure you are

With a model that passed, the performance period begins the day after the change. For every hour, the model is asked what the building would have used given that hour’s actual outdoor temperature and its place in the week; the meter says what it did use. The difference, summed over the period, is the verified saving. Priced at the tariff from your bills, it becomes dollars.

A single number is not enough for a reviewer, so the engine also reports how wrong it could be. The scatter of the baseline residuals, measured day by day, says how much the model is normally off; that scatter grows with the square root of the number of days, because errors accumulate. Consecutive days are not independent (a warm spell or a schedule change persists), so the engine measures the day-to-day correlation of the residuals and widens the band to match. The result is a confidence interval: a range that, with about 95% confidence, contains the true saving. The same arithmetic gives the fractional savings uncertainty, the half-width of a 90% interval as a share of the saving, the figure incentive programs most often ask for.

Two more guards. If more than 5% of the post-change hours had outdoor temperatures the model never saw in the baseline, the model is extrapolating, and the result is not incentive-grade. And if the performance period is shorter than about ten months, the engine refuses to annualise: a saving measured over a winter says nothing about summer, so it is reported for the window measured and no more.

How to read it. The line adds up the daily gap between the counterfactual and the meter, day after day, from the change onward. The grey band is the range the true total is likely to fall in (about 95% confidence); it widens with time because each day’s uncertainty accumulates, and it is widened further because one day’s error tends to carry into the next. A line that climbs steadily, with a band that stays clear of zero, is a saving you can defend. After 100 days: +15,604 kWh, give or take 988.

A known saving, recovered: cumulative gap with confidence band-2,00002,0004,0006,0008,00010,00012,00014,00016,00018,000Oct 27Nov 16Dec 6Dec 26Jan 15Feb 3kWh, cumulative+15,604 kWh
cumulative saving95% confidence band
Derived from the TOWT counterfactual Energy Agent fitted on a synthetic building (Synthetic office (the TOWT test building)); the product reports the total and the band, not this curve.

In the worked example the engine reports +15,604 kWh over 100 days, with a 95% interval of 14,627 to 16,581 kWh. The saving that was actually seeded into the data is 15,620 kWh, so the method recovered the truth within 0.1%. It is reported for the 100-day window only: 100 performance days is a seasonal window — annualizing linearly would misstate the year; savings reported for the measured window only.

Options A, B and C side by side

IPMVP’s Option D, calibrated simulation, is not offered; it belongs to new construction and deep retrofits where there is no before to meter.

OptionWhat is measuredWhat is stipulatedData neededConfidence intervalIncentive-gradeTypical use
Option CWhole-building meter, every hourNothingInterval data before and after; weatherYes, from the baseline residualsYes, when the four conditions below are metSchedule, setpoint, plant and controls fixes visible at the main meter
Option BAn isolating submeter, every hourNothingSubmeter interval data before and after; weatherYes, same methodYes, same conditionsChanges too small for the main meter: one plant, one panel, one system
Option AOne parameter, usually the kW reductionThe operating hoursA spot measurement and an agreed hours figureNoNeverLighting and constant-load equipment where hours are known and agreed

What “incentive-grade” means here

Energy Agent marks a verified saving as usable for an incentive application only when all four hold:

  1. The baseline model passes both Guideline 14 bars at hourly resolution.
  2. The baseline is at least 365 days long, so every season is represented.
  3. Outdoor temperature is in the model; a calendar-only counterfactual is not accepted.
  4. No more than 5% of the post-change hours fall outside the temperature range the model was trained on.

The first worked example below passes the fit bars comfortably but is not incentive-grade, because its baseline is 299 days. That is the gate working as intended.

Verification also says no

The second example is the natatorium on the Bay Palms demonstration campus. An “efficiency retrofit” changed the pool dehumidification setpoint the wrong way, and the building’s night floor stepped up by 48 kW. The baseline is a full year, the model passes every bar, and the verification is incentive-grade. It is also negative: the building used more after the retrofit than the counterfactual says it would have.

This is the point of verification. A promise is not a saving, and the same method that would have confirmed a real saving reports the loss with the same confidence band. The capital plan gets the truth either way.

Worked example 2: Ybor Athletics & Natatorium, 88,000 sq ft. Change on October 26, 2025; baseline October 25, 2024 to October 25, 2025; 309 days after.

How to read it. Each point is one day after the change. The navy line is what the meter recorded. The teal line is the baseline model’s answer to “what would this building have used on this day, with this weather, if nothing had changed?” The red area above the teal line is energy the building used that the model says it should not have: a verified loss. Over the 309 days shown the gap adds up to 363,071 kWh.

A retrofit that made it worse: metered use against the counterfactual5,5006,7508,0009,25010,500Oct 27Dec 28Feb 28May 1Jul 2Aug 31kWh/day
metered (actual)baseline model (counterfactual)verified loss
Computed by Energy Agent’s TOWT engine on a synthetic building (Ybor Athletics & Natatorium), not a customer result.

How to read it. The line adds up the daily gap between the counterfactual and the meter, day after day, from the change onward. The grey band is the range the true total is likely to fall in (about 95% confidence); it widens with time because each day’s uncertainty accumulates, and it is widened further because one day’s error tends to carry into the next. A line that falls steadily, with a band that stays clear of zero, is a loss the meter proves. After 309 days: −364,165 kWh, give or take 1,560. The band here is thinner than the line itself; on a synthetic building the model fits that well, and on a real one the band is wider.

A retrofit that made it worse: cumulative gap with confidence band-400,000-350,000-300,000-250,000-200,000-150,000-100,000-50,0000Oct 27Dec 28Feb 28May 1Jul 2Aug 31kWh, cumulative−364,165 kWh
cumulative loss95% confidence band
Derived from the TOWT counterfactual Energy Agent fitted on a synthetic building (Ybor Athletics & Natatorium); the product reports the total and the band, not this curve.
What the engine reported for “A retrofit that made it worse”
MeasureValueBarResult
Baseline length365 days365 for incentive gradefull year
CV(RMSE), hourly1.7%at most 30%pass
NMBE, hourly0.01%within ±10%pass
R² of the baseline0.984reported, not gated
Weather in the modelyesrequiredpass
Hours outside the trained temperature range0.1%at most 5%pass
Post-change period309 daysat least 14pass
Metered after the change2,538,071 kWh
Counterfactual (model)2,175,000 kWh
Verified loss−363,070 kWh (-16.7%), −$43,568 at $0.12/kWh
95% confidence interval-364,629 to -361,511 kWh (± 1,559)
Fractional savings uncertainty (90%)0.4%lower is better
Day-to-day autocorrelation of residuals0.32widens the band
Annualised−428,869 kWh/yrneeds 300+ post daysfull year measured
Incentive-gradeyes: every condition metyes

What you need to send

  • Interval electricity data from before the change: twelve months is the target, sixty days the minimum for a fit.
  • At least 14 days of interval data after the change; more is better, and 300 or more allows an annual figure.
  • The date of the change. Energy Agent’s change-point detection can suggest one if you are not sure, but a verification is never run on a guessed date.
  • Weather is fetched automatically from the building address. A weather file is optional.
  • For Option B, the interval data from the isolating submeter as well.

How to get your dataWhich analyses your data unlocksReadiness quizRequest a demo

Terms used on this page

Baseline period
The stretch of time before the change that the model learns the building from. Twelve months is the target so that every season is represented; sixty days is the least the engine will fit on.
Performance period
The stretch after the change over which the saving is measured. It starts the day after the change and must be at least 14 days; 300 or more days allow an annual figure.
Counterfactual
The model’s answer to “what would the meter have read in the performance period if nothing had changed?”, computed with the performance period’s real weather and calendar.
TOWT
Time-of-week and temperature: a baseline model with one term for every hour of the week (168) plus a segmented response to outdoor temperature. The standard baseline for interval-data verification.
Balance point
The outdoor temperature above which a building starts needing cooling, or below which it starts needing heating. Energy Agent finds it by search rather than assuming 65 °F.
HDD and CDD
Heating degree days and cooling degree days: for each day, how far the mean outdoor temperature sat below the heating balance point or above the cooling balance point. The traditional way to weather-normalise bills.
Change-point model
ASHRAE’s family of daily energy-versus-temperature models with two to five parameters (2P to 5P): flat below a balance point, sloped above it, or both. The simplest one that fits within 2% of the best is chosen.
Residual
For one hour or one day, the meter reading minus the model’s prediction. The fitness metrics and the confidence band are all computed from the residuals.
Degrees of freedom (n − p)
The number of observations minus the number of parameters the model learned. Dividing by it instead of by n stops a model with many terms from looking better than it is.
Coefficient of determination: the share of the variation in the meter data that the model explains, from 0 to 1. Reported for information; Guideline 14 does not set a bar on it.
Confidence interval
A range that contains the true saving with a stated probability, about 95% here. It is built from the scatter of the baseline residuals, grows with the square root of the number of performance days, and is widened for day-to-day autocorrelation.
Autocorrelation
The tendency of one day’s residual to resemble the previous day’s, because weather and occupancy persist. Ignoring it makes a confidence interval look narrower than it should; Energy Agent measures it and widens the interval accordingly.
Fractional savings uncertainty
The half-width of the 90% confidence interval divided by the saving, as a percentage. The single figure incentive programs most often ask for; lower is better.
Extrapolation
Asking the model about outdoor temperatures it never saw in the baseline. Energy Agent counts the share of performance hours that do this and refuses incentive grade above 5%.
Realization rate
Verified saving divided by the saving that was predicted before the fix. Energy Agent reports it for every implemented measure; it is how the estimate is held to account.

Common questions

Is a 95% confidence interval a guarantee?

No. It says that, given the model and the data, the true saving lies in that range with about 95% probability. It does not cover things the model cannot see, such as a change in occupancy that was not in the baseline. That is why the assumptions register lists what the analysis took as given.

Why does the engine refuse to annualise a short performance period?

A saving measured over one season depends on that season’s weather and schedule. Multiplying a winter figure by 365 ÷ days would claim summer savings that were never measured. The engine reports the window measured and annualises only once about ten months of performance data exist.

Can an Option A result be used for a utility incentive?

Not from Energy Agent. Option A stipulates the operating hours, so half the saving is agreed rather than measured, and there is no confidence interval. It is a useful calculator for lighting and constant loads; it is not marked incentive-grade.

Why do the worked examples use synthetic buildings?

So the true answer is known. The first example has a saving seeded into the data, which lets you see how close the method comes to the truth. Customer results are confidential and would not allow that check.

Written by Pelorus Engineering · Brent Waluzak, P.E. (FL)Examples computed

Screening-level estimates; confirm with measurement and verification before capital commitment. IPMVP is published by the Efficiency Valuation Organization; ASHRAE Guideline 14 by ASHRAE. Neither text is reproduced here.

See these rules run on your building.

Send twelve months of interval data and a BAS trend export. Energy Agent returns a ranked list of what is wasting money, what each item is worth per year at your tariff, and what to do about it.