# Why calorie estimates should be ranges: honest uncertainty as an API feature

September 29, 2026 · For Developers · 9 min read · https://burnweek.fit/blog/calorie-estimates-as-ranges-api/

> A single calorie number hides portion and label error. Return low, likely and high, and user corrections become cheap and easy to act on.

**Key takeaways**

- Portion estimates are the biggest error: people underestimated portions of beverages and medium energy-density foods by 30–46% in a lab study (Almiron-Roig 2013).
- Reference values drift too: reduced-energy restaurant foods measured 18% above stated energy on average (Urban 2010).
- Return low, likely and high per component; the width tells users which part of a meal to correct.
- Don't total a meal by adding every low and every high: independent errors partly offset, so a naive sum makes the meal range wider than it should be.
- Show the likely value on summary screens and the range in detail views; flag missing detail passively.
- Corrections that land outside the original range show where an estimate was too sure of itself.

## Portion error is the real error

Ask a nutrition API how many calories are in "a chicken and rice bowl" and most will answer with one number: 580. That number looks authoritative. It is also the output of a chain of guesses (how much chicken, how much rice, how much oil in the pan), and the single figure hides how wide those guesses are.

The largest guess is usually portion size. In a laboratory study, 32 adults estimated the number of portions in 33 foods and drinks; they tended to underestimate, errors differed by food type, and portions of beverages and medium energy-density foods were underestimated by 30 to 46 percent ([Almiron-Roig et al., 2013](https://pubmed.ncbi.nlm.nih.gov/23932948/)). The error is not a constant you can calibrate away, because it depends on what the food is and how it is presented.

The reference numbers are not exact either. Measured energy in 29 reduced-energy restaurant foods averaged 18 percent more than the stated values, and 10 supermarket frozen meals averaged 8 percent more, with some individual restaurant items reaching 200 percent of the stated value ([Urban et al., 2010](https://pubmed.ncbi.nlm.nih.gov/20102837/)). Even packaged snacks, where the conditions for accuracy are best, measured a median 4.3 percent above label once serving size was accounted for ([Jumpertz et al., 2013](https://pubmed.ncbi.nlm.nih.gov/23505182/)). Our [explainer on nutrition label accuracy](/blog/how-accurate-are-nutrition-labels) covers why labels are allowed to drift.

And people who self-report are not a clean signal. In a classic study of subjects who described themselves as diet-resistant, measured intake was substantially higher than reported intake, while their energy expenditure was close to predicted ([Lichtman et al., 1992](https://pubmed.ncbi.nlm.nih.gov/1454084/)). None of this means estimates are useless. It means a single number is a claim of precision that the inputs cannot support, and the [accuracy pillar](/blog/how-accurate-is-calorie-counting) makes the same case from the user's side.

If you are building a product on top of a nutrition API, this is your problem too. The number your UI displays is the number your users will trust, argue with, and eventually stop believing.

## Low, likely, high

The alternative is to return three numbers for every estimate: a low, a likely value, and a high. The likely value is the best single guess. The low and high bound the plausible range given what the input actually said. A photo of a plated meal with a visible portion gets a narrow range. "Some pasta" gets a wide one.

Our estimation API returns this shape for calories and protein, per component and per meal. A meal described as "chicken and rice with a bit of salad" comes back as a list of components, each with its own grams, calorie range and protein range, and the meal total is computed from them. Every component is independently editable, which matters later.

Three properties make ranges useful rather than decorative:

1. **The width carries information about the input.** A wide range on one component tells the user, and your UI, which part of the meal is doing the guessing. The fix is usually one detail ("it was about 150 g of rice"), not a re-log.
2. **Ranges are cheap to validate.** When a user corrects a component, you can see whether the correction landed inside the original range, which is a plain-language check on whether the ranges mean what they say.
3. **The likely value stays usable.** Nothing forces a UI to show three numbers everywhere. The range is data you can show when it helps.

## Adding ranges without inflating them

The obvious way to total a meal is to add all the lows, add all the highs and add all the likely values. That is correct for the likely value and wrong for the bounds. Adding lows assumes every component was simultaneously at its most optimistic, which is unlikely when the errors are roughly independent: overestimating the rice does not make you overestimate the chicken.

Project managers faced the same problem with three-point task-duration estimates in the 1950s, and the answer they wrote down is now textbook material: independent errors partly offset rather than stack, so the spread of a total grows more slowly than a simple sum of widths ([Malcolm et al., 1959](https://doi.org/10.1287/opre.7.5.646)).

The same idea applies to food. For a chicken and rice bowl with a likely 580 kcal, summing the component bounds produces a noticeably wider range than one that lets independent errors partly offset. The meal is no more precise than its parts, but its range no longer assumes every error went the same direction at once, and a day of five meals does not end with an absurdly wide band.

One caveat matters. Independence is an assumption, and some errors are correlated: if a cook is heavy-handed with oil, every fried component is high together. Even so, the alternative, summing bounds, produces ranges so wide that users learn to ignore them.

## What your UI should show

Ranges are an API feature first and a display decision second. What we have found works in our own app:

- **Headlines show the likely value.** A daily total or a meal card shows one plain number. We stopped prefixing it with a "≈" sign and stopped showing confidence badges on every card, because they added noise without changing anyone's decisions.
- **Detail views show the range.** When a user opens a meal, they see "680–780 cal" and the per-component breakdown with its ranges. That is where the width is actionable.
- **Flag missing detail passively.** A meal that lacks enough information for a tight range gets a quiet marker, and tapping it explains what would narrow the estimate. No pop-ups, no nagging.
- **Never show more digits than the range supports.** A likely value of 583 displayed next to a range of 480 to 690 invites the wrong kind of trust.

If your product computes something downstream (a remaining budget, a macro split, a reward), compute it from the likely value and keep the range available for explanation. Doing arithmetic on the bounds everywhere creates interfaces nobody can read.

## Ranges make corrections cheap

The best argument for ranges is what they do to corrections. A single-number API invites the user to reject the whole estimate when it looks wrong. A component-level range invites a targeted fix: the rice looked like 200 g, it was 150 g, so change that one component and the meal and day recompute.

That changes the economics of accuracy. Instead of trying to make every first estimate perfect, which the portion and label evidence above says is impossible, the system makes the first estimate honest about where it is unsure and makes the correction fast. The ranges then tell the user where their thirty seconds of attention would help most.

For an integrator, the practical checklist is short: store all three numbers per component, total them so independent errors can partly offset rather than with naive sums, show the likely value by default and the range on demand, and pay attention to corrections that land outside the range. That is the contract our API and the [nutrition SDK](/blog/stop-building-your-own-calorie-counter) are built around, and the [consumer-side explainer on ranges](/blog/why-calorie-counts-are-ranges) is a good page to link your users to.

## FAQ

### Why not just return one number and a confidence score?

A confidence score tells you how sure the estimator is but not in which direction or by how much. A low and high bound answers both questions in the units your users care about, and they can be summed across a day.

### Why not just add up the component ranges?

Adding bounds assumes every component was at its worst case at the same time, which rarely happens when errors are independent. Letting independent errors partly offset gives a more useful total, though correlated errors such as one cook's oil habits can still make a real total fall outside the combined range.

### Should my app show ranges on the home screen?

We show the likely value on summary screens and the range in detail views. Ranges everywhere tend to become visual noise; ranges nowhere hide the information users need to correct an estimate.

### How do I know whether an API's ranges are calibrated?

Log user corrections and check how often the corrected value lands inside the original low–high range. A well-calibrated estimator will contain most corrections; a narrow range that is routinely missed is overconfident.

## Sources

- [Almiron-Roig E, Solis-Trapala I, Dodd J, Jebb SA. Estimating food portions. Influence of unit number, meal type and energy density. Appetite. 2013.](https://pubmed.ncbi.nlm.nih.gov/23932948/)
- [Urban LE, Dallal GE, Robinson LM, Ausman LM, Saltzman E, Roberts SB. The accuracy of stated energy contents of reduced-energy, commercially prepared foods. J Am Diet Assoc. 2010.](https://pubmed.ncbi.nlm.nih.gov/20102837/)
- [Jumpertz R, et al. Food label accuracy of common snack foods. Obesity (Silver Spring). 2013.](https://pubmed.ncbi.nlm.nih.gov/23505182/)
- [Lichtman SW, et al. Discrepancy between self-reported and actual caloric intake and exercise in obese subjects. N Engl J Med. 1992.](https://pubmed.ncbi.nlm.nih.gov/1454084/)
- [Malcolm DG, Roseboom JH, Clark CE, Fazar W. Application of a Technique for Research and Development Program Evaluation. Operations Research. 1959.](https://doi.org/10.1287/opre.7.5.646)

Source: BurnWeek — "Why calorie estimates should be ranges: honest uncertainty as an API feature", https://burnweek.fit/blog/calorie-estimates-as-ranges-api/. Licensed CC BY 4.0: free to quote or reuse with a link to this page.
