A chatbot knows what a banana is. It does not know your banana.#
Ask ChatGPT, or any general-purpose AI chatbot, how many calories are in a medium banana, a cup of cooked rice, or a slice of cheddar, and the published tests say it will usually land close to what a nutrition database says. Ask it about your lunch, and it also has to judge how much food is on the plate, which is where the published errors grow. The tests split cleanly along one line: a chatbot is a good lookup table for standard foods and a weak measuring device for real meals.
That split matters more than any single accuracy figure, because the number you get back is always delivered with the same confidence. This article walks through what researchers have actually measured between 2023 and 2026, where the drift enters (portion size, cooking fat, mixed dishes), and how to use a chatbot as one input to a calorie estimate that was never exact to begin with.
Single foods typed as text: close to the database#
The first tests asked the easiest version of the question. A 2024 study in Nutrition queried ChatGPT for the nutrient content of common foods and compared the answers with US Department of Agriculture data. Energy was the best-matched value: 97 percent of the chatbot's calorie figures fell within 40 percent of the USDA number, and repeated queries returned fairly consistent answers, judged by low coefficients of variation1.
Read that bar carefully. "Within 40 percent" sounds reassuring, but on a 600-calorie meal it permits anything from about 360 to 840. It is a pass mark for a quiz, not for a food log.
A larger 2026 study in the Journal of Nutrition tightened the comparison. Researchers typed frequently consumed US foods into four chatbots and matched each item to a research-grade food composition database (the Nutrition Coordinating Center database used in dietary studies). Agreement was high for energy and for the three macronutrients across all four models. It was the micronutrients that wobbled: vitamin D, folate, and iron showed poor agreement for some models. And two food categories contributed disproportionately to the scatter: condiments and mixed dishes2.
For a named food with a stated amount, a chatbot's figure can agree with a database. That says nothing about whether the amount matches what you ate.
That is the first useful rule. If you type a named, weighed, single ingredient ("200 g cooked chicken breast"), the tests suggest the answer is usually in the right neighborhood of a table value.
Meals: much larger errors in controlled tests#
Meals are where the published numbers change character. In a Swedish study of 52 standardized food photographs (16 single foods and 36 complete meals) served in small, medium, and large portions, two leading chatbots (ChatGPT and Claude) missed the reference energy content by a mean absolute percentage error of 35.8 percent, and missed the food's weight by 36 to 37 percent. A third model did markedly worse, with errors of 64 to 110 percent across nutrients3.
A separate study from Emory tested three sizes of one open-weight model on 510 lunches from a Peruvian cookbook, with the meal name and a photo turned into a written description first. The best model's calorie estimate agreed with the recipe value only 45 percent of the time, with a mean absolute error of 108 kcal per serving. The best protein result came from a different model size: 70 percent agreement and a mean error of 6 g. The authors concluded that performance fell "short of the precision required" for clinical or consumer use4.
| Study | Input | Reference | Energy result |
|---|---|---|---|
| Haman 2024 | Food names (text) | USDA data | 97% of values within ±40% |
| Lawabni 2026 | Common US foods (text), 4 chatbots | Research food database | High agreement for energy and macros; mixed dishes and condiments noisiest |
| Fridolfsson 2025 | 52 standardized photos, 3 portion sizes | Weighed food | ~36% mean absolute error for the two best models |
| Carrillo-Larco 2026 | 510 Peruvian lunches, described in text | Recipe values | Best model agreed 45% of the time; 108 kcal mean error |
| Rodríguez-Jiménez 2025 | 195 dishes, four levels of context | Recipes and weighed meals | Error fell as ingredients and amounts were added |
The pattern across the table is consistent: the less a meal resembles a single standard food, the further the number drifts.
The drift has a direction: big portions come out small#
The most practically important finding is not the size of the error but its shape. In the Swedish test, every model underestimated more as portions grew, with bias slopes between −0.23 and −0.503. A 2025 evaluation of ChatGPT on 114 meal photographs found the same tilt: good agreement on weight for small meals, poor agreement for medium and large ones, and 11 of 16 nutrients underestimated6. The photo-specific evidence is covered in our piece on AI photo calorie counters.
The practical upshot: a chatbot is likeliest to be wrong exactly when being wrong costs the most. A modest salad may come back near its true value; a heaped plate of pasta at a restaurant is where the estimate sags. Random error averages out across a week. Error that consistently runs low on your largest meals does not.
It helps to know the human comparison before judging the machine too harshly. Observers estimating cafeteria portions from digital photographs, using standard servings as their reference, produced estimates that correlated highly with weighed food, and agreed with direct visual estimates to within 1.5 g7. With a reference portion, a photograph can support good estimates. At the other end, people reporting their own intake can miss badly: in a classic study of dieters who believed they ate under 1,200 kcal a day, self-reported intake fell short of measured intake by 47 percent on average8. These studies used different tasks and people, so they cannot rank a chatbot against either group; they only show that human estimates are far from a perfect baseline.
Context is the lever you control#
The most encouraging result comes from a Spanish study that tested one chatbot on 195 dishes under four escalating conditions: photo only; photo plus standardized non-visual details; photo plus a full ingredient list with amounts; and the ingredient list with no photo. Accuracy for calories and all three macros improved step by step as context was added, with the largest gains on home-prepared, dietitian-weighed meals. Removing the photo while keeping the ingredients made things worse again, which the authors read as evidence that the visual cues carried real information rather than the model simply doing arithmetic on a recipe5.
The Lawabni finding about condiments and mixed dishes points at the same thing from the other side. What a chatbot cannot see, or was not told, it fills in with an average. The average stir-fry is not your stir-fry; the average dressing quantity is not the pour you used. Those invisible ingredients (oil, butter, sauces) are exactly the ones with the highest calorie density, which is why mixed dishes resist estimation by any method.
What that means in practice, as practical guidance:
- Name the fat. "Cooked in about a tablespoon of olive oil" changes the answer more than any detail about the vegetables.
- Give a size anchor. Grams if you know them; otherwise a comparison ("a dinner plate, rice covering about a third").
- Split the dish. List the components instead of naming the dish, the way the ingredient-list condition did.
- Say what was left. A chatbot assumes you ate what was served.
How to use one without trusting one number#
The evidence supports a clear division of labor. For single, standard foods, a chatbot's answer is a reasonable stand-in for a database lookup, though a package label or a weighed amount is better when you have one. For meals, treat the answer as a point inside a range, and widen that range for large portions, restaurant food, and anything with sauce, keeping in mind that big portions tended to come out low in the tests. We lay out why a band describes a meal better than a point in why calorie counts are ranges.
Two habits keep the error from compounding. First, keep your context in the conversation: the same query phrased differently can return different numbers, so a consistent description of your usual breakfast gives you a consistent estimate, which matters more for tracking trends than any single value's accuracy. Second, check the chatbot's arithmetic. A total that does not equal the sum of the listed components is a sign to re-ask with the components spelled out.
None of the published tests suggest that a general-purpose chatbot is useless for calorie counting. They suggest it is about as reliable as the description you give it, that it drifts low on big plates, and that its confidence does not change with its accuracy. Speaking your meal instead of typing it does not change that equation, as covered in voice logging accuracy; the words are the input either way.
FAQ#
Is ChatGPT accurate enough to count calories for weight loss?#
For single foods typed with a quantity, published tests found energy values close to reference databases. For whole meals, errors averaged roughly 36 percent in a controlled photo study, and grew with portion size. These studies tested nutrient estimates, not weight-loss outcomes, so they do not show whether chatbot logs are accurate enough to steer a calorie deficit; they do show a single meal figure is not precise.
Why does a chatbot give different calorie numbers for the same meal?#
The answer depends heavily on what the description implies about portion size and preparation. When details are missing, the model fills them with typical values, and small wording changes shift those assumptions. Giving the same explicit amounts each time narrows the spread.
Are AI chatbots better with photos or with text descriptions?#
Neither wins alone. In one test, adding ingredient lists with amounts to a photo gave the best results, and removing the photo while keeping the list made the estimate worse again. The combination beat either input on its own.
Do chatbots overestimate or underestimate calories?#
In the studies that measured direction, they mostly underestimated, and more so as portions got bigger. That is an average tilt: any single estimate can still be too high, but for a large or restaurant meal, lean your adjustment upward.
Sources#
- Haman M, Školník M, Lošťák M. AI dietician: Unveiling the accuracy of ChatGPT's nutritional estimations. Nutrition. 2024;119:112325.
- Lawabni R, Solano OR, Kampen LL, et al. Evaluation of Energy and Nutrient Estimates from Large Language Models Using Text-Based Queries. J Nutr. 2026;156(9):101712.
- Fridolfsson J, Sjöberg E, Thiwång M, Pettersson S. Performance Evaluation of 3 Large Language Models for Nutritional Content Estimation from Food Images. Curr Dev Nutr. 2025;9(10):107556.
- Carrillo-Larco RM, Gallo Ruelas M, Matsuzaki M, Tarazona-Meza C. Evaluating Large Language Model Accuracy in Predicting Peruvian Meal Nutrition. J Nutr. 2026.
- Rodríguez-Jiménez M, Martín-Del-Campo-Becerra GD, Sumalla-Cano S, Crespo-Álvarez J, Elio I. Image-Based Dietary Energy and Macronutrients Estimation with ChatGPT-5: Cross-Source Evaluation Across Escalating Context Scenarios. Nutrients. 2025;17(22):3613.
- O'Hara C, et al. An Evaluation of ChatGPT for Nutrient Content Estimation from Meal Photographs. Nutrients. 2025;17(4):607.
- Williamson DA, Allen HR, Martin PD, Alfonso AJ, Gerald B, Hunt A. Comparison of digital photography to weighed and visual estimation of portion sizes. J Am Diet Assoc. 2003;103(9):1139-1145.
- Lichtman SW, Pisarska K, Berman ER, et al. Discrepancy between self-reported and actual caloric intake and exercise in obese subjects. N Engl J Med. 1992;327(27):1893-1898.
Source: BurnWeek — "Can ChatGPT count calories accurately?", https://burnweek.fit/blog/can-chatgpt-count-calories-accurately/. Licensed CC BY 4.0: free to quote or reuse with a link to this page.



