Sep 1, 2026
Can ChatGPT Count Calories? What 4 Studies Show (2026)
Peer-reviewed tests put ChatGPT's photo calorie error near 30%. Claude ties it. Gemini doubles it. What four studies found, how to shrink the error, and where a chatbot still fails.
ChatGPT can estimate calories, but peer-reviewed tests put its photo error around 30 to 36%, and it underestimated portions on 76% of the meal photos in the sharpest test. Claude performs about the same. Gemini is roughly twice as wrong.
How accurate is ChatGPT calorie counting?
From a photo, ChatGPT misses calories by 30 to 36% on average. That range comes from four tests published between 2025 and 2026. The table first, the detail after.
| Study | What was tested | Result |
|---|---|---|
| Three-model photo test, 2025 | 52 weighed food photos went through ChatGPT, Claude, and Gemini, with meals shot at three portion sizes. | ChatGPT missed food weight by 36.3% on average and Claude by 37.3%. Both missed calories by 35.8%. Gemini's errors ran from 64% on calories to 110% on protein. Every model underestimated more as portions grew. |
| Portion accuracy test, 2025 | ChatGPT analyzed 114 photos of 38 real meals, each served at three portion sizes and weighed for ground truth. | It identified the foods with 93% precision but underestimated the meal's weight on 76% of the photos. Small portions came out fine. Medium and large ones did not. 10 of 16 nutrients disagreed with the weighed values. |
| Four-model carb count, 2026 | ChatGPT, Gemini, Claude, and DeepSeek counted carbs in 124 typed meal descriptions, scored against clinicians using USDA and Italian food tables. | ChatGPT was the only model to land within 5% of the clinicians' counts, with the smallest bias and the steadiest answers. Claude drifted lowest and spread widest of the four. |
| Typed-context test, 2025 | ChatGPT-5 estimated 195 dishes, from photo alone up to photo plus a full typed ingredient list with amounts. | From the photo alone it missed calories by 30.5% on average. The typed ingredient list cut the error to 13.9%. |
If a number here has gone stale or a newer test contradicts one, tell us and we'll correct it.
The biggest head-to-head fed 52 photos of weighed foods to ChatGPT, Claude, and Gemini. ChatGPT missed food weight by 36.3% and calories by 35.8% on average. Claude landed at 37.3% and 35.8%. Gemini ran 64 to 110% off depending on the nutrient. The one pattern shared by all three: the bigger the portion, the bigger the underestimate.
The portion test explains where those misses come from. ChatGPT looked at 114 photos of 38 real meals and named the foods with 93% precision, then underestimated the meal's weight on 76% of the photos. Small plates came out fine. Medium and large plates didn't, and 10 of the 16 nutrients checked disagreed with the weighed values. It knows what's on the plate. It can't tell how much.
The newest test, published in 2026, dropped photos entirely. Four models counted carbs from 124 typed meal descriptions, and ChatGPT was the only one to land within 5% of clinicians' counts, with the steadiest answers of the four. Claude drifted lowest and spread widest. One catch: the test meals spelled out every food and amount, and the authors warn that vaguer input would do worse.
The fourth study measured what context buys you. ChatGPT-5 estimated 195 dishes and missed calories by 30.5% from the photo alone. A typed ingredient list with amounts cut that to 13.9%. That drop matters, and we'll come back to it, because it's the closest thing this literature has to a fix.
Can ChatGPT count calories from a photo?
It can name the food. It can't weigh it.
A photo carries no depth, no density, and no oil. A bowl of rice shot from above looks the same at 150 g and 250 g (5 oz and 9 oz). A tablespoon of oil in the pan is about 120 calories that show up in zero pixels. That's why the same test produced 93% precision on food identification and a 76% portion miss: identification is a vision problem, portioning is a measurement problem, and a flat image only solves the first one.
The misses aren't random either. Every model tested underestimated more as portions grew. Your smallest snacks get counted almost right, and your biggest meals lose the most.
Why the errors skew low
Say dinner is 700 calories and ChatGPT logs it 35% low. You write down 455 and eat 700. Those 245 phantom calories land on top of a standard 500 calorie deficit, and one dinner has silently erased half your day's progress while the log still looks perfect. Stack that across a week of big meals and "the diet that isn't working" is the tracker that never worked.
An error that skewed high would only make you eat a little less. An error that skews low tells you you're on track while you aren't. That's the expensive kind.
How to make ChatGPT more accurate at counting calories
The error isn't set in stone. Three things shrink it:
- Type the ingredients with amounts. This is the big one. It cut ChatGPT-5's error from 30.5% to 13.9%. "Chicken breast 180 g, white rice 200 g cooked, broccoli, olive oil 1 tbsp" beats any photo.
- Name the cooking method and the fat. Fried or grilled, butter or spray, dressing or dry. These are the calories a photo cannot see, so say them out loud.
- One meal per message, weights in grams. Give it one clean problem at a time and anchor the thing that dominates the plate. The carb study's meals were described exactly this precisely, and that precision is what its accuracy depended on.
My take: 13.9% is the most honest number in this whole literature. To earn it you have to type out every ingredient with its amount, and at that point you've done the estimating yourself.
And that's the ceiling. Even at 14%, every number lands in a chat reply. Nothing adds it to a running total, nothing compares it to a target, and by dinner your lunch is a message scrolled way up the thread. Accuracy was never the only problem, and the tracking half of it is a whole separate comparison.
Can Claude count calories?
About as well as ChatGPT, which is the problem.
On weighed photos Claude missed food weight by 37.3% against ChatGPT's 36.3%, and both missed calories by the same 35.8%. That's a statistical tie, with the same portion blindness and the same downward drift on bigger plates. On typed carb descriptions Claude came in last of the four models tested, with the lowest and widest guesses.
ChatGPT vs Claude vs Gemini for calorie counting
| Metric | ChatGPT | Claude | Gemini |
|---|---|---|---|
| Calorie error from a photo | It missed calories by 35.8% on average. | Also 35.8%. A statistical tie. | It missed by 64.2%, roughly double. |
| Portion weight error | It guessed food weight 36.3% off on average. | 37.3% off, effectively the same. | 65% off on average. |
| Behavior on bigger plates | It underestimated more as portions grew. | Same downward drift on bigger portions. | Same direction, roughly twice the size. |
| Consistency on typed carb counts | The only model within 5% of clinicians' counts, and the steadiest of the four tested. | The largest low bias and the widest spread of the four. | Landed within 10%, not 5%. |
Photo figures from a 2025 weighed-photo test, consistency figures from a 2026 carb-counting test. Both linked in the sources below.
The consistency row is the newest evidence on this page. In the 2026 carb test, ChatGPT had the smallest bias, the narrowest spread, and the highest agreement with clinicians across 124 meals. It was the only model of four to meet the 5% bar. Gemini and DeepSeek managed 10%. Claude missed both.
One caveat worth holding loosely: the photo head-to-head ran on 2024 versions of each model. But the ChatGPT-5 test is from late 2025 and still landed at 30% from a photo alone, so the ceiling is moving slowly.
Verdict: if you're going to use a chatbot for food math anyway, use ChatGPT. It beats Gemini by double and edges Claude on consistency. Least wrong is still 30% wrong.
Why no prompt fixes the error
Out of the box, a chatbot's calorie numbers are generated on the spot, not looked up. There's no database row behind the figure, no arithmetic you can audit, and no guarantee the same meal returns the same number tomorrow.
That's why the error survives every model upgrade and every prompt trick. Better context shrinks the guess, and 13.9% is a genuinely smaller guess. It's still a guess. Nothing you type gives the model a database to check.
What to use instead
Depends on the job the number has.
A one-off target. If you want to know what to eat per day and nothing more, a formula beats a chatbot. Our free calorie calculator runs the published equations and gives you workout-day and rest-day targets with no prompting.
Daily tracking. If you're logging food every day toward a goal, use a tracker with a real database. Any of them beats a chat thread. If you also lift, here's how the main one compares.
Tracking that adapts. APEX splits the job exactly where the studies say it breaks. The model only names what it sees: a food ID and a gram amount. It's structurally unable to output a calorie number. Code takes over from there, looks each food up in a catalog of 1,429 entries, 853 of them traced to a named food-composition record (733 USDA, 120 from national tables across 14 countries), and does the arithmetic. The same meal returns the same number every time. Nothing saves automatically: the photo comes back as one editable card per ingredient with a confidence score, and the app tells you straight that the amount is the likeliest thing to be off.
To be equally straight here: that gram amount is still a model guess, the same failure the portion study measured. The difference is where the error lives. In APEX it sits in one visible field you can fix in two taps before it touches your totals. In a chatbot it hides inside every number on the screen.
If you want the number checked against a database instead of guessed, download APEX and try the full experience free for a month.
Manual and saved-meal logging are free forever, and you get 10 free AI analyses to test the photo pipeline yourself. On the App Store today, with the Android waitlist open.
ChatGPT calorie counting FAQs
Why does ChatGPT underestimate portion sizes?
Because a photo carries no depth, no density, and no oil. The model sees the surface of the food, not how deep the bowl goes or what it was cooked in, so it guesses low. In a weighed test it underestimated the meal's weight on 76% of photos, and every model tested underestimated more as the portion grew. Small plates came out fine. Big ones lost the most.
Is Claude more accurate than ChatGPT for calories?
No. On weighed food photos Claude missed portion weight by 37.3% against ChatGPT's 36.3%, and both missed calories by the same 35.8%. That's a tie. On typed carb descriptions Claude was actually the least accurate of four models tested, with the lowest and widest guesses. Switching chatbots doesn't fix calorie counting.
Which AI is the most accurate calorie counter?
ChatGPT, mostly on consistency. In a 2026 test it was the only model of four to match clinicians' carb counts within 5%, and on photos it edged Claude while Gemini's errors ran roughly double. It's the least wrong of the big three, and still about 30 to 36% off from a photo alone.
Can ChatGPT count carbs?
Better than it counts calories from photos. Given precise typed descriptions with amounts, it was the only model of four to match clinicians within 5%. Two caveats. The test meals were described more precisely than most people type, and the researchers frame these tools as a help alongside diabetes education, never a replacement for it. If insulin depends on the number, keep a human in the loop.
Is ChatGPT accurate enough to run a cut on?
No. A 35% miss on a 700 calorie dinner is 245 calories, half of a standard 500 calorie daily deficit, gone in one estimate. And since the misses skew low, your log looks compliant while the scale stalls. Fine for a one-off sanity check. Too loose for twelve weeks of decisions.
Is a calorie tracking app more accurate than ChatGPT?
At the arithmetic, yes, because a tracker looks numbers up instead of generating them. In APEX the model only names the foods and gram amounts, then code does the math against a catalog of 1,429 foods. The gram estimate can still be off, exactly like ChatGPT's, but it sits in one visible field you can edit before saving instead of hiding inside every number.
Sources
- Fridolfsson J, Sjöberg E, Thiwång M, Pettersson S. Performance Evaluation of 3 Large Language Models for Nutritional Content Estimation from Food Images. Current Developments in Nutrition, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12513282/
- O'Hara C, Kent G, Flynn AC, Gibney ER, Timon CM. An Evaluation of ChatGPT for Nutrient Content Estimation from Meal Photographs. Nutrients, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC11858203/
- Zagaroli L, Caione N, Zara S, Guerra F, Zugaro A, Baroni MG, Iezzi ML, Delvecchio M. Accuracy of ChatGPT, Gemini, Claude and DeepSeek in Carbohydrate Counting. Diabetes, Obesity and Metabolism, 2026. https://pmc.ncbi.nlm.nih.gov/articles/PMC13243987/
- Rodríguez-Jiménez M, Martín-del-Campo-Becerra GD, Sumalla-Cano S, Crespo-Álvarez J, Elio I. Image-Based Dietary Energy and Macronutrients Estimation with ChatGPT-5: Cross-Source Evaluation Across Escalating Context Scenarios. Nutrients, 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12655113/
APEX is not affiliated with OpenAI, Anthropic, or Google. Study figures were checked against the published papers in September 2026 and model behavior can change.