TOPLINE:
Large language models (LLMs) demonstrate potential for automating the classification of ultraprocessed foods (UPFs), with ChatGPT o1 showing greater accuracy than DeepSeek-R1 when categorizing 1168 food items from the Brazilian Food Composition Table.
METHODOLOGY:
- Prior AI applications in dietary classification largely focused on computer vision and deep-learning approaches to identify foods in images and estimate nutrient content, but few studies have applied LLMs to NOVA classification.
- Researchers evaluated two LLMs (DeepSeek-R1 and ChatGPT o1) by tasking them with categorizing 1168 food items from version 7.0 of the Brazilian Food Composition Table according to the four NOVA categories.
- Both models received identical standardized prompts instructing them to assign each item to one of the four groups: unprocessed or minimally processed foods, processed culinary ingredients, processed foods, or UPFs.
- A trained nutritionist with experience applying the NOVA system manually classified all food items using the same standardized descriptions as the LLMs, which was used as a reference standard to compare the models’ results against.
- Agreement between models was assessed quantitatively using unweighted Cohen kappa; model performance was determined on the basis of accuracy, sensitivity, specificity, precision, and F1 score; discrepancies were determined qualitatively.
TAKEAWAY:
- Among the 1168 food items, the two models disagreed on 330 classifications; the unweighted Cohen kappa between DeepSeek-R1 and ChatGPT o1 was 0.768, indicating substantial agreement.
- ChatGPT o1 outperformed DeepSeek R1 for accuracy (98.0% vs 92.6%), sensitivity (94.7% vs 69.8%), and F1 score (95.6% vs 81.1%), with comparable results for specificity (99.0% vs 99.3%) and precision (96.5% vs 96.9%).
- DeepSeek R1 misclassified 86 items compared with 23 items by ChatGPT o1.
- Five recurring issues accounted for 67% of discrepancies: (1) DeepSeek mislabeling culinary preparations with UPF ingredients, (2) overlooking additives in processed meats, (3) classifying industrial sliced breads as artisanal, (4) misclassifying “diet/zero/light” products containing cosmetic additives, and (5) both models underestimating added sugars and flavorings in packaged and concentrated juices.
IN PRACTICE:
“The capacity to automate and accurately apply processing-based food classifications holds significant potential for enhancing dietary assessments, informing individualized counseling strategies, and improving the monitoring of UPF consumption in both research and practical settings,” the authors of the study wrote.
SOURCE:
The study was led by Marcus Vinicius de Oliveira Cattem, PhD, Nutrition Institute, Rio de Janeiro State University, Rio de Janeiro, Brazil. It was published online in Nutrition.
LIMITATIONS:
No widely accessible tool currently supports rapid clarification of NOVA classifications across diverse food contexts, limiting consistent application. Some food classifications remain challenging even for trained researchers because the Brazilian Food Composition Table lacks consistent, detailed product descriptions, ingredient lists, additive information, and preparation methods. The table primarily provides quantitative nutrient and energy content data rather than qualitative information on food processing.
DISCLOSURES:
The study received no specific funding. The authors reported having no known competing financial interests or personal relationships that could have influenced the work.
This article was created using several editorial tools, including AI, as part of the process. Human editors reviewed this content before publication.
Admin_Adham