IP Library › Granted Patent US 12,394,222
Granted Patent B2
US 12,394,222 · App. 18/090,391 · Granted Aug 19, 2025

Methods and apparatuses for analyzing food using image captioning

Inventors: Dae Hoon Kim (Seoul, KR); Jey Yoon Ru (Seoul, KR); Seung Woo Ji (Seoul, KR)
Assignee: NUVI LABS CO., LTD.
G06V20/68G06F16/583G06T7/11G06V10/82G06V20/70G06T2207/30128
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,222
App. No.
18/090,391
Granted
Aug 19, 2025
Kind
B2
Abstract

Provided are a method and an apparatus for analyzing food using image captioning. A method for analyzing food using image captioning according to one embodiment of the present disclosure comprises generating image captioning data using food image features extracted from a food image; and generating a food name for the food image using the generated image captioning data.

Claims (23)

1. A method for analyzing food executed by an apparatus for analyzing food, the method comprising:

training an image encoder, a second text encoder, and a third text encoder using contrastive learning using a contrastive language-image pre-learning structure;

generating image captioning data using food image features extracted from a food image,

wherein the generating image captioning data extracts a first embedding having food image features through the image encoder to which the food image is input and generates image captioning data including food ingredients for the food image by inputting the extracted first embedding to the first text decoder, and wherein the generating image captioning data extracts a second embedding having food ingredient features through the second text encoder to which the inferred food ingredients are input and generates image captioning data including food recipes for the food image by combining the first embedding having the extracted food image features and the extracted second embedding and inputting the combination to the second text decoder; and

generating a food name for the food image using the generated image captioning data,

wherein the generating a food name extracts a third embedding having food recipe features through a third text encoder to which the generated food recipe is input and generates a food name for the food image by combining the extracted first, second, and third embeddings and inputting the combination to the third text decoder.

2. The method of claim 1 , wherein the generating image captioning data generates image captioning data including food ingredients using food image features extracted from the food image.

3. The method of claim 1 , wherein the generating image captioning data infers food ingredients using food image features extracted from the food image and generates image captioning data including food recipes for the food image using the inferred food ingredients.

4. The method of claim 1 , wherein the image encoder, the second text encoder, and the third text encoder are learned so that a triplet of the first to third embeddings for the same food image, food ingredient, and food recipe is more similar to each other than a triplet of embeddings for different food images, food ingredients, and food recipes.

5. Method of claim 3 , further including analyzing food nutrients in the food image using at least one of the inferred food ingredient, the generated food recipe, and the generated food name.

6. An apparatus for analyzing food using image captioning comprising:

a memory storing one or more programs; and

a processor executing the one or more programs stored,

wherein the processor is configured to:

train an image encoder, a second text encoder, and a third text encoder using contrastive learning using a contrastive language-image pre-learning structure;

generate image captioning data using food image features extracted from a food image,

wherein the processor extracts a first embedding having food image features through the image encoder to which the food image is input and generates image captioning data including food ingredients for the food image by inputting the extracted first embedding to the first text decoder, and wherein the processor extracts a second embedding having food ingredient features through the second text encoder to which the inferred food ingredients are input and generates image captioning data including food recipes for the food image by combining the first embedding having the extracted food image features and the extracted second embedding and inputting the combination to the second text decoder; and

generate a food name for the food image using the generated image captioning data,

wherein the processor extracts a third embedding having food recipe features through the third text encoder to which the generated food recipe is input and generates a food name for the food image by combining the extracted first, second, and third embeddings and inputting the combination to the third text decoder.

7. The apparatus of claim 6 , wherein the processor generates image captioning data including food ingredients using food image features extracted from the food image.

8. The apparatus of claim 6 , wherein the processor infers food ingredients using food image features extracted from the food image and generates image captioning data including food recipes for the food image using the inferred food ingredients.

9. The apparatus of claim 6 , wherein the image encoder, the second text encoder, and the third text encoder are learned so that a triplet of the first to third embeddings for the same food image, food ingredient, and food recipe is more similar to each other than a triplet of embeddings for different food images, food ingredients, and food recipes.

10. The apparatus of claim 8 , wherein the processor analyzes food nutrients in the food image using at least one of the inferred food ingredient, the generated food recipe, and the generated food name.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2023
From: KIM, DAE HOON; RU, JEY YOON; JI, SEUNG WOO
To: NUVI LABS CO., LTD.
Reel/Frame 062263/0298 →
Priority Claims (1)
KR 10-2022-0145991 · Nov 4, 2022 · national
Continuity (1)
Related Publication 20240153287A1 · May 9, 2024
References Cited (50)
US 9311568B1 · Feller · 2016 [cited by examiner]
US 9659225B2 · Joshi · 2017 [cited by examiner]
US 9892501B2 · Dehais · 2018 [cited by examiner]
US 9977980B2 · Joshi · 2018 [cited by examiner]
US 10380174B2 · Bhagwan · 2019 [cited by examiner]
US 11322149B2 · Kim · 2022 [cited by examiner]
US 11594050B2 · DeSantola · 2023 [cited by examiner]
US 11672446B2 · Hadad · 2023 [cited by examiner]
US 11712633B2 · Yu · 2023 [cited by examiner]
US 11942208B2 · Starson · 2024 [cited by examiner]
US 12064697B2 · Yu · 2024 [cited by examiner]
US 20160163037A1 · Dehais · 2016 [cited by examiner]
US 20190290172A1 · Hadad · 2019 [cited by examiner]
US 20190295440A1 · Hadad · 2019 [cited by examiner]
US 20210118447A1 · Kim · 2021 [cited by examiner]
US 20210166077A1 · Yu · 2021 [cited by examiner]
US 20210365687A1 · Starson · 2021 [cited by examiner]
US 20220292853A1 · DeSantola · 2022 [cited by examiner]
US 20230196802A1 · Gong · 2023 [cited by examiner]
US 20230222821A1 · Delp, III · 2023 [cited by examiner]
US 20230321550A1 · Yu · 2023 [cited by examiner]
US 20240087345A1 · Estrada Diaz · 2024 [cited by examiner]
CN 105512501A · 2016 [cited by examiner]
CN 108830154A · 2018 [cited by examiner]
Channam et al., “Extraction of Recipes from Food Images by Using CNN Algorithm,”  2021 Fifth International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC), Palladam, India, 2021, pp. 1308-1315… [cited by examiner]
Kumari et al., “Food Image to Cooking Instructions Conversion Through Compressed Embeddings Using Deep Learning,” 2019 IEEE 35th International Conference on Data Engineering Workshops (ICDEW), Macao, China, 2019, pp. 81… [cited by examiner]
Jelodar et al., “Calorie Aware Automatic Meal Kit Generation from an Image.” arXiv preprint arXiv:2112.09839 (2021). (Year: 2021). [cited by examiner]
Li et al., “Picture-to-amount (pita): Predicting relative ingredient amounts from food images.” In 2020 25th International Conference on Pattern Recognition (ICPR), pp. 10343-10350. IEEE, 2021. (Year: 2020). [cited by examiner]
Mezgec et al., “NutriNet: A Deep Learning Food and Drink Image Recognition System for Dietary Assessment.” Nutrients 9, No. 7 (2017): 657. (Year: 2017). [cited by examiner]
Ruede et al., “Multi-Task Learning for Calorie Prediction on a Novel Large-Scale Recipe Dataset Enriched with Nutritional Information,” 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 202… [cited by examiner]
Yang et al., “Yum-me: a personalized nutrient-based meal recommender system.” ACM Transactions on Information Systems (TOIS) 36, No. 1 (2017): 1-31. (Year: 2017). [cited by examiner]
Vasiloglou et al., “Assessing Mediterranean Diet Adherence with the Smartphone: The Medipiatto Project. Nutrients.” Dec. 7, 2020; 12(12):3763. doi: 10.3390/nu12123763. PMID: 33297550; PMCID: PMC7762404. (Year: 2020). [cited by examiner]
Machine translation of CN 105512501 A (Year: 2016). [cited by examiner]
Machine translation of CN 108830154 A (Year: 2018). [cited by examiner]
Min et al., “A survey on food computing.” ACM Computing Surveys (CSUR) 52, No. 5 (2019): 1-36. (Year: 2019). [cited by examiner]
Shao et al., “Towards the creation of a nutrition and food group based image database.” arXiv preprint arXiv:2206.02086 (2022). (Year: 2022). [cited by examiner]
Min et al., “Being a Supercook: Joint Food Attributes and Multimodal Content Modeling for Recipe Retrieval and Exploration,” in IEEE Transactions on Multimedia, vol. 19, No. 5, pp. 1100-1113, May 2017 (Year: 2017). [cited by examiner]
Wang et al., “Cross-Modal Food Retrieval: Learning a Joint Embedding of Food Images and Recipes With Semantic Consistency and Attention Mechanism,” in IEEE Transactions on Multimedia, vol. 24, pp. 2515-2525, 2022 (Year:… [cited by examiner]
Hakguder et al., “Smart Diet Management through Food Image and Cooking Recipe Analysis,” 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Las Vegas, NV, USA, 2022, pp. 2603-2610 (Year: 2022). [cited by examiner]
Chu et al., “Food image description based on deep-based joint food category, ingredient, and cooking method recognition,” 2017 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Hong Kong, China, 2017… [cited by examiner]
Carvalho et al., “Images and Recipes: Retrieval in the Cooking Context,” 2018 IEEE 34th International Conference on Data Engineering Workshops (ICDEW), Paris, France, 2018, pp. 169-174 (Year: 2018). [cited by examiner]
Zhang et al., “Sequential Learning for Ingredient Recognition From Images,” in IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, No. 5, pp. 2162-2175, May 2023 (date of publication: Nov. 1, 2022).… [cited by examiner]
Mao et al., “Visual Aware Hierarchy Based Food Recognition.” arXiv e-prints (2020): arXiv-2012. (Year: 2020). [cited by examiner]
Wu et al., “A large-scale benchmark for food image segmentation.” In Proceedings of the 29th ACM international conference on multimedia, pp. 506-515. 2021. (Year: 2021). [cited by examiner]
Wang et al., “Learning Structural Representations for Recipe Generation and Food Retrieval.” arXiv preprint arXiv:2110.01209 ( 2021). (Year: 2021). [cited by examiner]
Salvador, A. et al., “Inverse Cooking: Recipe Generation From Food Images,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10445-10454 (2019). [cited by applicant]
Radford, A. et al., “Learning Transferable Visual Models From Natural Language Supervision,” Proceedings of the 38th International Conference on Machine Learning, PMLR, vol. 139, (2021). [cited by applicant]
Chhikara, P. et al., “FIRE: Food Image to Recipe Generation,” arXiv, (2023). [cited by applicant]
Ma, Z. et al., “Food-500 Cap: A Fine-Grained Food Caption Benchmark for Evaluating Vision-Language Models,” Proceedings of the 31st ACM International Conference on Multimedia, pp. 5674-5685 (2023). [cited by applicant]
Extended European Search Report dated Apr. 12, 2024 as received in Application No. 23210766.4. [cited by applicant]