IP Library › Granted Patent US 12,511,902
Granted Patent B2
US 12,511,902 · App. 18/813,620 · Granted Dec 30, 2025

Method and device for providing recipe video

Inventors: Jaewook Shin (Suwon-si, KR); Haedong Yeo (Suwon-si, KR); Junho Rim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06V20/46G06F40/20G06V10/25G06V10/774G06V20/41G06V20/49G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,902
App. No.
18/813,620
Granted
Dec 30, 2025
Kind
B2
Abstract

A device and method for providing a recipe video are provided, the method including obtaining text feature information corresponding to each step of a cooking recipe, the cooking recipe comprising a plurality of steps corresponding to a plurality of coking actions; obtaining image feature information corresponding to each frame of a plurality of frames constituting a cooking video; matching each step of the cooking recipe with a respective section of the cooking video based on a correlation between the obtained text feature information and the obtained image feature information obtained using a machine learned matching model, wherein the respective section of the cooking video corresponds to a respective step among the plurality of steps; and generating the recipe video by using the respective section of the cooking video matched with each step of the cooking recipe.

Claims (54)

1 . A method of providing a recipe video, the method comprising:

obtaining text feature information corresponding to each step of a cooking recipe, the cooking recipe comprising a plurality of steps corresponding to a plurality of coking actions;

obtaining image feature information corresponding to each frame of a plurality of frames constituting a cooking video;

matching each step of the cooking recipe with a respective section of the cooking video based on a correlation between the obtained text feature information and the obtained image feature information obtained using a machine learned matching model, wherein the respective section of the cooking video corresponds to a respective step among the plurality of steps; and

generating the recipe video by using the respective section of the cooking video matched with each step of the cooking recipe.

2 . The method of claim 1 , wherein

the obtaining of the image feature information comprises:

obtaining information about a target of interest comprising a section of interest and a region of interest;

determining, among the plurality of frames, a frame corresponding to the section of interest and a patch within the frame corresponding to the region of interest according to the information about the target of interest; and

obtaining the image feature information corresponding to the determined frame and the patch within the frame from the cooking video.

3 . The method of claim 2 , wherein the obtaining of the information about the target of interest comprises obtaining information about the section of interest based on an operational situation of equipment in a room and information about the region of interest based on a location of the equipment in the room.

4 . The method of claim 2 , wherein the obtaining of the information about the target of interest comprises obtaining information about the section of interest and obtaining information about the region of interest based on tracking of a cooking action and equipment in a room.

5 . The method of claim 2 , wherein the obtaining of the information about the target of interest comprises obtaining information about at least one video of interest among a plurality of cooking videos, and obtaining information about the section of interest and the region of interest for the at least one video of interest.

6 . The method of claim 1 , wherein

the generating of the recipe video comprises:

selecting, based on saliency according to a predetermined criterion, at least one key frame from among frames constituting the respective section of the cooking video matched with each step of the cooking recipe; and

generating the recipe video comprising the selected at least one key frame.

7 . The method of claim 6 , wherein

the selecting of the at least one key frame comprises:

dividing, based on similarity, the frames constituting the respective section of the cooking video matched with each step, into clusters; and

determining the at least one key frame that satisfies a predetermined condition for each cluster of frames, based on saliency estimated according to at least one of a result of scoring an image quality, a result of scoring a degree to which a main action is performed, or a result of scoring a degree of cooking progress.

8 . The method of claim 1 , wherein the machine learned matching model receives the obtained text feature information and the obtained image feature information as inputs and outputs the correlation between the obtained image feature information and the obtained text feature information.

9 . The method of claim 8 , wherein the machine learned matching model is trained based on a dataset including training text feature information and training image feature information to calculate the correlation between the obtained text feature information and the obtained image feature information through a matrix multiplication between the obtained text feature information and the obtained image feature information.

10 . A non-transitory computer-readable medium storing instructions that, when executed by at least one processor of a recipe video device, cause the at least one processor to:

obtain text feature information corresponding to each step of a cooking recipe, the cooking recipe comprising a plurality of steps corresponding to a plurality of coking actions;

obtain image feature information corresponding to each frame of a plurality of frames constituting a cooking video;

match each step of the cooking recipe with a respective section of the cooking video based on a correlation between the obtained text feature information and the obtained image feature information obtained using a machine learned matching model, wherein the respective section of the cooking video corresponds to a respective step among the plurality of steps; and

generate the recipe video by using the respective section of the cooking video matched with each step of the cooking recipe.

11 . A device for providing a recipe video, the device comprising:

a memory storing one or more instructions; and

a processor configured to execute the one or more instructions to:

obtain text feature information corresponding to each step of a cooking recipe, the cooking recipe comprising a plurality of steps corresponding to a plurality of coking actions,

obtain image feature information corresponding to each frame of a plurality of frames constituting a cooking video,

match each step of the cooking recipe with a respective section of the cooking video based on a correlation between the obtained text feature information and the obtained image feature information obtained using a machine learned matching model, wherein the respective section of the cooking video corresponds to a respective step among the plurality of steps, and

generate the recipe video by using the respective section of the cooking video matched with each step of the cooking recipe.

12 . The device of claim 11 , wherein the processor is further configured to execute the one or more instructions to:

obtain information about a target of interest comprising a section of interest and a region of interest,

determine, among the plurality of frames, a frame corresponding to the section of interest and a patch within the frame corresponding to the region of interest according to the information about the target of interest, and

obtain the image feature information corresponding to the determined frame and the patch within the frame from the cooking video.

13 . The device of claim 12 , wherein the processor is further configured to execute the one or more instructions to obtain information about the section of interest based on an operational situation of equipment in a room and information about the region of interest based on a location of the equipment in the room.

14 . The device of claim 12 , wherein the processor is further configured to execute the one or more instructions to obtain information about the section of interest and obtain information about the region of interest based on tracking of a cooking action and equipment in a room.

15 . The device of claim 12 , wherein the processor is further configured to execute the one or more instructions to obtain information about at least one video of interest among a plurality of cooking videos, and obtain information about the section of interest and the region of interest for the at least one video of interest.

16 . The device of claim 11 , wherein the processor is further configured to execute the one or more instructions to:

select, based on saliency according to a predetermined criterion, at least one key frame from among frames constituting the respective section of the cooking video matched with each step of the cooking recipe, and

generate the recipe video comprising the selected at least one key frame.

17 . The device of claim 16 , wherein the processor is further configured to execute the one or more instructions to:

divide, based on similarity, the frames constituting the respective section of the cooking video matched with each step, into clusters, and

determine the at least one key frame that satisfies a predetermined condition for each cluster of frames, based on saliency estimated according to at least one of a result of scoring an image quality, a result of scoring a degree to which a main action is performed, or a result of scoring a degree of cooking progress.

18 . The device of claim 11 , wherein the machine learned matching model receives the obtained text feature information and the obtained image feature information as inputs and outputs the correlation between the obtained image feature information and the obtained text feature information.

19 . The device of claim 18 , wherein the machine learned matching model is trained based on a dataset including training text feature information and training image feature information to calculate the correlation between the obtained text feature information and the obtained image feature information through a matrix multiplication between the obtained text feature information and the obtained image feature information.

20 . The device of claim 19 , further comprising

a communication interface,

wherein the processor is further configured to execute the one or more instructions to,

via the communication interface, receive the cooking video and transmit the generated recipe video to an external device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2024
From: SHIN, JAEWOOK; YEO, HAEDONG; RIM, JUNHO
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 068384/0287 →
Priority Claims (1)
KR 10-2023-0110767 · Aug 23, 2023 · national
Continuity (2)
Continuation PCTKR2024010743 · Jul 24, 2024
Related Publication 20250069395A1 · Feb 27, 2025
References Cited (24)
US 11551575B1 · Knighton et al. · 2023 [cited by applicant]
US 12041384B2 · Tak et al. · 2024 [cited by applicant]
US 20160217329A1 · Kang et al. · 2016 [cited by applicant]
US 20210118447A1 · Kim · 2021 [cited by examiner]
US 20240404284A1 · Hashimoto · 2024 [cited by examiner]
JP 6391078B1 · 2018 [cited by applicant]
JP 6465328B1 · 2019 [cited by applicant]
JP 201936886A · 2019 [cited by applicant]
JP 2019201396A · 2019 [cited by applicant]
JP 202176876A · 2021 [cited by applicant]
KR 1020190043830A · 2019 [cited by applicant]
KR 101968908B1 · 2019 [cited by applicant]
KR 1020190100525A · 2019 [cited by applicant]
KR 1020210046170A · 2021 [cited by applicant]
KR 102313279B1 · 2021 [cited by applicant]
KR 1020230100531A · 2023 [cited by applicant]
Alec Radford et al., “CLIP: Connecting text and images,” OpenAI, Jan. 5, 2021, total 12 pages. [cited by applicant]
Anonymous, “Recipe, A Schema.org Type,” Schema.org, V27.02, Jul. 1, 2024, total 9 pages. [cited by applicant]
Anonymous, “CVPR 2022: TubeR: Tubelet Transformer for Video Action Detection,” CEES Snoek Research on Video and Image AI, Jun. 13, 2022, total 3 pages. [cited by applicant]
Kemal Erdem (burnpiro), “Understanding Region of Interest—(Rol Aligh and Rol Warp),” Towards Data Science, Feb. 10, 2020, total 16 pages. [cited by applicant]
Anonymous, “Object Detection in the Wild via Grounded Language Image Pre-training,” Microsoft, Project Florence-VL, Jun. 17, 2022, total 7 pages. [cited by applicant]
WonJun Moon et al., “Query-Dependent Video Representation for Moment Retrieval and Highlight Detection,” arXiv:2303.13874v1 [cs.CV], Mar. 24, 2023, total 12 pages. [cited by applicant]
Jie Lei et al., “QVHIGHLIGHTS: Detecting Moments and Highlights in Videos via Natural Language Queries,” 35th Conference on Neural Information Processing Systems (NeurIPS 2021), 2021, total 13 pages. [cited by applicant]
International Search Report (PCT/ISA/210) issued by the International Searching Authority on Oct. 24, 2024 in corresponding International Application No. PCT/KR2024/010743. [cited by applicant]