IP Library Granted Patent US 12,307,817
Granted Patent B2
US 12,307,817 · App. 17/885,907 · Granted May 20, 2025

Method and system for automatically capturing and processing an image of a user

Inventors: Aditya Kumar (Noida, IN); Natasha Meena (Noida, IN)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06V40/174G06T7/254G06T7/70G06V40/20G06T2207/20224
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,817
App. No.
17/885,907
Granted
May 20, 2025
Kind
B2
Abstract

A method for automatically capturing and processing an image of a user is provided. The method includes determining a level of an emotion identified from a multimedia content; determining an adjusted emotion level by adjusting the level of the emotion based on user information; capturing a plurality of images of the user over a period of time, based on the adjusted emotion level being greater than a first threshold and a reaction of the user determined based on the plurality of images, being less than a second threshold; prioritizing the plurality of images based on at least one of a frame emotion level of the plurality of images, facial features of the user in the plurality of images, or a number of faces in the plurality of images; and processing the prioritized images to generate an output.

Claims (49)

1. A method for automatically capturing and processing an image of a user watching, reading, or listening to a multimedia content, the method comprising:

determining a level of an emotion identified from the multimedia content;

determining an adjusted emotion level of the multimedia content by adjusting the level of the emotion based on user information;

capturing a plurality of images of the user over a period of time from when the adjusted emotion level of the multimedia content is greater than a first threshold to when a reaction of the user determined based on the plurality of images is less than a second threshold;

prioritizing the plurality of images based on at least one of a frame emotion level of the plurality of images, facial features of the user in the plurality of images, or a number of faces in the plurality of images; and

processing the prioritized images to generate an output.

2. The method according to claim 1 , wherein the determining the level of the emotion of the multimedia content comprises:

determining, by using a pre-configured determination model, an emotion probability value based on the emotion identified from the multimedia content; and

determining the level of the emotion based on the emotion probability value.

3. The method according to claim 1 , wherein the determining the adjusted emotion level comprises:

determining an adjustment factor for the emotion based on the user information, wherein the user information comprises at least one of demographic data of the user, a past usage history of the user, or past sensor biological data; and

determining the adjusted emotion level based on the adjustment factor.

4. The method according to claim 1 , wherein the first threshold is determined by retrieving, from an emotion threshold table, a value corresponding to the identified emotion.

5. The method according to claim 1 , wherein the first threshold is a minimum expression intensity for at least one of static and dynamic facial expressions of the emotion to be detected.

6. The method according to claim 1 , wherein the second threshold is determined by:

determining a position of the user in a current captured image and a previous captured image; and

determining a difference in the current captured image and the previous captured image, based on an area corresponding to the position of the user in the current captured image and the previous captured image, and determining a change in the reaction of the user based on the difference, to determine the second threshold.

7. The method according to claim 1 , wherein the prioritizing the plurality of images comprises:

categorizing the plurality of images into at least one set of frames, wherein each set of frames comprises images of the user corresponding to a predefined emotion category;

obtaining, for each of the at least one set of frames, a frame emotion level from images included in each set of frames;

generating a priority value by applying a predetermined function to the obtained frame emotion level in each set of frames, wherein the predetermined function includes weighted summation of the frame emotion level of each set of frames; and

prioritizing the images in each set of frames based on the priority value.

8. The method according to claim 7 , wherein the frame emotion level of a set of frames is obtained by:

determining the frame emotion level based on a weighted sum of emotion values of images included in the set of frames.

9. The method according to claim 7 , wherein the predetermined function further includes determining a facial priority value by:

determining a presence of at least one predefined facial feature in the images of each set of frame, wherein each of the at least one predefined facial feature is assigned a predetermined weight; and

determining the facial priority value based on the presence of the at least one predefined facial feature and the predetermined weight.

10. The method according to claim 1 , wherein the processing of the prioritized images comprises at least one of:

enhancing the prioritized images, wherein enhancing comprises adjusting at least one of blur, lighting, sharpness, or color content in the prioritized images;

generating a new media content from the prioritized images, wherein the new media content comprises at least one of an emoji, an avatar, a video, an audio, or an animated image of the user;

obtaining a health index of the user, wherein the health index indicates a mood of the user; or

editing the multimedia content, wherein editing comprises at least one of zooming, trimming, or animating the multimedia content.

11. The method according to claim 1 , wherein the multimedia content comprises at least one of an audio and video content, a text content, an audio content or a video content.

12. The method according to claim 1 , wherein the multimedia content comprises at least one of a live content or a stored content.

13. An electronic device for automatically capturing and processing an image of a user watching, reading, or listening to a multimedia content, the electronic device comprising at least one processor configured to:

determine a level of an emotion identified from the multimedia content;

determine an adjusted emotion level of the multimedia content by adjusting the level of the emotion based on user information;

capture a plurality of images of the user over a period of time from when the adjusted emotion level of the multimedia content is greater than a first threshold to when a reaction of the user determined based on the plurality of images is less than a second threshold;

prioritize the plurality of images based on at least one of a frame emotion level of the plurality of images, facial features of the user in the plurality of images, or a number of faces in the plurality of images; and

process the prioritized images to generate an output.

14. The electronic device according to claim 13 , wherein the at least one processor is further configured to determine the level of the emotion by:

determining, by using a pre-configured determination model, an emotion probability value based on the emotion identified from the multimedia content; and

determining the level of the emotion based on the emotion probability value.

15. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of an electronic device for automatically capturing and processing an image of a user watching, reading, or listening to a multimedia content, cause the one or more processors to:

determining a level of an emotion identified from the multimedia content;

determining an adjusted emotion level of the multimedia content by adjusting the level of the emotion based on user information;

capturing a plurality of images of the user over a period of time from when the adjusted emotion level of the multimedia content is greater than a first threshold to when a reaction of the user determined based on the plurality of images is less than a second threshold;

prioritizing the plurality of images based on at least one of a frame emotion level of the plurality of images, facial features of the user in the plurality of images, or a number of faces in the plurality of images; and

processing the prioritized images to generate an output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 11, 2022
From: KUMAR, ADITYA; MEENA, NATASHA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061149/0321 →
Priority Claims (1)
IN 202111036374 · Aug 11, 2021 · national
Continuity (2)
Continuation PCTKR2022011301 · Aug 1, 2022
Related Publication 20230066331A1 · Mar 2, 2023
References Cited (48)
US 9159151B2 · Perez et al. · 2015 [cited by applicant]
US 9519989B2 · Perez et al. · 2016 [cited by applicant]
US 9576190B2 · Shaburov · 2017 [cited by examiner]
US 9760767B1 · Bonazzoli et al. · 2017 [cited by applicant]
US 9900498B2 · Kim et al. · 2018 [cited by applicant]
US 10176619B2 · Jiao · 2019 [cited by examiner]
US 10387717B2 · Li et al. · 2019 [cited by applicant]
US 10529379B2 · Chintalapoodi et al. · 2020 [cited by applicant]
US 10755087B2 · Rajvanshi · 2020 [cited by examiner]
US 20040081338A1 · Takenaka · 2004 [cited by examiner]
US 20050223237A1 · Barletta · 2005 [cited by examiner]
US 20160050169A1 · Ben Atar et al. · 2016 [cited by applicant]
US 20160088355A1 · Zubarieva et al. · 2016 [cited by applicant]
US 20160358013A1 · Carter · 2016 [cited by examiner]
US 20170178287A1 · Anderson · 2017 [cited by examiner]
US 20170364484A1 · Hayes · 2017 [cited by applicant]
US 20180082313A1 · Duggin et al. · 2018 [cited by applicant]
US 20180150722A1 · Du et al. · 2018 [cited by applicant]
US 20180330152A1 · Mittelstaedt · 2018 [cited by examiner]
US 20190199663A1 · Liu · 2019 [cited by examiner]
US 20190205626A1 · Kim · 2019 [cited by examiner]
US 20200074156A1 · Janumpally · 2020 [cited by examiner]
US 20200139077A1 · Biradar · 2020 [cited by examiner]
US 20200296480A1 · Chappell, III et al. · 2020 [cited by applicant]
US 20210241444A1 · Cyrus · 2021 [cited by examiner]
CN 104333688B · 2018 [cited by applicant]
EP 3934268A1 · 2022 [cited by applicant]
GB 2585261A · 2021 [cited by applicant]
JP 4891802B2 · 2012 [cited by applicant]
JP 5917841B2 · 2016 [cited by applicant]
KR 101700468B1 · 2017 [cited by applicant]
KR 101704848B1 · 2017 [cited by applicant]
KR 101838792B1 · 2018 [cited by applicant]
KR 1020200078705A · 2020 [cited by applicant]
KR 1020210012528A · 2021 [cited by applicant]
KR 1020210078863A · 2021 [cited by applicant]
KR 1020210089248A · 2021 [cited by applicant]
US 11,388,126 B2, 07/2022, Kennedy (withdrawn) [cited by applicant]
Affective computing: Emotion sensing using 3D images. Madhusudan et al. (Year: 2016). [cited by examiner]
“What is AR Zone on the Galaxy S20”, Samsung, Oct. 15, 2021, 8 pages total. [cited by applicant]
Andre Violante, “Simple Reinforcement Learning: Q-learning”, Towards Data Science, Mar. 19, 2019, 6 pages total. [cited by applicant]
Manuel G. Calvo et al., “Recognition Thresholds for Static and Dynamic Emotional Faces”, Emotion, 2016, vol. 16, No. 8, American Psychological Association, Jun. 30, 2016, 15 pages total. [cited by applicant]
M. Pantic et al., “An Expert System for Multiple Emotional Classification of Facial Expressions”, ResearchGate, IEEE Xplore, DOI: 10.1109/TAI.1999.809775, Feb. 1999, 10 pages total. [cited by applicant]
Shichuan Du et al., “Wait, are you sad or angry? Large exposure time differences required for the categorization of facial expressions of emotion”, Journal of Vision, 13, 4, 13, doi: 10.1167/13.4.13, Mar. 18, 2013, 14 p… [cited by applicant]
International Search Report (PCT/ISA/220 and PCT/ISA/210) and Written Opinion (PCT/ISA/237) issued by the International Searching Authority in International Application No. PC/TKR2022/011301, on Nov. 4, 2022. [cited by applicant]
Extended European Search Report dated Sep. 17, 2024 issued by the European Patent Office in European Application No. 22856083.5. [cited by applicant]
Kaliouby et al., ; “FAIM: Intergrating Automated Facial Affect Analysis in Instant Messaging”, pp. 244-246 (3 pages total) 2004. [cited by applicant]
Teixeira et al., “Determination of emotional content of video clips by low-level audiovisual features ; a dimensional and categorical experimental approach”, Multimed Tools Appl, 2011, vol. 61, pp. 21-49 (29 pages total… [cited by applicant]