IP Library Granted Patent US 12,379,778
Granted Patent B2
US 12,379,778 · App. 18/627,695 · Granted Aug 5, 2025

Smart windowing for reducing power consumption of a head-mounted camera used for detecting facial expressions

Inventors: Gil Thieberger (Kiryat Tivon, IL); Ari M Frank (Haifa, IL)
Assignee: Facense Ltd.
G06F3/013A61B5/0205A61B5/02427A61B5/02438A61B5/14546A61B5/1455A61B5/6803G06V10/141G06V40/166G06V40/174H04N23/611H04N23/651H04N23/951A61B5/02416A61B5/7221H04N25/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,379,778
App. No.
18/627,695
Granted
Aug 5, 2025
Kind
B2
Abstract

System and method that utilize windowing for efficient capturing of facial landmarks include an inward-facing head-mounted camera that captures images of a region on a user's face utilizing a sensor supporting changing of its region of interest (ROI). The system also includes a computer that detects, based on the images, a type of facial expression expressed by the user, which belongs to a group comprising first and second facial expressions. Responsive to detecting that the user expresses the first facial expression, the computer reads from the camera a first ROI that covers a first subset of facial landmarks relevant to the first facial expression. Responsive to detecting that the user expresses the second facial expression, the computer reads from the camera a second ROI that covers a second subset of facial landmarks relevant to the second facial expression, with the first and second ROIs being different.

Claims (33)

1. A system configured to utilize windowing for efficient capturing of facial landmarks, comprising:

an inward-facing head-mounted camera configured to capture images of a region on a user's face utilizing a sensor that supports changing of its region of interest (ROI); and

a computer configured to:

detect, based on the images, a type of facial expression expressed by the user, which belongs to a group comprising first and second facial expressions;

responsive to detecting that the user expresses the first facial expression, read from the camera a first ROI that covers a first subset of facial landmarks relevant to the first facial expression; and

responsive to detecting that the user expresses the second facial expression, read from the camera a second ROI that covers a second subset of facial landmarks relevant to the second facial expression; wherein the first and second ROIs are different.

2. The system of claim 1 , wherein the computer is further configured to select the first subset as follows: calculate first relevance scores for facial landmarks extracted from a first subset of the images, select a first proper subset of the facial landmarks whose relevance scores reach a first threshold, and set the first ROI to cover the first proper subset of the facial landmarks.

3. The system of claim 2 , wherein the computer is further configured to select the second subset as follows: calculate second relevance scores for facial landmarks extracted from a second subset of the images, select a second proper subset of the facial landmarks whose relevance scores reach a second threshold, and set the second ROI to cover the second proper subset of the facial landmarks.

4. The system of claim 1 , wherein the computer is further configured to select the first and second ROIs based on a pre-calculated function and/or a lookup table that maps between types of facial expressions and their corresponding ROIs.

5. The system of claim 1 , wherein total power consumed from head-mounted components for a process of rendering an avatar based on the first and second ROIs is lower than total power that would have been consumed from the head-mounted components for a process of rendering the avatar based on images of the region.

6. The system of claim 1 , wherein the system further comprises a head-mounted acoustic sensor configured to take audio recordings of the user and a head-mounted movement sensor configured to measure movements of the user's head; and the computer is further configured to (i) generate feature values based on data read from the camera, the audio recordings, and the movements, and (ii) utilize a machine learning-based model to render an avatar of the user based on the feature values.

7. The system of claim 1 , wherein each of the first and second ROIs covers less than half of the region; and wherein the computer is further configured to detect changes in locations of the facial landmarks in the first and second subsets due to facial movements and/or movements of the camera relative to the user's face, and to update each of the first and second ROIs according to the changes.

8. The system of claim 1 , wherein the sensor further supports at least two different binning values for at least two different ROIs, respectively, and the computer is further configured to (i) select, based on performance metrics of facial expression analysis configured to detect the type of facial expression expressed by the user, first and second resolutions for the first and second ROIs, respectively, and (ii) set different binning values for the first and second ROIs according to the first and second resolutions.

9. The system of claim 1 , wherein the sensor further supports changing its binning value, wherein the computer is further configured to calculate relevance scores for facial landmarks extracted from overlapping sub-regions having at least two different binning values; wherein the sub-regions are subsets of the region, and a relevance score per facial landmark at a binning value increases as accuracy of facial expression detection based on the facial landmark at the binning value increases and power consumption used for the facial expression detection decreases; and set the binning values according to a function that optimizes the relevance scores.

10. The system of claim 9 , wherein the computer is configured to increase the relevance scores in proportion to an expected magnitude of movement of the facial landmarks, in order to prefer a higher binning for facial expressions causing larger movements of their respective facial landmarks.

11. The system of claim 1 , wherein the camera is physically coupled to a frame configured to be worn on the user's head, the camera is located less than 15 cm away from the user's face, and the computer is further configured to render an avatar of the user based on data read from the camera.

12. The system of claim 11 , wherein the system is further configured to reduce power consumption of its head-mounted components by checking quality of predictions of locations of facial landmarks using a model, and if the locations of the facial landmarks are closer than a threshold to their expected locations, then a bitrate at which the camera is read is reduced.

13. The system of claim 12 , wherein the computer is further configured to identify that the locations of the facial landmarks are not closer than the threshold to their expected locations, and then increase the bitrate at which the camera is read.

14. A method comprising:

capturing images of a region on a user's face utilizing an inward-facing head-mounted camera comprising a sensor that supports changing of its region of interest (ROI);

detecting, based on the images, a type of facial expression expressed by the user, which belongs to a group comprising first and second facial expressions;

responsive to detecting that the user expresses the first facial expression, reading from the camera a first ROI that covers a first subset of facial landmarks relevant to the first facial expression; and

responsive to detecting that the user expresses the second facial expression, reading from the camera a second ROI that covers a second subset of facial landmarks relevant to the second facial expression; wherein the first and second ROIs are different.

15. The method of claim 14 , further comprising: calculating first relevance scores for facial landmarks extracted from a first subset of the images, selecting a first proper subset of the facial landmarks whose relevance scores reach a first threshold, and setting the first ROI to cover the first proper subset of the facial landmarks.

16. The method of claim 14 , wherein the sensor further supports at least two different binning values for at least two different ROIs, respectively; and further comprising (i) selecting, based on performance metrics of facial expression analysis for detecting the type of facial expression expressed by the user, first and second resolutions for the first and second ROIs, respectively, and (ii) setting different binning values for the first and second ROIs according to the first and second resolutions.

17. The method of claim 14 , further comprising reducing power consumption of head-mounted components related to the camera by checking quality of predictions of locations of facial landmarks using a model, and if the locations of the facial landmarks are closer than a threshold to their expected locations, then reducing bitrate at which the camera is read.

18. A non-transitory computer readable medium storing one or more computer programs configured to cause a processor based system to execute steps comprising:

capturing images of a region on a user's face utilizing an inward-facing head-mounted camera comprising a sensor that supports changing of its region of interest (ROI);

detecting, based on the images, a type of facial expression expressed by the user, which belongs to a group comprising first and second facial expressions;

responsive to detecting that the user expresses the first facial expression, reading from the camera a first ROI that covers a first subset of facial landmarks relevant to the first facial expression; and

responsive to detecting that the user expresses the second facial expression, reading from the camera a second ROI that covers a second subset of facial landmarks relevant to the second facial expression; wherein the first and second ROIs are different.

19. The non-transitory computer readable medium of claim 18 , further comprising: calculating first relevance scores for facial landmarks extracted from a first subset of the images, selecting a first proper subset of the facial landmarks whose relevance scores reach a first threshold, and setting the first ROI to cover the first proper subset of the facial landmarks.

20. The non-transitory computer readable medium of claim 18 , wherein the sensor further supports at least two different binning values for at least two different ROIs, respectively; and further comprising (i) selecting, based on performance metrics of facial expression analysis for detecting the type of facial expression expressed by the user, first and second resolutions for the first and second ROIs, respectively, and (ii) setting different binning values for the first and second ROIs according to the first and second resolutions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: THIEBERGER, GIL, MR.; FRANK, ARI M,, DR.; THIEBERGER-NAVON, TAL, MRS.
To: FACENSE LTD.
Reel/Frame 067725/0524 →
Continuity (6)
Continuation 18105829 · Feb 4, 2023
Continuation 17524411 · Nov 11, 2021
Provisional Application 63140453 · Jan 22, 2021
Provisional Application 63122961 · Dec 9, 2020
Provisional Application 63113846 · Nov 14, 2020
Related Publication 20240256038A1 · Aug 1, 2024
References Cited (39)
US 8270473B2 · Chen et al. · 2012 [cited by applicant]
US 8401248B1 · Moon et al. · 2013 [cited by applicant]
US 9836484B1 · Bialynicka-Birula et al. · 2017 [cited by applicant]
US 20170007165A1 · Jain · 2017 [cited by examiner]
US 20180137678A1 · Kaehler · 2018 [cited by applicant]
JP WO2017006872 · 2017 [cited by applicant]
U.S. Appl. No. 10/074,024, filed Sep. 11, 2018, el Kaliouby et al. [cited by applicant]
U.S. Appl. No. 10/335,045, filed Jul. 2, 2019, Sebe et al. [cited by applicant]
U.S. Appl. No. 10/665,243, filed May 26, 2020, Whitmire et al. [cited by applicant]
Al Chanti, Dawood and Alice Caplier. “Spontaneous facial expression recognition using sparse representation.” arXiv preprint arXiv:1810.00362 (2018). [cited by applicant]
Barrett, Lisa Feldman, et al. “Emotional expressions reconsidered: Challenges to inferring emotion from human facial movements.” Psychological science in the public interest 20.1 (2019): 1-68. [cited by applicant]
Canedo, Daniel, and António JR Neves. “Facial expression recognition using computer vision: a systematic review.” Applied Sciences 9.21 (2019): 4678. [cited by applicant]
Cha, Jaekwang, Jinhyuk Kim, and Shiho Kim. “Hands-free user interface for AR/VR devices exploiting wearer's facial gestures using unsupervised deep learning.” Sensors 19.20 (2019): 4441. [cited by applicant]
Chew, Sien W., et al. “Sparse temporal representations for facial expression recognition.” Pacific-Rim Symposium on Image and Video Technology. Springer, Berlin, Heidelberg, 2011. [cited by applicant]
Corneanu, Ciprian Adrian, et al. “Survey on rgb, 3d, thermal, and multimodal approaches for facial expression recognition: History, trends, and affect-related applications.” IEEE transactions on pattern analysis and mac… [cited by applicant]
Gao, Xinbo, et al. “A review of active appearance models.” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) 40.2 (2010): 145-158. [cited by applicant]
Hickson, Steven, et al. “Eyemotion: Classifying facial expressions in VR using eye-tracking cameras.” 2019 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 2019. [cited by applicant]
Jia, Shan, et al. “Detection of Genuine and Posed Facial Expressions of Emotion: Databases and Methods.” Frontiers in Psychology 11 (2020): 3818. [cited by applicant]
Kerst, Ariane, Jürgen Zielasek, and Wolfgang Gaebel. “Smartphone applications for depression: a systematic literature review and a survey of health care professionals' attitudes towards their use in clinical practice.” … [cited by applicant]
Li, Hao, et al. “Facial performance sensing head-mounted display.” ACM Transactions on Graphics (ToG) 34.4 (2015): 1-9. [cited by applicant]
Li, Shan, and Weihong Deng. “Deep facial expression recognition: A survey.” IEEE transactions on affective computing (2020). [cited by applicant]
Masai, Katsutoshi, et al. “Evaluation of facial expression recognition by a smart eyewear for facial direction changes, repeatability, and positional drift.” ACM Transactions on Interactive Intelligent Systems (TiiS) 7.… [cited by applicant]
Masai, Katsutoshi, Yuta Sugiura, and Maki Sugimoto. “Facerubbing: Input technique by rubbing face using optical sensors on smart eyewear for facial expression recognition.” Proceedings of the 9th Augmented Human Interna… [cited by applicant]
Masai, Katsutoshi. “Facial Expression Classification Using Photo-reflective Sensors on Smart Eyewear.”, PhD Thesis, (2018). [cited by applicant]
Milborrow, Stephen, and Fred Nicolls. “Locating facial features with an extended active shape model.” European conference on computer vision. Springer, Berlin, Heidelberg, 2008. [cited by applicant]
Nakamura, Fumihiko, et al. “Automatic Labeling of Training Data by Vowel Recognition for Mouth Shape Recognition with Optical Sensors Embedded in Head-Mounted Display.” ICAT-EGVE. 2019. [cited by applicant]
Nakamura, Hiromi, and Homei Miyashita. “Control of augmented reality information volume by glabellar fader.” Proceedings of the 1st Augmented Human international Conference. 2010. [cited by applicant]
Perusquía-Hernández, Monica. “Are people happy when they smile?: Affective assessments based on automatic smile genuineness identification.” Emotion Studies 6.1 (2021): 57-71. [cited by applicant]
Ringeval, Fabien, et al. “AVEC 2019 workshop and challenge: state-of-mind, detecting depression with AI, and cross-cultural affect recognition.” Proceedings of the 9th International on Audio/Visual Emotion Challenge and… [cited by applicant]
Si, Jiaxin, Yanchen Wan, and Teng Zhang. “Facial expression recognition based on ASM and Finite Automata Machine.” (2016): 1306-1310. Proceedings of the 2016 2nd Workshop on Advanced Research and Technology in Industry … [cited by applicant]
Sugiura, Yuta, et al. “Behind the palm: Hand gesture recognition through measuring skin deformation on back of hand by using optical sensors.” 2017 56th Annual Conference of the Society of Instrument and Control Enginee… [cited by applicant]
Suk, Myunghoon, and Balakrishnan Prabhakaran. “Real-time facial expression recognition on smartphones.” 2015 IEEE Winter Conference on Applications of Computer Vision. IEEE, 2015. [cited by applicant]
Suzuki, Katsuhiro, et al. “Recognition and mapping of facial expressions to avatar by embedded photo reflective sensors in head mounted display.” 2017 IEEE Virtual Reality (VR). IEEE, 2017. [cited by applicant]
Terracciano, Antonio, et al. “Personality predictors of longevity: activity, emotional stability, and conscientiousness.” Psychosomatic medicine 70.6 (2008): 621. [cited by applicant]
Umezawa, Akino, et al. “e2-MaskZ: a Mask-type Display with Facial Expression Identification using Embedded Photo Reflective Sensors.” Proceedings of the Augmented Humans International Conference. 2020. [cited by applicant]
Valstar, Michel, et al. “Avec 2016: Depression, mood, and emotion recognition workshop and challenge.” Proceedings of the 6th international workshop on audio/visual emotion challenge. 2016. [cited by applicant]
Yamashita, Koki, et al. “CheekInput: turning your cheek into an input surface by embedded optical sensors on a head-mounted display.” Proceedings of the 23rd ACM Symposium on Virtual Reality Software and Technology. 201… [cited by applicant]
Yamashita, Koki, et al. “DecoTouch: Turning the Forehead as Input Surface for Head Mounted Display.” International Conference on Entertainment Computing. Springer, Cham, 2017. [cited by applicant]
Zeng, Nianyin, et al. “Facial expression recognition via learning deep sparse autoencoders.” Neurocomputing 273 (2018): 643-649. [cited by applicant]