IP Library Granted Patent US 12,205,196
Granted Patent B2
US 12,205,196 · App. 17/847,689 · Granted Jan 21, 2025

Method and apparatus for video conferencing

Inventors: Sachi Mizobuchi (Markham, CA); Xuan Zhou (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06T11/00G06V20/40G06V40/174G06V40/20H04L65/403
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,196
App. No.
17/847,689
Granted
Jan 21, 2025
Kind
B2
Abstract

The present disclosure relates to a video conferencing system which provides non-obtrusive feedback to a presenter during a video conference to improve interactivity among participants of video conference. An image sensor, such as a camera, captures a sequence of images (i.e., a video) of the participant while a participant of the video conference is moving their port or a part of their body and sends the images to video conferencing server. The video conferencing server processes the images to recognize a type of gesture performed by the participant and selects an ambient graphic that corresponds to the recognized gesture. The video conferencing server sends the ambient graphic to a client device associated with the presenter. The client device associated with the presenter renders or displays the ambient graphic on a display screen of the client device without obscuring information displayed on the display screen of the client device.

Claims (36)

1. A video conferencing server comprising:

a processor, and

a non-transitory storage medium storing instructions to recognize facial expressions and body movements, the instructions being executable by the processor such that the video conferencing server is configured to:

receive participant video information from a participant client device;

generate a bounding box for each participant detected in the received participant video information;

for each bounding box corresponding to a detected participant, perform object recognition on the bounding box to recognize a facial expression and a body movement of the detected participant;

determine a number of similar facial expressions and/or similar body movements amongst the detected participants;

provide a semi-transparent ambient graphic based on attributes of the recognized facial expression and body movement for each detected participant and the number of similar facial expressions and/or similar body movements amongst the detected participants; and

transmit the ambient graphic to a presenter display associated with a presenter client device for rendering the ambient graphic configured to be overlaid over at least a portion of a current content being displayed on the presenter display without obscuring the current displayed content.

2. The video conferencing server of claim 1 , further configured to:

transmit the ambient graphic to a participant display associated with the participant client device for displaying the ambient graphic over at least a portion of a current content being displayed on the participant display.

3. The video conferencing server of claim 1 , further configured to:

receive participant video information from a second participant client device;

generate a bounding box for each participant detected in the received second participant video information;

for each bounding box corresponding to a detected participant of the second participant video information, perform object recognition on the bounding box to recognize a facial expression and a body movement of each detected participant of the second participant video information

provide a semi-transparent ambient graphic representative of attributes of the recognized facial expression and/or body movement for each detected participant; and

transmit the ambient graphic to the presenter display associated with the presenter client device for rendering the ambient graphic configured to be overlaid over at least a portion of the current content being displayed on the presenter display without obscuring the current content.

4. The video conferencing server of claim 3 , further configured to:

transmit the ambient graphic to a participant display associated with the second participant client device for displaying the ambient graphic over at least a portion of a current content being displayed on the participant display associated with the second participant client device.

5. The video conferencing server of claim 1 , wherein, the participant information includes audio and video information associated with at least one participant, the video information is captured by an image sensor and the audio information is captured by a microphone.

6. The video conferencing server of claim 1 , wherein the facial expression includes at least one of a laughing, smiling or nodding and the body movement includes at least one of a head nodding, head tilting, raising hands, waving hands, pointing hands or applauding.

7. The video conferencing server of claim 1 , wherein providing the ambient graphic further includes generating the ambient graphic.

8. The video conferencing server of claim 1 , wherein the current content includes at least one of a digital document stored on the presenter client device or a digital document accessed online.

9. A video conferencing method for a server to recognize facial expressions and body movements, the method comprising: receiving participant video information from a participant client device; generating a bounding box for each participant detected in the received participant video information; for each bounding box corresponding to a detected participant, performing object recognition on the bounding box to recognize a facial expression and body movement of each detected participant; determining a number of similar facial expressions and/or similar body movements amongst the detected participants; providing a semi-transparent ambient graphic representative of attributes of the recognized facial expression and/or body movement for each detected participant and the number of similar facial expressions and/or similar body movements amongst the detected participants; and transmitting the ambient graphic to a presenter display associated with a presenter client device for rendering the ambient graphic configured to be overlaid over at least a portion of a current content being displayed on the presenter display without obscuring the current displayed content.

10. The method of claim 9 further comprising transmitting the ambient graphic to a participant display associated with the participant client device for displaying the ambient graphic over at least a portion of a current content being displayed on the participant display associated with the participant client device.

11. The method of claim 9 further comprising:

receiving participant video information from a second participant client device;

generating a bounding box for each participant detected in the received second participant video information;

for each bounding box corresponding to a detected participant of the second participant video information, performing object recognition on the bounding box to recognize a facial expression and/or body movement of each detected participant of the second participant video information;

providing a semi-transparent ambient graphic representative of attributes of the recognized facial expression and/or body movement for each detected participant; and

transmitting the ambient graphic to the presenter display associated with the presenter client device for rendering the ambient graphic configured to be overlaid over at least a portion of the current content being displayed on the presenter display associated with the presenter client device without obscuring the current content.

12. The method of claim 11 further comprising transmitting the ambient graphic to a participant display associated with the second participant client device for displaying the ambient graphic over at least a portion of a current content being displayed on the participant display associated with the second participant client device.

13. The method of claim 9 , wherein the participant information includes audio and video information associated with at least one participant and the video information is captured by an image sensor and the audio information is captured by a microphone.

14. The method of claim 9 , wherein the facial expression includes at least one of a laughing, smiling or nodding and the body movement includes at least one of a head nodding, head tilting, raising hands, waving hands, pointing hands or applauding.

15. The method of claim 9 , wherein the providing of the ambient graphic further includes generating the ambient graphic.

16. The method of claim 9 , wherein the current content includes at least one of a digital document stored on the presenter client device or a digital document accessed online.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2022
From: MIZOBUCHI, SACHI; ZHOU, XUAN
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 060365/0440 →
Continuity (2)
Continuation PCTCN2020102492 · Jul 16, 2020
Related Publication 20220319063A1 · Oct 6, 2022
References Cited (18)
US 20100207874A1 · Yuxin · 2010 [cited by examiner]
US 20110029893A1 · Roberts · 2011 [cited by examiner]
US 20130019187A1 · Hind · 2013 [cited by examiner]
US 20140270388A1 · Lucey · 2014 [cited by examiner]
US 20170310927A1 · West · 2017 [cited by examiner]
US 20180063482A1 · Goesnar · 2018 [cited by examiner]
US 20190373216A1 · Cutler et al. · 2019 [cited by applicant]
US 20200099890A1 · Tanaka · 2020 [cited by examiner]
US 20200184203A1 · Anders · 2020 [cited by examiner]
US 20200312331A1 · Gustafson · 2020 [cited by examiner]
US 20200349429A1 · Vendrow · 2020 [cited by examiner]
CN 103607556A · 2014 [cited by applicant]
CN 104349111A · 2015 [cited by applicant]
CN 106548517A · 2017 [cited by applicant]
CN 108881784A · 2018 [cited by applicant]
WO 2018214746A1 · 2018 [cited by applicant]
International Search Report of PCT/CN2020/102492; ISA/CN; Jia Meng; Apr. 19, 2021. [cited by applicant]
Shami et al., Enhancing distributed corporate meetings with lightweight avatars, Apr. 14-15, 2010, In CHI'10 Extended Abstracts on Human Factors in Computing Systems, Atlanta, USA, pp. 3829-3834. [cited by applicant]