IP Library Granted Patent US 9,430,695
Granted Patent B2
US 9,430,695 · App. 14/740,498 · Granted Aug 30, 2016

Determining which participant is speaking in a videoconference

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,430,695
App. No.
14/740,498
Granted
Aug 30, 2016
Kind
B2
Abstract

Aspects herein describe methods and systems of receiving, by one or more cameras, images in which the images comprise facial images of individuals. Aspects of the disclosure describe extracting the facial images from the images received, sorting the extracted facial images into separate groups wherein each group corresponds to the facial images of each individual, and selecting, for each individual, a preferred facial image from each group. The preferred facial images selected are transmitted to a client for display. Aspects of the disclosure also describe selecting either a facial recognition algorithm or an audio triangulation algorithm to use to determine which individual is speaking wherein the selection is based on whether lip movement of one or more of the individuals is visible in the images received from the cameras.

Claims (60)

1. A system comprising:

one or more cameras; and

a device configured to be connected to the one or more cameras, the device comprising one or more processors and memory storing instructions that, when executed by one of the processors, cause the device to:

obtain, from at least one of the cameras, a set of images of a plurality of individuals at a location,

select, from the set of images, for each individual, a preferred facial image for the individual,

determine whether lip movement of one of the individuals is visible in the set of images, and

select, based on whether lip movement of one of the individuals is visible in the set of images, at least one of a facial recognition algorithm and an audio triangulation algorithm to determine which individual is speaking;

wherein the instructions, when executed by one of the processors, further cause the device to provide a graphical window which is rendered on a display, the graphical window including multiple viewports, the graphical window being divided into a grid pattern, each viewport of the multiple viewports residing in a respective location of the grid pattern and including a facial image for an individual, the graphical window further including viewport highlighting which highlights viewports corresponding to individuals that are speaking.

2. The system of claim 1 , wherein:

the viewport highlighting comprises flashing of one or more viewport borders.

3. The system of claim 1 , wherein:

the viewport highlighting comprises shading a background of one or more viewport portions.

4. The system of claim 1 , wherein:

the instructions, when executed by one of the processors, cause the device to:

determine that multiple individuals are speaking simultaneously, and

cause multiple viewports that present facial images of the multiple individuals to provide indications that the multiple individuals are speaking.

5. The system of claim 1 , wherein:

determining whether lip movement of one of the individuals is visible in the set of images comprises determining whether a lip area of the individual is visible in the set of images.

6. The system of claim 1 , wherein:

the set of images comprises video of the individuals participating in a videoconference at the location.

7. A device comprising:

one or more processors; and

memory storing instructions that, when executed by one of the processors, cause the device to:

obtain, from at least one camera, a set of images of a plurality of individuals at a location,

select, from the set of images, for each individual, a preferred facial image for the individual,

determine whether lip movement of one of the individuals is visible in the set of images, and

select, based on whether lip movement of one of the individuals is visible in the set of images, at least one of a facial recognition algorithm and an audio triangulation algorithm to determine which individual is speaking;

wherein the instructions, when executed by one of the processors, further cause the device to provide a graphical window which is rendered on a display, the graphical window including multiple viewports, the graphical window being divided into a grid pattern, each viewport of the multiple viewports residing in a respective location of the grid pattern and including a facial image for an individual, the graphical window further including viewport highlighting which highlights viewports corresponding to individuals that are speaking.

8. The device of claim 7 , wherein:

the viewport highlighting comprises flashing of one or more viewport borders.

9. The device of claim 7 , wherein:

the viewport highlighting comprises shading a background of one or more viewport portions.

10. The device of claim 7 , wherein:

the instructions, when executed by one of the processors, cause the device to:

determine that multiple individuals are speaking simultaneously, and

cause multiple viewports that present facial images of the multiple individuals to provide indications that the multiple individuals are speaking.

11. The device of claim 7 , wherein:

determining whether lip movement of one of the individuals is visible in the set of images comprises determining whether a lip area of the individual is visible in the set of images.

12. The device of claim 7 , wherein:

the set of images comprises video of the individuals participating in a videoconference at the location.

13. A non-transitory computer-readable medium storing instructions, that when executed by a processor of a device, cause the device to:

obtain, from at least one camera, video of a plurality of individuals participating in a video conference,

select, from the video, for each individual, a preferred facial image for the individual,

determine whether lip movement of one of the individuals is visible in the video, and

select, based on whether lip movement of one of the individuals is visible in the video, at least one of a facial recognition algorithm and an audio triangulation algorithm to determine which individual is speaking;

wherein the instructions, when executed, further cause the device to provide a graphical window which is rendered on a display, the graphical window including multiple viewports, the graphical window being divided into a grid pattern, each viewport of the multiple viewports residing in a respective location of the grid pattern and including a facial image for an individual, the graphical window further including viewport highlighting which highlights viewports corresponding to individuals that are speaking.

14. The computer-readable medium of claim 13 , wherein:

the viewport highlighting comprises shading a background of one or more viewport portions.

15. The computer-readable medium of claim 13 , wherein:

determining whether lip movement of one of the individuals is visible in the video comprises determining whether a lip area of the individual is visible in the video.

16. A method of managing a videoconference which includes multiple participants, the method comprising:

for each participant of the videoconference, selecting a preferred video feed that includes that participant from a respective set of video feeds that include that participant;

for each participant of the videoconference, extracting a series of facial images of that participant from the selected preferred video feed that includes that participant; and

provide a graphical window which is rendered on a display, the graphical window including multiple viewports, the graphical window being divided into a grid pattern, each viewport of the multiple viewports residing in a respective location of the grid pattern and including one of the series of facial images extracted for a particular participant, the graphical window further including viewport highlighting which highlights viewports corresponding to participants that are speaking.

17. A method as in claim 16 , further comprising:

identifying exactly one participant that is speaking based on (i) detected lip movement and (ii) audio triangulation; and

wherein the viewport highlighting highlights exactly one viewport which corresponds to the exactly one participant that is speaking.

18. A method as in claim 16 , further comprising:

based on (i) detected lip movement and (ii) audio triangulation, identifying a first participant and a second participant that are simultaneously speaking; and

wherein the viewport highlighting concurrently highlights (i) a first viewport which corresponds to the first participant and (ii) a second viewport which corresponds to the second participant, the first viewport and the second viewport residing in different locations of the grid pattern.

Assignments (12)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 053667/0169, REEL/FRAME 060450/0171, REEL/FRAME 063341/0051) Recorded Mar 15, 2024
From: BARCLAYS BANK PLC, AS COLLATERAL AGENT
To: GOTO GROUP, INC. (F/K/A LOGMEIN, INC.)
Reel/Frame 066800/0145 →
SECURITY INTEREST Recorded Feb 16, 2024
From: GOTO COMMUNICATIONS, INC.,; GOTO GROUP, INC., A; LASTPASS US LP,
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS THE NOTES COLLATERAL AGENT
Reel/Frame 066614/0402 →
SECURITY INTEREST Recorded Feb 16, 2024
From: GOTO COMMUNICATIONS, INC.; GOTO GROUP, INC.; LASTPASS US LP
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS THE NOTES COLLATERAL AGENT
Reel/Frame 066614/0355 →
SECURITY INTEREST Recorded Feb 7, 2024
From: GOTO GROUP, INC.,; GOTO COMMUNICATIONS, INC.; LASTPASS US LP
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 066508/0443 →
CHANGE OF NAME Recorded Apr 8, 2022
From: LOGMEIN, INC.
To: GOTO GROUP, INC.
Reel/Frame 059644/0090 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS (SECOND LIEN) Recorded Feb 16, 2021
From: BARCLAYS BANK PLC, AS COLLATERAL AGENT
To: LOGMEIN, INC.
Reel/Frame 055306/0200 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Sep 1, 2020
From: LOGMEIN, INC.
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 053667/0079 →
NOTES LIEN PATENT SECURITY AGREEMENT Recorded Sep 1, 2020
From: LOGMEIN, INC.
To: U.S. BANK NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 053667/0032 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Sep 1, 2020
From: LOGMEIN, INC.
To: BARCLAYS BANK PLC, AS COLLATERAL AGENT
Reel/Frame 053667/0169 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2019
From: GETGO, INC.
To: LOGMEIN, INC.
Reel/Frame 050732/0725 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2016
From: CITRIX SYSTEMS, INC.
To: GETGO, INC.
Reel/Frame 039970/0670 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2016
From: SUMMERS, JACOB JARED
To: CITRIX SYSTEMS, INC.
Reel/Frame 037931/0779 →