IP Library Granted Patent US 11,082,465
Granted Patent B1
US 11,082,465 · App. 16/998,979 · Granted Aug 3, 2021

Intelligent detection and automatic correction of erroneous audio settings in a video conference

Inventors: David Chavez (Broomfield, CO); Pushkar Yashavant Deole (Pune, IN); Sandesh Chopdekar (Pune, IN); Navin Daga (Silapathar, IN)
Assignee: Avaya Management L.P.
H04L65/4038G06F3/165G06K9/00315H04N7/152
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,082,465
App. No.
16/998,979
Granted
Aug 3, 2021
Kind
B1
Abstract

Systems, methods, and software to provide intelligent detection and automatic correction of erroneous audio settings in a video conference. Electronic conferences can often be the source of frustration and wasted resources as participants may be forced to contend with extraneous sounds, such as background/ambient noises, or conversations not intended for the conference, provided by an endpoint that should be muted. Similarly, participants may speak with the intention of providing their speech to the conference while their associated endpoint is muted. As a result, the conference may be awkward and lack a productive flow while endpoints are erroneously muted or non-muted. By intelligently processing at least the video portion of a video conference, endpoints/participants may be prompted to mute/unmute or automatically muted/unmuted.

Claims (51)

1. A video conference server, comprising:

a network interface to a network;

a storage component comprising a non-transitory storage device;

a processor, comprising at least one microprocessor; and

wherein the processor, upon accessing machine-executable instructions, is caused to:

broadcast conference content, via the network, to each of a plurality of endpoints, wherein the broadcasted conference content comprises an audio portion and a video portion received from each of the plurality of endpoints;

process at least the video portion from at least one endpoint to determine whether a corresponding audio portion is extraneous to the broadcasted conference content and determine a confidence score associated with the determination whether the corresponding audio portion is extraneous to the broadcasted conference content, wherein the confidence score is based on analysis of the video portion and at least one of natural language processing and/or analysis of the audio portion; and

upon determining that the corresponding audio portion is extraneous to the broadcasted conference content, execute a muting action to exclude the corresponding audio portion from the broadcasted conference content.

2. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

automatically mute an endpoint associated with the corresponding audio portion; and

transmit a message to the automatically muted endpoint indicating that the endpoint was automatically muted.

3. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

signal an endpoint associated with the corresponding audio portion to cause the associated endpoint to prompt a participant to mute their audio.

4. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

automatically mute an endpoint associated with the corresponding audio portion when the confidence score is above a threshold.

5. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

determine that a participant in the at least the video portion is speaking but not looking at their screen.

6. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

determine that a participant in the at least the video portion is not speaking and/or the corresponding audio portion does not comprise speech.

7. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

determine that there is no person in the at least the video portion.

8. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

determine at least one of: a participant's lips are not moving, the participant's other facial parts do not indicate speech, and/or the participant's facial expressions do not indicate speech.

9. The video conference server of claim 1 , wherein additional instructions, when executed further cause the processor to:

process the at least the video portion from the at least one endpoint to determine whether the at least one endpoint is unintentionally muted; and

upon determining that the at least one endpoint may be unintentionally muted, execute signaling to the at least one endpoint that may be unintentionally muted to cause the at least one endpoint to prompt participant associated with the at least one endpoint to unmute their audio.

10. The video conference server of claim 9 , wherein additional instructions, when executed further cause the processor to:

determine that the participant associated with the at least one endpoint is looking at a camera and/or screen, and at least one of: the participant's lips are moving, the participant's other facial parts indicate speech, and/or the participant's facial expressions indicate speech.

11. The video conference server of claim 9 , wherein additional instructions, when executed further cause the processor to:

process at least the audio portion from at least one endpoint to determine a name associated with a particular conference participant was spoken; and

upon determining that the name associated with the particular conference participant was spoken, transmitting to an endpoint associated with the particular conference participant a prompt to unmute their audio.

12. The video conference server of claim 9 , wherein the prompt comprises at least one of: a textual, visual, and/or audible alert.

13. A method of muting an endpoint in a video conference, the method comprising:

broadcasting conference content to each of a plurality of endpoints, wherein the broadcasted conference content comprises an audio portion and a video portion received from each of the plurality of endpoints;

processing at least the video portion from at least one endpoint to determine whether a corresponding audio portion is extraneous to the broadcasted conference content and determine a confidence score associated with the determination whether the corresponding audio portion is extraneous to the broadcasted conference content, wherein the confidence score is based on analysis of the video portion and at least one of natural language processing and/or analysis of the audio portion; and

upon determining that the corresponding audio portion is extraneous to the broadcasted conference content, executing a muting action to exclude the corresponding audio portion from the broadcasted conference content.

14. The method of claim 13 , wherein executing the muting action to exclude the corresponding audio portion from the broadcasted conference content comprises sending a signal to an endpoint associated with the corresponding audio portion to cause the associated endpoint to prompt a participant to mute their audio.

15. The method of claim 13 , wherein executing the muting action to exclude the corresponding audio portion from the broadcasted conference content comprises automatically muting an endpoint associated with the corresponding audio portion when the confidence score is above a threshold.

16. The method of claim 13 , wherein processing the at least the video portion from the at least one endpoint comprises:

determining that a participant in the at least the video portion is speaking but their gaze is not direct to their device.

17. The method of claim 13 , wherein processing the at least the video portion from the at least one endpoint comprises:

determining that a participant in the at least the video portion is not speaking and/or the corresponding audio portion does not comprise speech.

18. A method of unmuting an endpoint in a video conference, the method comprising:

broadcasting conference content to each of a plurality of endpoints, wherein the broadcasted conference content comprises an audio portion and a video portion received from each of the plurality of endpoints;

processing at least the video portion from at least one endpoint to determine whether the at least one endpoint is unintentionally muted and determine a confidence score associated with the determination whether the at least one endpoint is unintentionally muted, wherein the confidence score is based on analysis of the video portion from the at least one endpoint and at least one of natural language processing and/or analysis of the audio portion from the at least one endpoint; and

upon determining that the at least one endpoint may be unintentionally muted, executing signaling to the unintentionally muted at least one endpoint to prompt a participant associated with the unintentionally muted at least one endpoint to unmute their audio.

19. The method of claim 18 , wherein processing the at least the video portion from the at least one endpoint to determine whether the at least one endpoint may be unintentionally muted comprises:

determining that the at least one endpoint is muted, the participant associated with the at least one endpoint is looking at a camera and/or screen, and at least one of: the participant's lips are moving, the participant's other facial parts indicate speech, and/or the participant's facial expressions indicate speech.

20. The method of claim 18 , wherein processing the at least the video portion from the at least one endpoint to determine whether the at least one endpoint may be unintentionally muted further comprises:

processing at least the audio portion from at least one endpoint to determine a name associated with a particular conference participant was spoken; and

upon determining that the name associated with the particular conference participant was spoken, signaling an endpoint associated with the particular conference participant to prompt particular conference participant to unmute their audio.

Assignments (12)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2024
From: AVAYA MANAGEMENT LP
To: ARLINGTON TECHNOLOGIES, LLC
Reel/Frame 066983/0138 →
INTELLECTUAL PROPERTY RELEASE AND REASSIGNMENT Recorded Mar 25, 2024
From: CITIBANK, N.A.
To: AVAYA LLC; AVAYA MANAGEMENT L.P.
Reel/Frame 066894/0117 →
INTELLECTUAL PROPERTY RELEASE AND REASSIGNMENT Recorded Mar 25, 2024
From: WILMINGTON SAVINGS FUND SOCIETY, FSB
To: AVAYA LLC; AVAYA MANAGEMENT L.P.
Reel/Frame 066894/0227 →
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 53955/0436) Recorded May 18, 2023
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: AVAYA MANAGEMENT L.P.; AVAYA INC.; INTELLISIST, INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC
Reel/Frame 063705/0023 →
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 61087/0386) Recorded May 18, 2023
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
To: AVAYA MANAGEMENT L.P.; AVAYA INC.; INTELLISIST, INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC
Reel/Frame 063690/0359 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded May 4, 2023
From: AVAYA INC.; AVAYA MANAGEMENT L.P.; INTELLISIST, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 063542/0662 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded May 3, 2023
From: AVAYA MANAGEMENT L.P.; AVAYA INC.; INTELLISIST, INC.; KNOAHSOFT INC.
To: WILMINGTON SAVINGS FUND SOCIETY, FSB [COLLATERAL AGENT]
Reel/Frame 063742/0001 →
RELEASE OF SECURITY INTEREST IN PATENTS AT REEL 57700/FRAME 0935 Recorded Apr 26, 2023
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: AVAYA HOLDINGS CORP.; AVAYA INC.; AVAYA MANAGEMENT L.P.
Reel/Frame 063458/0303 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Aug 5, 2022
From: AVAYA INC.; INTELLISIST, INC.; AVAYA MANAGEMENT L.P.; AVAYA CABINET SOLUTIONS LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 061087/0386 →
SECURITY INTEREST Recorded Oct 4, 2021
From: AVAYA MANAGEMENT LP
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 057700/0935 →
SECURITY INTEREST Recorded Sep 25, 2020
From: AVAYA INC.; AVAYA MANAGEMENT L.P.; INTELLISIST, INC.; AVAYA INTEGRATED CABINET SOLUTIONS LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 053955/0436 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2020
From: CHAVEZ, DAVID; DEOLE, PUSHKAR YASHAVANT; CHOPDEKAR, SANDESH; DAGA, NAVIN
To: AVAYA MANAGEMENT L.P.
Reel/Frame 053560/0975 →
Cited By (9)
US 12,186,672 US 12,242,364 US 12,287,956 US 12,353,796 US 12,443,388 US 12,477,018 US 12,489,653 US 12,563,159 US 12,684,089