IP Library Granted Patent US 11,405,584
Granted Patent B1
US 11,405,584 · App. 17/212,475 · Granted Aug 2, 2022

Smart audio muting in a videoconferencing system

Inventors: Jonathan Grover (San Jose, CA); Cary Arnold Bran (Vashon, WA)
Assignee: PLANTRONICS, INC.
H04N7/15G06V40/168G10L15/25H04L65/403H04L65/60H04R3/04H04R2430/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,405,584
App. No.
17/212,475
Granted
Aug 2, 2022
Kind
B1
Abstract

A smart muting method for a teleconference or videoconference participant includes detecting audio of an audio-video stream; analyzing video data of the audio-video stream with respect to the detected audio; determining that the audio corresponds to an intended communication, based on the analyzing video data of the audio-video stream with respect to the detected audio; and rendering the audio, responsive to determining that the audio corresponds to an intended communication.

Claims (33)

1. A smart muting method, comprising:

detecting, using a processor, audio of an audio-video stream;

analyzing, using the processor, video data of the audio-video stream with respect to the detected audio;

determining, using the processor, that the audio corresponds to an intended communication, based on the analyzing video data of the audio-video stream with respect to the detected audio; and

rendering, using the processor, the audio, responsive to determining that the audio corresponds to an intended communication,

wherein analyzing video data of the audio-video stream with respect to the detected audio comprises determining a first set of text corresponding to lip movements.

2. The method of claim 1 , wherein detecting audio of the audio-video stream comprises buffering the audio of the audio-video stream.

3. The method of claim 2 , wherein rendering the audio, responsive to determining that the audio corresponds to an intended communication comprises transmitting the buffered audio of the audio-video stream to a remote endpoint.

4. The method of claim 1 , wherein analyzing video data of the audio-video stream with respect to the detected audio further comprises determining a second set of text corresponding to the audio of the audio-video stream.

5. The method of claim 4 , wherein determining that the audio corresponds to an intended communication comprises determining that a degree to which the first set of text and the second set of text correspond exceeds a predetermined threshold.

6. The method of claim 5 , wherein the predetermined threshold corresponds to a degree of seventy-five percent.

7. A teleconferencing system, comprising:

a processor configured to:

detect audio of an audio-video stream;

analyze video data of the audio-video stream with respect to the detected audio;

determine that the audio corresponds to an intended communication, based on the analyzing video data of the audio-video stream with respect to the detected audio;

and render the audio, responsive to determining that the audio corresponds to an intended communication,

wherein analyzing video data of the audio-video stream with respect to the detected audio comprises determining a first set of text corresponding to lip movements.

8. The teleconferencing system of claim 7 , wherein the processor is further configured to detect audio of the audio-video stream by buffering the audio of the audio-video stream.

9. The teleconferencing system of claim 8 , wherein the processor is further configured to render the audio, responsive to determining that the audio corresponds to an intended communication by transmitting the buffered audio of the audio-video stream to a remote endpoint.

10. The teleconferencing system of claim 7 , wherein the processor is further configured to analyze video data of the audio-video stream with respect to the detected audio by determining a second set of text corresponding the audio of the audio-video stream.

11. The teleconferencing system of claim 10 , wherein the processor is further configured to determine that the audio corresponds to an intended communication by determining that a degree to which the first set of text and the second set of text correspond exceeds a predetermined threshold.

12. The teleconferencing system of claim 11 , wherein the predetermined threshold corresponds to a degree of seventy-five percent.

13. A non-transitory computer readable memory storing instructions executable by a processor, wherein the instructions comprise instructions to:

detect audio of an audio-video stream;

analyze video data of the audio-video stream with respect to the detected audio;

determine that the audio corresponds to an intended communication, based on the analyzing video data of the audio-video stream with respect to the detected audio; and

render the audio, responsive to determining that the audio corresponds to an intended communication,

wherein the instructions to analyze video data of the audio-video stream with respect to the detected audio comprise instructions to determine a first set of text corresponding to lip movements.

14. The memory of claim 13 , wherein the instructions to detect audio of the audio-video stream comprise instructions to buffer the audio of the audio-video stream.

15. The memory of claim 14 , wherein the instructions to render the audio, responsive to determining that the audio corresponds to an intended communication comprise instructions to transmit the buffered audio of the audio-video stream to a remote endpoint.

16. The memory of claim 13 , wherein the instructions to analyze video data of the audio-video stream with respect to the detected audio comprise instructions to determine a second set of text corresponding to the audio of the audio-video stream.

17. The memory of claim 16 , wherein the instructions to determine that the audio corresponds to an intended communication comprise instructions to determine that a degree to which the first set of text and the second set of text correspond exceeds a confidence value established responsive detection of one or more user selections.

Assignments (4)
NUNC PRO TUNC ASSIGNMENT Recorded Nov 13, 2023
From: PLANTRONICS, INC.
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 065549/0065 →
RELEASE OF PATENT SECURITY INTERESTS Recorded Aug 30, 2022
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: PLANTRONICS, INC.; POLYCOM, INC.
Reel/Frame 061356/0366 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Oct 6, 2021
From: PLANTRONICS, INC.; POLYCOM, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 057723/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2021
From: GROVER, JONATHAN; BRAN, CARY ARNOLD
To: PLANTRONICS, INC.
Reel/Frame 055719/0711 →
Cited By (2)
US 12,307,012 US 12,386,581