IP Library Granted Patent US 11,176,923
Granted Patent B1
US 11,176,923 · App. 17/222,952 · Granted Nov 16, 2021

System and method for noise cancellation

Inventors: Vlad Vendrow (Reno, NV); Ilya Vladimirovich Mikhailov (Saint-Petersburg, RU)
Assignee: RingCentral, Inc.
G10K11/17823G06F3/165G06K9/00335G06K9/00362G06K9/00711G10K11/17873G10L17/06G10L25/57H04N7/15G10K2210/108G10K2210/3027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,176,923
App. No.
17/222,952
Granted
Nov 16, 2021
Kind
B1
Abstract

A method includes receiving a video data associated with a user in an electronic conference. The method further includes receiving an audio data associated with the user in the electronic conference. It is appreciated that the video data is processed to determine one or more actions taken by the user, and wherein the processing identifies a physical surrounding of the user. The method further includes identifying a portion of the audio data to be suppressed based on the one or more actions taken by the user during the electronic conference and further based on the identification of the physical surrounding of the user.

Claims (45)

1. A method, comprising:

receiving a video data associated with a user in an electronic conference;

receiving an audio data associated with the user in the electronic conference;

processing the video data to determine one or more actions taken by the user, and wherein the processing identifies a physical surrounding of the user; and

identifying a portion of the audio data to be suppressed based on the one or more actions taken by the user during the electronic conference and further based on the identification of the physical surrounding of the user.

2. The method as described in claim 1 , wherein the electronic conference is a video conference, an audio conference, or a webinar.

3. The method as described in claim 1 , further comprising: suppressing the portion of the audio data to be suppressed.

4. The method as described in claim 1 , wherein the identifying is further based on a contextual data associated with the electronic conference.

5. The method as described in claim 4 , wherein the contextual data includes a title of the electronic conference and a content exchanged between users of the electronic conference.

6. The method as described in claim 1 , wherein the one or more actions taken by the user is a direction at which the user is looking, whether the user is typing, or whether lips of the user are moving.

7. The method as described in claim 1 , wherein the one or more actions taken by the user is a pattern of muting/unmuting during the electronic conference.

8. The method as described in claim 1 , wherein the identifying is further based on a direction from which the audio data is received.

9. The method as described in claim 1 , further comprising:

accessing a depository of users voice profile, and wherein the identifying the portion of the audio data is based on the portion of the audio data mismatch with a voice profile associated with the user.

10. The method as described in claim 1 , further comprising:

processing the received audio data using a machine learning algorithm to identify a pattern associated with unwanted audio data; and

identifying the portion of the audio data to be suppressed based on the portion of the audio data matching the pattern associated with unwanted audio data.

11. The method as described in claim 1 , wherein the identifying the portion of the audio data further includes applying a machine learning algorithm to the audio data, wherein the machine learning algorithm is based on one or more actions of the user in one or more electronic conferences other than the electronic conference, wherein applying the machine learning algorithm identifies whether the one or more actions taken by the user during the electronic conference matches a pattern associated with the one or more actions of the user in the one or more electronic conferences other than the electronic conference.

12. The method as described in claim 1 , wherein the identifying is further based on a direction from which a portion of the audio data is received, wherein the portion of the audio is generated by the user.

13. A web-based server for determining a portion of an audio data to be suppressed, comprising:

a memory storing a set of instructions; and

at least one processor configured to execute the instructions to:

facilitate an electronic conference between a first user and a second user;

receive a first audio data from the first user during the electronic conference;

receive a first video data from the first user during the electronic conference;

process the first video data to determine one or more actions taken by the first user;

identify a physical surrounding of the first user; and

identify a portion of the first audio data to be suppressed based on the one or more actions taken by the first user during the electronic conference and further based on the identification of the physical surrounding of the first user.

14. The web-based server as described in claim 13 , wherein the processor is configured to suppress the portion of the first audio data.

15. The web-based server as described in claim 14 , wherein the processor is configured to cause transmission of the first audio data except for the portion of the first audio data to the second user.

16. The web-based server as described in claim 13 , wherein the processor is further configured to identify the portion of the first audio data based on a contextual data associated with the electronic conference, wherein the contextual data includes a title of the electronic conference and a content exchanged between users of the electronic conference.

17. The web-based server as described in claim 13 , wherein the one or more actions taken by the first user is a direction at which the first user is looking, whether the first user is typing, whether lips of the first user are moving, or a pattern of muting/unmuting during the electronic conference.

18. The web-based server as described in claim 13 , wherein the processor is configured to access a voice profile of the first user, and wherein the identifying the portion of the first audio data is based on the portion of the first audio data mismatch the voice profile of the first user.

19. A method, comprising:

facilitating an electronic conference between a first user and a second user;

receiving a first audio data from the first user during the electronic conference;

receiving a first video data from the first user during the electronic conference;

processing the first video data to determine one or more actions taken by the first user;

identifying a physical surrounding of the first user; and

identifying a portion of the first audio data to be suppressed based on the one or more actions taken by the first user during the electronic conference and further based on the identification of the physical surrounding of the first user.

20. The method as described in claim 19 , further comprising: suppressing the portion of the first audio data.

21. The method as described in claim 20 , further comprising: transmitting the first audio data except for the portion of the first audio data to the second user.

22. The method as described in claim 19 , wherein the identifying the portion of the first audio data is further based on a contextual data associated with the electronic conference, wherein the contextual data includes a title of the electronic conference and a content exchanged between users of the electronic conference.

23. The method as described in claim 19 , wherein the one or more actions taken by the first user is a direction at which the first user is looking, whether the first user is typing, whether lips of the first user are moving, or a pattern of muting/unmuting during the electronic conference.

24. The method as described in claim 19 , further comprising: accessing a voice profile of the first user, and wherein the identifying the portion of the first audio data is further based on the portion of the first audio data mismatch the voice profile of the first user.

Assignments (2)
SECURITY INTEREST Recorded Feb 14, 2023
From: RINGCENTRAL, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 062973/0194 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 5, 2021
From: VENDROW, VLAD; MIKHAILOV, ILYA VLADIMIROVICH
To: RINGCENTRAL, INC.
Reel/Frame 055830/0241 →
Priority Claims (1)
WO PCT/RU2020/000790 · Dec 30, 2020 · international
Cited By (2)
US 12,443,389 US 12,658,197