IP Library Granted Patent US 12,418,628
Granted Patent B2
US 12,418,628 · App. 17/977,726 · Granted Sep 16, 2025

Removing undesirable speech from within a conference audio stream

Inventor: Nick Swerdlow (Santa Clara, CA)
Assignee: Zoom Communications, Inc.
H04N7/152G10L15/1815G10L15/22H04L12/1818H04L12/1822H04N7/147
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,418,628
App. No.
17/977,726
Granted
Sep 16, 2025
Kind
B2
Abstract

A portion of speech represented in an audio stream of a conference participant is removed based on a determination that the portion of the speech corresponds to language identified as undesirable. An audio stream representing speech of a user of a participant device connected to a conference is obtained. The audio stream is processed to detect that a portion of the speech corresponds to language identified as undesirable within an audio profile. A modified audio stream is produced by removing the portion of the speech from the audio stream. An output, within the conference, of the modified audio stream is then caused in place of the audio stream.

Claims (45)

1. A method, comprising:

obtaining an audio stream of speech of a user of a first participant device connected to a conference;

detecting that a portion of the speech corresponds to language identified as undesirable within an audio profile, wherein the language identified as undesirable is predefined based on an organization chart of an entity associated with the first participant device;

determining that a second participant device connected to the conference corresponds to a different domain than the first participant device;

producing, based on determining that the second participant device corresponds to the different domain, a modified audio stream by removing the portion of the speech from the audio stream; and

causing an output, within the conference, of the modified audio stream in place of the audio stream.

2. The method of claim 1 , wherein the audio profile corresponds to the user of the first participant device and identifies, as undesirable, a plurality of language spoken by the user of the first participant device in one or more past conferences.

3. The method of claim 1 , wherein the audio profile corresponds to an entity with which the user of the first participant device is associated and identifies, as undesirable, a plurality of language spoken by one or more participant device users associated with the entity in one or more past conferences.

4. The method of claim 1 , wherein the audio profile is configured by the user to identify select language as undesirable.

5. The method of claim 1 , wherein the audio profile is stored at the first participant device.

6. The method of claim 1 , wherein producing the modified audio stream by removing the portion of the speech from the audio stream comprises:

obfuscating the portion of the speech.

7. The method of claim 1 , comprising:

presenting, based on detecting that the portion of the speech corresponds to language identified as undesirable, a prompt to the user of the first participant device recommending to assert a filter against future audio stream content obtained from the first participant device.

8. The method of claim 1 , comprising:

storing, based on a modification of the portion of the speech, configuration data indicating to remove the language identified as undesirable from a future audio stream during one or more future conferences.

9. The method of claim 1 , wherein the portion of the speech includes profanity.

10. The method of claim 1 , wherein the portion of the speech includes a filler word.

11. A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:

obtaining an audio stream of speech of a user of a first participant device connected to a conference;

detecting that a portion of the speech corresponds to language identified as undesirable within an audio profile, wherein the language identified as undesirable is predefined based on an organization chart of an entity associated with the first participant device;

determining that a second participant device connected to the conference corresponds to a different domain than the first participant device;

producing, based on determining that the second participant device corresponds to the different domain, a modified audio stream by removing the portion of the speech from the audio stream; and

causing an output, within the conference, of the modified audio stream in place of the audio stream.

12. The non-transitory computer readable medium of claim 11 , the operations comprising:

applying a filter to automate a removal of further instances of the language identified as undesirable from the audio stream.

13. The non-transitory computer readable medium of claim 11 , wherein the operations for detecting that the portion of the speech corresponds to the language identified as undesirable within the audio profile comprise:

determining a hash value for the portion of the speech; and

comparing the hash value against hash values of the audio profile.

14. The non-transitory computer readable medium of claim 11 , wherein portions of the speech other than the portion of the speech remain within the modified audio stream.

15. An apparatus, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to:

obtain an audio stream of speech of a user of a first participant device connected to a conference;

detect that a portion of the speech corresponds to language identified as undesirable within an audio profile, wherein the language identified as undesirable is predefined based on an organization chart of an entity associated with the first participant device;

determine that a second participant device connected to the conference corresponds to a different domain than the first participant device;

produce, based on determining that the second participant device corresponds to the different domain, a modified audio stream by removing the portion of the speech from the audio stream; and

cause an output, within the conference, of the modified audio stream in place of the audio stream.

16. The apparatus of claim 15 , wherein, to detect that the portion of the speech corresponds to the language identified as undesirable within the audio profile, the processor is configured to execute the instructions to:

obtain, from a client application running at the first participant device, an indication that the portion of the speech corresponds to the language identified as undesirable.

17. The apparatus of claim 15 , wherein the audio profile is stored at the first participant device and updated by an administrator of a user account of the user of the first participant device.

18. The apparatus of claim 15 , wherein the language is identified as undesirable based on past speech from the user of the first participant device.

19. The apparatus of claim 15 , wherein the modified audio stream is output at one or more participant devices associated with the different domain.

20. The apparatus of claim 15 , wherein the processor is configured to execute the instructions to:

notify the user of the first participant device of the removal of the portion of the speech from the audio stream.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2022
From: SWERDLOW, NICK
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 061600/0017 →
Continuity (1)
Related Publication 20240146875A1 · May 2, 2024
References Cited (23)
US 8121845B2 · Kirby · 2012 [cited by applicant]
US 8537978B2 · Jaiswal et al. · 2013 [cited by applicant]
US 9413891B2 · Dwyer et al. · 2016 [cited by applicant]
US 9443518B1 · Gauci · 2016 [cited by applicant]
US 10432687B1 · Hanes et al. · 2019 [cited by applicant]
US 10764534B1 · Shevchenko · 2020 [cited by examiner]
US 11450334B2 · Pichaimurthy et al. · 2022 [cited by applicant]
US 11563855B1 · Spivak et al. · 2023 [cited by applicant]
US 20040263636A1 · Cutler et al. · 2004 [cited by applicant]
US 20070230372A1 · He et al. · 2007 [cited by applicant]
US 20130139259A1 · Tegreene · 2013 [cited by examiner]
US 20130329866A1 · Mai et al. · 2013 [cited by applicant]
US 20140028784A1 · Deyerle et al. · 2014 [cited by applicant]
US 20150012270A1 · Reynolds · 2015 [cited by applicant]
US 20150149173A1 · Korycki · 2015 [cited by applicant]
US 20160063097A1 · Brown · 2016 [cited by examiner]
US 20220051652A1 · Winsvold et al. · 2022 [cited by applicant]
US 20220199102A1 · Ostrand et al. · 2022 [cited by applicant]
US 20230013497A1 · Aher et al. · 2023 [cited by applicant]
US 20230117129A1 · Mouline et al. · 2023 [cited by applicant]
MyFone, 6 Popular Real-Time Voice Changers for Zoom [2022 List], Karen William, Sep. 10, 2021 (Updated Jul. 5, 2022), 8 pages. [cited by applicant]
Voicemod, Voice Changer for Video Calls: Zoom, Hangouts, Facetime, Sep. 2022, 2 pages. [cited by applicant]
Accent Conversion using Pre-trained Model and Synthesized Data from Voice Conversion, Tuan Nam Nguyen, Ngoc Quan Pham, Alexander Waibel, Karlsruhe Institute of Technology and Carnegie Mellon University, Sep. 2022, 5 pag… [cited by applicant]