IP Library Granted Patent US 12,380,910
Granted Patent B2
US 12,380,910 · App. 17/364,583 · Granted Aug 5, 2025

Systems and methods for virtual meeting speaker separation

Inventors: Prashant Kukde (Milpitas, CA); Sushant Shivram Hiray (Maharashtra, IN)
Assignee: RingCentral, Inc.
G10L21/0272G10L25/51G10L25/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,910
App. No.
17/364,583
Granted
Aug 5, 2025
Kind
B2
Abstract

A computer-implemented machine learning method for improving speaker separation is provided. The method comprises processing audio data to generate prepared audio data and determining feature data and speaker data from the prepared audio data through a clustering iteration to generate an audio file. The method further comprises re-segmenting the audio file to generate a speaker segment and causing to display the speaker segment through a client device.

Claims (41)

1. A computer-implemented machine learning method for improving speaker separation, the method comprising:

processing audio data to generate prepared audio data, wherein processing the audio data comprises eliminating one or more non-speech audio segments;

determining feature data and speaker data from the prepared audio data through a clustering iteration to generate an audio file;

determining a plurality of environments associated with the audio data based on the feature data, wherein the feature data includes background sounds used to distinguish environments of the plurality of environments, and wherein the clustering iteration comprises sequentially performing:

applying a divisive hierarchical clustering to divide out the audio file based on the determined plurality of environments to form a plurality of audio files, wherein each audio file of the plurality of audio files is associated with an environment of the plurality of environments;

subsequent to forming the plurality of audio files, applying agglomerative hierarchical clustering to divide out speakers and obtain the speaker data within each respective audio file of the plurality of audio files;

subsequent to performing the clustering iteration, re-segmenting the plurality of audio files to generate a speaker segment based on the agglomerative hierarchical clustering applied to each audio file of the plurality of audio files; and

causing to display the speaker segment through a client device.

2. The computer-implemented machine learning method of claim 1 , wherein determining feature data comprises determining a gender using machine learning.

3. The computer-implemented machine learning method of claim 1 , wherein determining speaker data comprises extracting speaker data from the determined environment using machine learning.

4. The computer-implemented machine learning method of claim 1 , wherein processing the audio data comprises generating normalized audio data.

5. The computer-implemented machine learning method of claim 1 , wherein processing the audio data comprises eliminating one or more non-speech audio segments before processing remaining audio data to generate prepared audio data.

6. The computer-implemented machine learning method of claim 1 , wherein processing the audio data comprises separating an overlapping speech segment.

7. A non-transitory, computer-readable medium storing a set of instructions that, when executed by a processor, cause:

processing audio data to generate prepared audio data, wherein processing the audio data comprises eliminating one or more non-speech audio segments;

determining feature data and speaker data from the prepared audio data through a clustering iteration to generate an audio file;

determining a plurality of environments associated with the audio data based on the feature data, wherein the feature data includes background sounds used to distinguish environments of the plurality of environments, and wherein the clustering iteration comprises sequentially performing:

applying a divisive hierarchical clustering to divide out the audio file based on the determined plurality of environments to form a plurality of audio files, wherein each audio file of the plurality of audio files is associated with an environment of the plurality of environments;

subsequent to forming the plurality of audio files, applying agglomerative hierarchical clustering to divide out speakers and obtain the speaker data within each respective audio file of the plurality of audio files;

subsequent to performing the clustering iteration, re-segmenting the plurality of audio files to generate a speaker segment based on the agglomerative hierarchical clustering applied to each audio file of the plurality of audio files; and

causing to display the speaker segment through a client device.

8. The non-transitory, computer-readable medium of claim 7 , wherein determining feature data comprises determining a gender using machine learning.

9. The non-transitory, computer-readable medium of claim 7 , wherein determining speaker data comprises extracting speaker data from the determined environment using machine learning.

10. The non-transitory, computer-readable medium of claim 7 , wherein processing the audio data comprises generating normalized audio data.

11. The non-transitory, computer-readable medium of claim 7 , wherein processing the audio data comprises eliminating one or more non-speech audio segments before processing remaining audio data to generate prepared audio data.

12. The non-transitory, computer-readable medium of claim 7 , wherein processing the audio data comprises separating an overlapping speech segment.

13. A machine learning system for improving speaker separation, the system comprising:

a processor;

a memory storing instructions that, when executed by the processor, cause:

processing audio data to generate prepared audio data, wherein processing the audio data comprises eliminating one or more non-speech audio segments;

determining feature data and speaker data from the prepared audio data through a clustering iteration to generate an audio file;

determining a plurality of environments associated with the audio data based on the feature data, wherein the feature data includes background sounds used to distinguish environments of the plurality of environments, and wherein the clustering iteration comprises sequentially performing:

applying a divisive hierarchical clustering to divide out the audio file based on the determined plurality of environments to form a plurality of audio files, wherein each audio file of the plurality of audio files is associated with an environment of the plurality of environments;

subsequent to forming the plurality of audio files, applying agglomerative hierarchical clustering to divide out speakers and obtain the speaker data within each respective audio file of the plurality of audio files;

subsequent to performing the clustering iteration, re-segmenting the plurality of audio files to generate a speaker segment based on the agglomerative hierarchical clustering applied to each audio file of the plurality of audio files; and

causing to display the speaker segment through a client device.

14. The machine learning system of claim 13 , wherein determining feature data comprises determining a gender using machine learning.

15. The machine learning system of claim 13 , wherein determining speaker data comprises extracting speaker data from the determined environment using machine learning.

16. The machine learning system of claim 13 , wherein processing the audio data comprises generating normalized audio data.

17. The machine learning system of claim 13 , wherein processing the audio data comprises eliminating one or more non-speech audio segments before processing remaining audio data to generate prepared audio data.

18. The machine learning system of claim 13 , wherein processing the audio data comprises separating an overlapping speech segment.

Assignments (3)
SECURITY INTEREST Recorded Nov 12, 2025
From: RINGCENTRAL, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 072876/0523 →
SECURITY INTEREST Recorded Feb 14, 2023
From: RINGCENTRAL, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 062973/0194 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2021
From: KUKDE, PRASHANT; HIRAY, SUSHANT SHIVRAM
To: RINGCENTRAL, INC.
Reel/Frame 056725/0222 →