IP Library Granted Patent US 11,790,921
Granted Patent B2
US 11,790,921 · App. 17/169,843 · Granted Oct 17, 2023

Speaker separation based on real-time latent speaker state characterization

Inventors: Valentin Alain Jean Perret (Zurich, CH); Nándor Kedves (Adliswil, CH); Nicolas Lucien Perony (Zurich, CH)
Assignee: OTO Systems Inc.
G10L17/06G06N3/045G06N3/049G06N3/08G10L17/02G10L17/04G10L17/18G10L21/0272
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,790,921
App. No.
17/169,843
Granted
Oct 17, 2023
Kind
B2
Abstract

Systems, methods, and non-transitory computer-readable media can obtain a stream of audio waveform data that represents speech involving a plurality of speakers. As the stream of audio waveform data is obtained, a plurality of audio chunks can be determined. An audio chunk can be associated with one or more identity embeddings. The stream of audio waveform data can be segmented into a plurality of segments based on the plurality of audio chunks and respective identity embeddings associated with the plurality of audio chunks. A segment can be associated with a speaker included in the plurality of speakers. Information describing the plurality of segments associated with the stream of audio waveform data can be provided.

Claims (39)

1. A computer-implemented method comprising:

obtaining, by a computing system, a stream of audio waveform data that represents speech involving a plurality of speakers;

as the stream of audio waveform data is obtained, determining, by the computing system, a plurality of audio chunks, wherein the plurality of audio chunks is associated with one or more identity embeddings;

generating one or more identity-based pretrained markers from the one or more identity embeddings;

segmenting, by the computing system, the stream of audio waveform data into a plurality of segments based on the plurality of audio chunks and the one or more identity-based pretrained markers, wherein each segment of the plurality of segments can be associated with a respective speaker included in the plurality of speakers, the segmenting including determining, based on the one or more identity-based pretrained markers, that a first audio chunk in the plurality of audio chunks matches a speaker in a first state and a second audio chunk of the plurality of audio chunks matches the speaker in a second state; and

providing, by the computing system, information describing the plurality of segments associated with the stream of audio waveform data.

2. The computer-implemented method of claim 1 , wherein the segmenting is performed based on a computational graph.

3. The computer-implemented method of claim 1 , wherein each audio chunk in the plurality of audio chunks corresponds to a fixed length of time.

4. The computer-implemented method of claim 1 , wherein the one or more identity embeddings associated with the audio chunk are generated by a temporal convolutional network that pre-processes the audio chunk and outputs the one or more identity embeddings.

5. The computer-implemented method of claim 1 , wherein segmenting the stream of audio waveform data into the plurality of segments further comprises: assigning, by the computing system, the first audio chunk to the speaker, the speaker being included in a speaker inventory.

6. The computer-implemented method of claim 5 , wherein a temporal convolutional network evaluates at least one identity embedding associated with the first audio chunk and at least one identity embedding associated with the second audio chunk to determine whether the first audio chunk matches the second audio chunk.

7. The computer-implemented method of claim 5 , wherein the speaker inventory maintains associations between speakers identified in the stream of audio waveform data, audio chunks, and identity embeddings.

8. The computer-implemented method of claim 5 , wherein the speaker inventory is refreshed at regular time intervals to reconcile a first speaker in the speaker inventory and a second speaker in the speaker inventory as a same speaker.

9. The computer-implemented method of claim 1 , wherein segmenting the stream of audio waveform data into the plurality of segments further comprises:

determining, by the computing system, that an audio chunk does not match any audio chunks associated with speakers included in a speaker inventory; and

updating, by the computing system, the speaker inventory to include a new speaker associated with the audio chunk.

10. The computer-implemented method of claim 1 , wherein the information describing the plurality of segments provides labels for the plurality of segments, and wherein a label can indicate that a segment represents a particular speaker.

11. A system comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, cause the system to perform operations, the operations comprising:

obtaining a stream of audio waveform data that represents speech involving a plurality of speakers;

as the stream of audio waveform data is obtained, determining a plurality of audio chunks, wherein the plurality of audio chunks is associated with one or more identity embeddings;

generating one or more identity-based pretrained markers from the one or more identity embeddings;

segmenting, by the computing system, the stream of audio waveform data into a plurality of segments based on the plurality of audio chunks and the one or more identity-based pretrained markers, wherein each segment of the plurality of segments can be associated with a respective speaker included in the plurality of speakers, the segmenting including determining, based on the one or more identity-based pretrained markers, that a first audio chunk in the plurality of audio chunks matches a speaker in a first state and a second audio chunk of the plurality of audio chunks matches the speaker in a second state; and

providing information describing the plurality of segments associated with the stream of audio waveform data.

12. The system of claim 11 , wherein the segmenting is performed based on a computational graph.

13. The system of claim 11 , wherein each audio chunk in the plurality of audio chunks corresponds to a fixed length of time.

14. The system of claim 11 , wherein the one or more identity embeddings associated with the audio chunk are generated by a temporal convolutional network that pre-processes the audio chunk and outputs the one or more identity embeddings.

15. The system of claim 11 , wherein segmenting the stream of audio waveform data into the plurality of segments further causes the system to perform: assigning the first audio chunk to the speaker included in the speaker inventory, the speaker being included in a speaker inventory.

16. A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform:

obtaining a stream of audio waveform data that represents speech involving a plurality of speakers;

as the stream of audio waveform data is obtained, determining a plurality of audio chunks, wherein the plurality of audio chunks is associated with one or more identity embeddings;

generating one or more identity-based pretrained markers from the one or more identity embeddings;

segmenting, by the computing system, the stream of audio waveform data into a plurality of segments based on the plurality of audio chunks and the one or more identity-based pretrained markers, wherein each segment of the plurality of segments can be associated with a respective speaker included in the plurality of speakers, the segmenting including determining, based on the one or more identity-based pretrained markers, that a first audio chunk in the plurality of audio chunks matches a speaker in a first state and a second audio chunk of the plurality of audio chunks matches the speaker in a second state; and

providing information describing the plurality of segments associated with the stream of audio waveform data.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the segmenting is performed based on a computational graph.

18. The non-transitory computer-readable storage medium of claim 16 , wherein each audio chunk in the plurality of audio chunks corresponds to a fixed length of time.

19. The non-transitory computer-readable storage medium of claim 16 , wherein the one or more identity embeddings associated with the audio chunk are generated by a temporal convolutional network that pre-processes the audio chunk and outputs the one or more identity embeddings.

20. The non-transitory computer-readable storage medium of claim 16 , wherein segmenting the stream of audio waveform data into the plurality of segments further causes the computing system to perform: assigning the first audio chunk to the speaker included in the speaker inventory, the speaker being included in a speaker inventory.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2023
From: OTO SYSTEMS INC.
To: UNITY TECHNOLOGIES SF
Reel/Frame 065565/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: PERRET, VALENTIN ALAIN JEAN; KEDVES, NÁNDOR; PERONYV, NICOLAS LUCIEN
To: OTO SYSTEMS INC.
Reel/Frame 064980/0739 →
Continuity (3)
Continuation In Part 17115382 · Dec 8, 2020
Provisional Application 63061018 · Aug 4, 2020
Related Publication 20220044687A1 · Feb 10, 2022
Cited By (1)
US 12,315,516