IP Library Granted Patent US 12,315,516
Granted Patent B2
US 12,315,516 · App. 18/368,459 · Granted May 27, 2025

Speaker separation based on real-time latent speaker state characterization

Inventors: Valentin Alain Jean Perret (Zurich, CH); Nándor Kedves (Adliswil, CH); Nicolas Lucien Perony (Zurich, CH)
Assignee: Unity Technologies SF
G10L17/06G06N3/045G06N3/049G06N3/08G10L17/02G10L17/04G10L17/18G10L21/0272
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,516
App. No.
18/368,459
Granted
May 27, 2025
Kind
B2
Abstract

Systems, methods, and non-transitory computer-readable media can obtain a stream of audio waveform data that represents speech involving a plurality of speakers. As the stream of audio waveform data is obtained, a plurality of audio chunks can be determined. An audio chunk can be associated with one or more identity embeddings. The stream of audio waveform data can be segmented into a plurality of segments based on the plurality of audio chunks and respective identity embeddings associated with the plurality of audio chunks. A segment can be associated with a speaker included in the plurality of speakers. Information describing the plurality of segments associated with the stream of audio waveform data can be provided.

Claims (32)

1. A method comprising:

obtaining a stream of audio waveform data that represents speech involving a plurality of speakers;

determining a plurality of audio chunks, wherein the plurality of audio chunks is associated with one or more identity embeddings;

generating one or more identity-based pretrained markers from the one or more identity embeddings; and

segmenting the stream of audio waveform data into a plurality of segments, the segmenting including assigning a first speaker of the plurality of speakers to a first audio chunk of the plurality of audio chunks or adding a new speaker to the plurality of speakers to assign to the first audio chunk, wherein the assigning is based on a degree of matching of an identity-based pretrained marker of the one or more identity-based pretrained markers to a previously-generated identity-based pretrained marker, the previously-generated identity-based pretrained marker being stored in a speaker inventory.

2. The method of claim 1 , wherein the segmenting is performed based on a computational graph.

3. The method of claim 1 , wherein the one or more identity-based pretrained markers include speaker state information, and wherein the degree of matching is based in part on the speaker state information.

4. The method of claim 1 , wherein the one or more identity embeddings are generated by a temporal convolutional network that pre-processes the plurality of audio chunks and outputs the one or more identity embeddings.

5. The method of claim 1 , wherein the degree of matching is based on output from a temporal convolutional network that takes the identity-based pretrained marker of the one or more identity-based pretrained markers and the previously-generated identity-based pretrained marker as inputs.

6. The method of claim 1 , wherein the speaker inventory maintains associations between the plurality of speakers, the plurality of audio chunks, and the one or more identity embeddings.

7. The method of claim 1 , wherein the speaker inventory is refreshed at regular time intervals to reconcile a first speaker in the speaker inventory and a second speaker in the speaker inventory as a same speaker.

8. The method of claim 1 , wherein the adding of the new speaker is based on a determination that the identity-based pretrained marker of the one or more identity-based pretrained markers does not match any of a plurality of previously-generated identity-based pretrained markers.

9. The method of claim 2 , further comprising providing information describing the plurality of segments associated with the stream of audio waveform data.

10. The method of claim 9 , wherein the information includes a label for a segment of the plurality of segments indicating that the segment represents a particular speaker of the plurality of speakers.

11. A system comprising:

obtaining a stream of audio waveform data that represents speech involving a plurality of speakers;

determining a plurality of audio chunks, wherein the plurality of audio chunks is associated with one or more identity embeddings;

generating one or more identity-based pretrained markers from the one or more identity embeddings; and

segmenting the stream of audio waveform data into a plurality of segments, the segmenting including assigning a first speaker of the plurality of speakers to a first audio chunk of the plurality of audio chunks or adding a new speaker to the plurality of speakers to assign to the first audio chunk, wherein the assigning is based on a degree of matching of an identity-based pretrained marker of the one or more identity-based pretrained markers to a previously-generated identity-based pretrained marker, the previously-generated identity-based pretrained marker being stored in a speaker inventory.

12. The system of claim 11 , wherein the one or more identity-based pretrained markers include speaker state information, and wherein the degree of matching is based in part on the speaker state information.

13. The system of claim 11 , wherein the speaker inventory maintains associations between the plurality of speakers, the plurality of audio chunks, and the one or more identity embeddings.

14. The system of claim 11 , wherein the speaker inventory is refreshed at regular time intervals to reconcile a first speaker in the speaker inventory and a second speaker in the speaker inventory as a same speaker.

15. The system of claim 11 , wherein the adding of the new speaker is based a determination that the identity-based pretrained marker of the one or more identity-based pretrained markers does not match any of a plurality of previously-generated identity-based pretrained markers.

16. A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform:

obtaining a stream of audio waveform data that represents speech involving a plurality of speakers;

determining a plurality of audio chunks, wherein the plurality of audio chunks is associated with one or more identity embeddings;

generating one or more identity-based pretrained markers from the one or more identity embeddings; and

segmenting the stream of audio waveform data into a plurality of segments, the segmenting including assigning a first speaker of the plurality of speakers to a first audio chunk of the plurality of audio chunks or adding a new speaker to the plurality of speakers to assign to the first audio chunk, wherein the assigning is based on a degree of matching of an identity-based pretrained marker of the one or more identity-based pretrained markers to a previously-generated identity-based pretrained marker, the previously-generated identity-based pretrained marker being stored in a speaker inventory.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the one or more identity-based pretrained markers include speaker state information, and wherein the degree of matching is based in part on the speaker state information.

18. The non-transitory computer-readable storage medium of claim 16 , wherein the speaker inventory maintains associations between the plurality of speakers, the plurality of audio chunks, and the one or more identity embeddings.

19. The non-transitory computer-readable storage medium of claim 16 , wherein the speaker inventory is refreshed at regular time intervals to reconcile a first speaker in the speaker inventory and a second speaker in the speaker inventory as a same speaker.

20. The non-transitory computer-readable storage medium of claim 16 , wherein the adding of the new speaker is based a determination that the identity-based pretrained marker of the one or more identity-based pretrained markers does not match any of a plurality of previously-generated identity-based pretrained markers.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2024
From: OTO SYSTEMS INC.
To: UNITY TECHNOLOGIES SF
Reel/Frame 068832/0433 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2023
From: OTO SYSTEMS INC.
To: UNITY TECHNOLOGIES SF
Reel/Frame 065565/0285 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: PERRET, VALENTIN ALAIN JEAN; KEDVES, NÁNDOR; PERONY, NICOLAS LUCIEN
To: OTO SYSTEMS INC.
Reel/Frame 064980/0888 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: PERRET, VALENTIN ALAIN JEAN; KEDVES, NÁNDOR; PERONY, NICOLAS LUCIEN
To: OTO SYSTEMS INC.
Reel/Frame 064987/0778 →
Continuity (4)
Continuation 17169843 · Feb 8, 2021
Continuation In Part 17115382 · Dec 8, 2020
Provisional Application 63061018 · Aug 4, 2020
Related Publication 20240153509A1 · May 9, 2024
References Cited (41)
US 10706857B1 · Ramasubramanian et al. · 2020 [cited by applicant]
US 11646037B2 · Perret · 2023 [cited by examiner]
US 11790921B2 · Perret · 2023 [cited by examiner]
US 20080040110A1 · Pereg · 2008 [cited by examiner]
US 20180358003A1 · Calle et al. · 2018 [cited by applicant]
US 20200211544A1 · Mikhailov · 2020 [cited by examiner]
US 20200219517A1 · Wang · 2020 [cited by examiner]
US 20200301959A1 · Anorga et al. · 2020 [cited by applicant]
US 20200349921A1 · Jansen et al. · 2020 [cited by applicant]
US 20210020161A1 · Gao · 2021 [cited by applicant]
US 20210272571A1 · Balasubramaniam et al. · 2021 [cited by applicant]
US 20220044687A1 · Perret et al. · 2022 [cited by applicant]
US 20220044688A1 · Perret et al. · 2022 [cited by applicant]
US 20220115020A1 · Bradley et al. · 2022 [cited by applicant]
US 20220188577A1 · Chopde et al. · 2022 [cited by applicant]
US 20220199094A1 · El Shafey et al. · 2022 [cited by applicant]
US 20220206698A1 · Jang · 2022 [cited by applicant]
US 20230352031A1 · Perret et al. · 2023 [cited by applicant]
US 20240153509A1 · Perret · 2024 [cited by examiner]
WO WO2019084419A1 · 2019 [cited by applicant]
WO WO2019132690A1 · 2019 [cited by examiner]
“U.S. Appl. No. 17/115,382, Examiner Interview Summary mailed Nov. 8, 2022”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/115,382, Non Final Office Action mailed Jun. 28, 2022”, 10 pgs. [cited by applicant]
“U.S. Appl. No. 17/115,382, Notice of Allowance mailed Dec. 23, 2022”, 10 pgs. [cited by applicant]
“U.S. Appl. No. 17/115,382, Response filed Nov. 28, 2022 to Non Final Office Action mailed Jun. 28, 2022”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/115,382, Supplemental Notice of Allowability mailed Apr. 4, 2023”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 17/115,382, Supplemental Notice of Allowability mailed Apr. 7, 2023”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 17/169,843, Examiner Interview Summary mailed Nov. 8, 2022”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/169,843, Final Office Action mailed Feb. 28, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/169,843, Non Final Office Action mailed Jul. 8, 2022”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/169,843, Notice of Allowance mailed Jun. 14, 2023”, 10 pgs. [cited by applicant]
“U.S. Appl. No. 17/169,843, Response filed Jan. 9, 2023 to Non Final Office Action mailed Jul. 8, 2022”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/169,843, Response filed May 30, 2023 to Final Office Action mailed Feb. 28, 2023”, 12 pgs. [cited by applicant]
Alzantot, Moustafa, et al., “Deep residual neural networks for audio spoofing detection”, arXiv preprint arXiv:1907.00501, (2019), 5 pgs. [cited by applicant]
Hadsell, Raia, et al., “Dimensionality Reduction by Learning an Invariant Mapping”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06) vol. 2, (2006), 1735-1742. [cited by applicant]
Qian, Qi, et al., “SoftTriple Loss: Deep Metric Learning Without Triplet Sampling”, Proceedings of the IEEE International Conference on Computer Vision, (2019), 6450-6458. [cited by applicant]
Sohn, Kihyuk, et al., “Improved Deep Metric Learning with Multi-class N-pair Loss Objective”, 30th Conference on Neural Information Processing Systems (NIPS 2016), (2016), 1857-1865. [cited by applicant]
Song, Hyun Oh, et al., “Deep Metric Learning Via Lifted Structured Feature Embedding”, Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR), (Jun. 27-30, 2016), 4004-4012. [cited by applicant]
Turpault, Nicolas, et al., “Semi-supervised triplet loss based learning of ambient audio embeddings”, ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, (2019), 5 p… [cited by applicant]
Wang, Jian, et al., “Deep Metric Learning with Angular Loss”, Proceedings of the IEEE International Conference on Computer Vision, (2017), 4321-4329. [cited by applicant]
Zhang, Chunlei, et al., “Text-independent speaker verification based on triplet convolutional neural network embeddings”, IEEE/ACM Transactions on Audio, Speech, and Language Processing 26.9, (2018), 1633-1644. [cited by applicant]