IP Library Granted Patent US 12,074,720
Granted Patent B2
US 12,074,720 · App. 17/732,826 · Granted Aug 27, 2024

Automated language identification during virtual conferences

Inventors: Awni Yusuf Hannun (Los Altos, CA); Sebastian Stüker (Karlsruhe, DE)
Assignee: Zoom Video Communications, Inc.
H04L12/1818G06F40/58G10L15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,074,720
App. No.
17/732,826
Filed
Apr 29, 2022
Granted
Aug 27, 2024
Kind
B2
Art Unit
2451
USPC
709/204
Abstract

In some aspects, a computing device may access audio information comprising an audio stream from a client device. The computing device may provide an audio segment from the audio stream to a language identification process of the computing device comprising a machine learning model that is trained to identify a language of a plurality of languages within recorded speech. The computing device may identify an identified-language of the plurality of languages for the speech based at least in part on the audio segment. The computing device may provide the identified-language to the client device. Numerous other aspects are described.

Claims (45)

1. A computer-implemented method, comprising:

accessing, by a computing device of a video conference provider system, audio information comprising an audio stream from a client device;

providing, by the computing device, a first audio segment from the audio stream to a language identification process of the computing device comprising a machine learning model that is trained to identify a language of a plurality of languages within recorded speech, wherein the plurality of languages comprises an unidentified language indicator and one or more languages, wherein the language identification process assigns a first confidence score to the first audio segment;

identifying, by the language identification process of the computing device, a first identified-language corresponding to the first audio segment based at least in part on the first confidence score exceeding a confidence threshold;

initiating, by the computing device, a change timer in response to identifying the first identified-language, wherein the confidence threshold is an increased confidence threshold until a conclusion of the change timer; and

before the conclusion of the change timer:

providing, by the computing device, a second audio segment from the audio stream to the language identification process of the computing device, wherein the language identification process assigns a second confidence score to the second audio segment; and

identifying, by the language identification process of the computing device, a second identified-language corresponding to the second audio segment based at least in part on the second confidence score exceeding the increased confidence threshold, wherein the first identified-language and the second identified-language are different languages.

2. The method of claim 1 , wherein the machine learning model generates a plurality of confidence score for the audio segment, wherein a corresponding confidence score is assigned for each language of the plurality of languages.

3. The method of claim 2 , wherein the plurality of confidence scores for the audio segment sum to one.

4. The method of claim 1 , wherein the providing further comprises:

sending, by the computing device, a notification containing the identified-language to the client device.

5. The method of claim 1 , further comprising:

providing the audio information to a selected transcription process, where the transcription process is selected, from a plurality of transcription processes, based at least in part on the identified-language.

6. A non-transitory computer-readable medium storing a set of instructions, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a computing device of a video conference provider system, cause the computing device to:

access audio information comprising an audio stream from a client device;

provide a first audio segment from the audio stream to a language identification process of the computing device comprising a machine learning model that is trained to identify a language of a plurality of languages within recorded speech, wherein the plurality of languages comprises an unidentified language indicator and one or more languages, wherein the language identification process assigns a first confidence score to the first audio segment;

identify a first identified-language corresponding to the first audio segment based at least in part on the first confidence score exceeding a confidence threshold;

initiate a change timer in response to identifying the first identified-language, wherein the confidence threshold is an increased confidence threshold until a conclusion of the change timer; and

before the conclusion of the change timer:

provide a second audio segment from the audio stream to the language identification process, wherein the language identification process assigns a second confidence score to the second audio segment; and

identify, by the language identification process, a second identified-language corresponding to the second audio segment based at least in part on the second confidence score exceeding the increased confidence threshold, wherein the first identified-language and the second identified-language are different languages.

7. The non-transitory computer-readable medium of claim 6 , wherein the machine learning model generates a plurality of confidence score for the audio segment, wherein a corresponding confidence score is assigned for each language of the plurality of languages.

8. The non-transitory computer-readable medium of claim 7 , wherein the plurality of confidence scores for the audio segment sum to one.

9. The non-transitory computer-readable medium of claim 6 , wherein the providing further comprises:

send a notification containing the identified-language to the client device.

10. The non-transitory computer-readable medium of claim 6 , wherein the one or more instructions further cause the computing device to:

provide the audio information to a selected transcription process, where the transcription process is selected, from a plurality of transcription processes, based at least in part on the identified-language.

11. A computing device, comprising:

one or more memories; and

one or more processors, communicatively coupled to the one or more memories, configured to:

access audio information comprising an audio stream from a client device;

provide a first audio segment from the audio stream to a language identification process of a computing device of a video conference provider system comprising a machine learning model that is trained to identify a language of a plurality of languages within recorded speech, wherein the plurality of languages comprises an unidentified language indicator and one or more languages, wherein the language identification process assigns a first confidence score to the first audio segment;

identify a first identified-language corresponding to the first audio segment based at least in part on the first confidence score exceeding a confidence threshold;

initiate a change timer in response to identifying the first identified-language, wherein the confidence threshold is an increased confidence threshold until a conclusion of the change timer; and

before the conclusion of the change timer:

providing a second audio segment from the audio stream to the language identification process of the computing device, wherein the language identification process assigns a second confidence score to the second audio segment; and

identify, by the language identification process, a second identified-language corresponding to the second audio segment based at least in part on the second confidence score exceeding the increased confidence threshold, wherein the first identified-language and the second identified-language are different languages.

12. The computing device of claim 11 , wherein the machine learning model generates a plurality of confidence score for the audio segment, wherein a corresponding confidence score is assigned for each language of the plurality of languages.

13. The computing device of claim 12 , wherein the plurality of confidence scores for the audio segment sum to one.

14. The computing device of claim 11 , wherein the providing further comprises:

send a notification containing the identified-language to the client device.

15. The computing device of claim 11 , wherein the one or more processors are further configured to:

provide the audio information to a selected transcription process, where the transcription process is selected, from a plurality of transcription processes, based at least in part on the identified-language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2022
From: HANNUN, AWNI; STUEKER, SEBASTIAN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 059796/0689 →
Continuity (1)
Related Publication 20230353399A1 · Nov 2, 2023