IP Library Granted Patent US 11,410,645
Granted Patent B2
US 11,410,645 · App. 16/348,689 · Granted Aug 9, 2022

Techniques for language independent wake-up word detection

Inventors: Xiao-Lin Ren (Shanghai, CN); Jianzhong Teng (Shanghai, CN)
Assignee: Cerence Operating Company
G10L15/22G10L15/005G10L15/02G10L15/08G10L15/18G10L15/1815G10L2015/025G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,410,645
App. No.
16/348,689
Granted
Aug 9, 2022
Kind
B2
Abstract

A user device configured to perform wake-up word detection in a target language. The user device comprises at least one microphone ( 430 ) configured to obtain acoustic information from the environment of the user device, at least one computer readable medium ( 435 ) storing an acoustic model ( 150 ) trained on a corpus of training data ( 105 ) in a source language different than the target language, and storing a first sequence of speech units obtained by providing acoustic features ( 110 ) derived from audio comprising the user speaking a wake-up word in the target language to the acoustic model ( 150 ), and at least one processor ( 415,425 ) coupled to the at least one computer readable medium ( 435 ) and programmed to perform receiving, from the at least one microphone ( 430 ), acoustic input from the user speaking in the target language while the user device is operating in a low-power mode, applying acoustic features derived from the acoustic input to the acoustic model ( 150 ) to obtain a second sequence of speech units corresponding to the acoustic input, determining if the user spoke the wake-up word at least in part by comparing the first sequence of speech units to the second sequence of speech units, and exiting the low-power mode if it is determined that the user spoke the wake-up word.

Claims (31)

1. A method of enabling wake-up word detection in a target language on a user device, the method comprising:

receiving acoustic input of a user speaking a wake-up word in the target language when the user device is in a low-power mode;

providing acoustic features derived from the acoustic input to an acoustic model stored on the user device to obtain a first sequence of speech units corresponding to the wake-up word spoken by the user in the target language, the acoustic model trained on a corpus of training data in a source language different than the target language;

comparing the first sequence of speech units with a reference sequence of speech units to recognize the wake-up word in the target language, wherein the reference sequence of speech units is obtained by applying acoustic features derived from audio comprising the user speaking the wake-up word in the target language to the acoustic model;

responsive to recognizing the wake-up word, transitioning the user device from the low-power mode to an active mode; and

adapting the acoustic model to the user using both the reference sequence of speech units and the first sequence of speech units.

2. The method of claim 1 , further comprising adapting the acoustic model to the user using the acoustic input of the user speaking the wake-up word in the target language and the sequence of speech units obtained therefrom via the acoustic model.

3. The method of claim 1 , wherein the sequence of speech unit comprises a phoneme sequence corresponding to the wake-up word spoken by the user in the target language.

4. The method of claim 1 , further comprising associating at least one task with the wake-up word to be performed when it is determined that the user has spoken the wake-up word.

5. A user device configured to enable wake-up word detection in a target language, the user device comprising:

at least one microphone configured to obtain acoustic information from an environment of the user device;

at least one computer readable medium storing an acoustic model trained on a corpus of training data in a source language different than the target language; and

at least one processor coupled to the at least one computer readable medium and programmed to perform:

receiving, from the at least one microphone, acoustic input from the user speaking a wake-up word in the target language when the user device is in a low-power mode;

providing acoustic features derived from the acoustic input to the acoustic model to obtain a sequence of speech units corresponding to the wake-up word spoken by the user in the target language;

comparing the sequence of speech units to a reference sequence of speech units, wherein the reference sequence of speech units obtained by applying acoustic features derived from audio comprising the user speaking the wake-up word in the target language to the acoustic model; and

adapting the acoustic model to the user using the sequence of speech units and the reference sequence of the speech units.

6. The user device of claim 5 , further comprising adapting the acoustic model to the user using the acoustic input of the user speaking the wake-up word in the target language and the sequence of speech units obtained therefrom via the acoustic model.

7. The user device of claim 5 , wherein the sequence of speech units comprises a phoneme sequence corresponding to the wake-up word spoken by the user in the target language.

8. The user device of claim 5 , further comprising associating at least one task with the wake-up word to be performed when it is determined that the user has spoken the wake-up word in the target language.

9. A method of performing wake-up word detection on a user device, the method comprising:

while the user device is operating in a low-power mode:

receiving acoustic input from a user speaking in a target language;

providing acoustic features derived from the acoustic input to an acoustic model stored on the user device to obtain a first sequence of speech units corresponding to the acoustic input, the acoustic model trained on a corpus of training data in a source language different than the target language;

determining if the user spoke the wake-up word at least in part by comparing the first sequence of speech units to a reference sequence of speech units stored on the user device, wherein the reference sequence of speech units obtained by applying acoustic features derived from audio comprising the user speaking the wake-up word in the target language to the acoustic model;

exiting the low-power mode if it is determined that the user spoke the wake-up word; and

adapting the acoustic model to the user using both the reference sequence of speech units and the first sequence of speech units.

10. The method of claim 9 , wherein the acoustic model was adapted to the user using the audio comprising the user speaking the wake-up word in the target language and the reference sequence of speech units obtained therefrom via the acoustic model.

11. The method of claim 9 , further comprising performing at least one task associated with the wake-up word if it is determined that the user spoke the wake-up word.

12. The method of claim 9 , wherein the first sequence of speech units comprises a phoneme sequence corresponding to the acoustic input, and wherein the reference sequence of speech units comprises a phoneme sequence corresponding to the user speaking the wake-up word in the target language.

13. The method of claim 9 , wherein exiting the low-power mode comprises transitioning the user device from an idle mode to an active mode.

Assignments (5)
RELEASE (REEL 067417 / FRAME 0303) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0422 →
SECURITY AGREEMENT Recorded Apr 15, 2024
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 067417/0303 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2023
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 064723/0519 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2021
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 055927/0620 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2019
From: REN, XIAO-LIN; TENG, JIANZHONG
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 049661/0149 →