IP Library › Granted Patent US 12,494,197
Granted Patent B2
US 12,494,197 · App. 18/161,871 · Granted Dec 9, 2025

Speech recognition using word or phoneme time markers based on user input

Inventor: Dongeek Shin (San Jose, CA)
Assignee: Google LLC
G10L15/20G10L15/08G10L15/22G10L25/87G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,494,197
App. No.
18/161,871
Granted
Dec 9, 2025
Kind
B2
Abstract

A method for separating target speech from background noise contained in an input audio signal includes receiving the input audio signal captured by a user device, wherein the input audio signal corresponds to target speech of multiple words spoken by a target user and containing background noise in the presence of the user device while the target user spoke the multiple words in the target speech. The method also includes receiving a sequence of time markers input by the target user in cadence with the target user speaking the multiple words in the target speech, and correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal. The method also includes processing, using a speech recognition model, the enhanced audio features to generate a transcription of the target speech.

Claims (40)

1 . A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving an input audio signal captured by a user device, the input audio signal corresponding to target speech of multiple words spoken by a target user and containing background noise in the presence of the user device while the target user spoke the multiple words in the target speech;

receiving a sequence of time markers input by the target user as the target speaker speaks the target speech and in cadence with the target user speaking the multiple words in the target speech, wherein each time marker in the sequence of time markers is received responsive to the user device detecting, via a sensor, a respective user input indication from the target speaker directed toward the user device;

correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal; and

processing, using a speech recognition model, the enhanced audio features to generate a transcription of the target speech.

2 . The computer-implemented method of claim 1 , wherein correlating the sequence of time markers with the input audio signal comprises:

computing, using the sequence of time markers, a sequence of word time stamps each designating a respective time corresponding to one of the multiple words in the target speech that was spoken by the target user; and

separating, using the sequence of computed word time stamps, the target speech from the background noise in the input audio signal to generate the enhanced audio features.

3 . The computer-implemented method of claim 2 , wherein separating the target speech from the background noise in the input audio signal comprises removing, from inclusion in the enhanced audio features, the background noise.

4 . The computer-implemented method of claim 2 , wherein separating the target speech from the background noise in the input audio signal comprises designating the sequence of word time stamps to corresponding audio segments of the enhanced audio features to differentiate the target speech from the background noise.

5 . The computer-implemented method of claim 1 , wherein receiving the sequence of time markers input by the target user comprises receiving each time marker in the sequence of time markers in response to the target user touching or pressing a predefined region of the user device or another device in communication with the data processing hardware.

6 . The computer-implemented method of claim 5 , wherein the predefined region of the user device or the other device comprises a physical button disposed on the user device or the other device.

7 . The computer-implemented method of claim 5 , wherein the predefined region of the user device or the other device comprises a graphical button displayed on a graphical user interface of the user device.

8 . The computer-implemented method of claim 1 , wherein receiving the sequence of time markers input by the target user comprises receiving each time marker in the sequence of time markers in response to a sensor in communication with the data processing hardware detecting the target user performing a predefined gesture.

9 . The computer-implemented method of claim 1 , wherein a number of time markers in the sequence of time markers input by the user is equal to a number of the multiple words spoken by the target user in the target speech.

10 . The computer-implemented method of claim 1 , wherein the data processing hardware resides on the user device associated with the target user.

11 . The computer-implemented method of claim 1 , wherein the data processing hardware resides on a remote server in communication with the user device associated with the target user.

12 . The computer-implemented method of claim 1 , wherein the background noise in contained in the input audio signal comprises competing speech spoken by one or more other users.

13 . The computer-implemented method of claim 1 , wherein the target speech spoken by the target user comprises a query directed toward a digital assistant executing on the data processing hardware, the query specifying an operation for the digital assistant to perform.

14 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform comprising:

receiving an input audio signal captured by a user device, the input audio signal corresponding to target speech of multiple words spoken by a target user and containing background noise in the presence of the user device while the target user spoke the multiple words in the target speech;

receiving a sequence of time markers input by the target user as the target speaker speaks the target speech and in cadence with the target user speaking the multiple words in the target speech, wherein each time marker in the sequence of time markers is received responsive to the user device detecting, via a sensor, a respective user input indication from the target speaker directed toward the user device;

correlating the sequence of time markers with the input audio signal to generate enhanced audio features that separate the target speech from the background noise in the input audio signal; and

processing, using a speech recognition model, the enhanced audio features to generate a transcription of the target speech.

15 . The system of claim 14 , wherein correlating the sequence of time markers with the input audio signal comprises:

computing, using the sequence of time markers, a sequence of word time stamps each designating a respective time corresponding to one of the multiple words in the target speech that was spoken by the target user; and

separating, using the sequence of computed word time stamps, the target speech from the background noise in the input audio signal to generate the enhanced audio features.

16 . The system of claim 15 , wherein separating the target speech from the background noise in the input audio signal comprises removing, from inclusion in the enhanced audio features, the background noise.

17 . The system of claim 15 , wherein separating the target speech from the background noise in the input audio signal comprises designating the sequence of word time stamps to corresponding audio segments of the enhanced audio features to differentiate the target speech from the background noise.

18 . The system of claim 14 , wherein receiving the sequence of time markers input by the target user comprises receiving each time marker in the sequence of time markers in response to the target user touching or pressing a predefined region of the user device or another device in communication with the data processing hardware.

19 . The system of claim 18 , wherein the predefined region of the user device or the other device comprises a physical button disposed on the user device or the other device.

20 . The system of claim 18 , wherein the predefined region of the user device or the other device comprises a graphical button displayed on a graphical user interface of the user device.

21 . The system of claim 14 , wherein receiving the sequence of time markers input by the target user comprises receiving each time marker in the sequence of time markers in response to a sensor in communication with the data processing hardware detecting the target user performing a predefined gesture.

22 . The system of claim 14 , wherein a number of time markers in the sequence of time markers input by the user is equal to a number of the multiple words spoken by the target user in the target speech.

23 . The system of claim 14 , wherein the data processing hardware resides on the user device associated with the target user.

24 . The system of claim 14 , wherein the data processing hardware resides on a remote server in communication with the user device associated with the target user.

25 . The system of claim 14 , wherein the background noise in contained in the input audio signal comprises competing speech spoken by one or more other users.

26 . The system of claim 14 , wherein the target speech spoken by the target user comprises a query directed toward a digital assistant executing on the data processing hardware, the query specifying an operation for the digital assistant to perform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2023
From: SHIN, DONGEEK
To: GOOGLE LLC
Reel/Frame 062563/0291 →
Continuity (2)
Provisional Application 63267436 · Feb 2, 2022
Related Publication 20230306965A1 · Sep 28, 2023
References Cited (35)
US 8712770B2 · Fukuda · 2014 [cited by examiner]
US 11398230B2 · Jang · 2022 [cited by examiner]
US 12039098B2 · Lee · 2024 [cited by examiner]
US 20040049388A1 · Roth et al. · 2004 [cited by applicant]
US 20100241963A1 · Kulis · 2010 [cited by examiner]
US 20120245936A1 · Treglia · 2012 [cited by examiner]
US 20140012587A1 · Park · 2014 [cited by examiner]
US 20140229167A1 · Wolff · 2014 [cited by examiner]
US 20140278441A1 · Ton et al. · 2014 [cited by applicant]
US 20160275954A1 · Park · 2016 [cited by examiner]
US 20170069321A1 · Toiyama · 2017 [cited by applicant]
US 20190268465A1 · Broidy · 2019 [cited by examiner]
US 20200105264A1 · Jang · 2020 [cited by examiner]
US 20200243072A1 · Park · 2020 [cited by examiner]
US 20210020181A1 · Gorny · 2021 [cited by examiner]
US 20210249027A1 · Zeghidour et al. · 2021 [cited by applicant]
US 20210389868A1 · Crowder · 2021 [cited by examiner]
US 20220293125A1 · Maddika · 2022 [cited by examiner]
US 20230018784A1 · Lee · 2023 [cited by examiner]
US 20230035941A1 · Herman · 2023 [cited by examiner]
US 20230215439A1 · Kanda · 2023 [cited by examiner]
US 20230237266A1 · Hoarau · 2023 [cited by examiner]
US 20230298612A1 · Caroselli · 2023 [cited by examiner]
US 20230368812A1 · Marchi · 2023 [cited by examiner]
US 20240055017A1 · Maddika · 2024 [cited by examiner]
US 20240257836A1 · Sexton · 2024 [cited by examiner]
EP 1739546A2 · 2007 [cited by applicant]
EP 2683147A1 · 2014 [cited by applicant]
JP S6184771A · 1986 [cited by applicant]
JP S62166399A · 1987 [cited by applicant]
JP 2001142485A · 2001 [cited by applicant]
WO 2021004309A1 · 2021 [cited by applicant]
International Search Report and Written Opinion for the related Application No. PCT/US2023/061610, dated May 8, 2023, 93 pages. [cited by applicant]
“A Database for Research on Detection and Enhancement of Speech Transmitted Over Hf Links” Heitkaemper et al. 2021. [cited by applicant]
Japanese Office Action for the related Application No. 2024-545989 dated Sep. 2, 2025. [cited by applicant]