IP Library › Granted Patent US 12,283,278
Granted Patent B2
US 12,283,278 · App. 18/615,621 · Granted Apr 22, 2025

Alphanumeric sequence biasing for automatic speech recognition using a rendered system prompt

Inventors: Benjamin Haynor (New York, NY); Petar Aleksic (Jersey City, NJ)
Assignee: GOOGLE LLC
G10L15/26G10L15/16G10L15/193G10L15/22G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,278
App. No.
18/615,621
Granted
Apr 22, 2025
Kind
B2
Abstract

Speech processing techniques are disclosed that enable determining a text representation of alphanumeric sequences in captured audio data. Various implementations include determining a contextual biasing finite state transducer (FST) based on contextual information corresponding to the captured audio data. Additional or alternative implementations include modifying probabilities of one or more candidate recognitions of the alphanumeric sequence using the contextual biasing FST.

Claims (64)

1. A client device comprising:

one or more microphones;

memory;

one or more processors that execute instructions, stored in the memory, to:

generate a text representation of audio data capturing a spoken utterance, including an alphanumeric sequence, using an automatic speech recognition (“ASR”) engine, wherein in generating the text representation of the audio data capturing the spoken utterance, including the alphanumeric sequence, using the ASR engine, comprises:

determine contextual information for the alphanumeric sequence, wherein in determining the contextual information for the alphanumeric sequence is based on a rendered system prompt that immediately preceded the spoken utterance, and wherein determining the contextual information based on the rendered system prompt one or more of the processors are to:

determine the contextual information based on at least one predicted response to the rendered system prompt;

select, based on the contextual information, one or more contextual finite state transducers for the alphanumeric sequence;

generate a set of candidate recognitions of the spoken utterance based on processing the audio data using an ASR model portion of the ASR engine; and

generate the text representation of the spoken utterance, wherein the text representation includes the alphanumeric sequence, and wherein generating the text representation is based on the generated set of candidate recognitions and the one or more selected contextual finite state transducers.

2. The client device of claim 1 , wherein the ASR model is a recurrent neural network transducer (RNN-T) model, and wherein the ASR engine further comprises a beam search portion.

3. The client device of claim 2 , wherein in generating the text representation of the spoken utterance based on the generated set of candidate recognitions and the one or more selected contextual finite state transducers one or more of the processors are to:

modify the beam search portion of the ASR engine using the one or more contextual finite state transducers.

4. The client device of claim 3 , wherein in generating the text representation of the spoken utterance based on the generated set of candidate recognitions and the one or more selected contextual finite state transducers one or more of the processors are further to:

determine a corresponding probability measure for each candidate recognition in the set of candidate recognitions;

modify the corresponding probability measures using the beam search portion of the ASR engine modified using the one or more contextual finite state transducers;

select a candidate recognition, from the set of candidate recognitions, based on determining that the corresponding probability measure for the candidate recognition satisfies one or more conditions; and

generate the text representation of the spoken utterance based on the selected candidate recognition.

5. The client device of claim 1 , wherein the audio data is captured via the one or more microphones of the client device.

6. The client device of claim 1 , wherein determining the contextual information for the alphanumeric sequence is based on the audio data capturing the spoken utterance, and wherein in determining the contextual information for the alphanumeric sequence based on the audio data one or more of the processors are to:

generate the contextual information for the alphanumeric sequence based on one or more recognized terms of the audio data, the one or more recognized terms being in addition to the alphanumeric sequence.

7. The client device of claim 1 , wherein one or more of the processors are further to generate at least a given contextual finite state transducer, of the one or more contextual finite state transducers, wherein in generating the given contextual finite state transducer one or more of the processors are to:

select an alphanumeric grammar finite state transducer corresponding to an alphanumeric sequence;

select a speller finite state transducer which maps wordpieces to constitute graphemes; and

generate an unweighted wordpiece based acceptor grammar based on the alphanumeric grammar finite state transducer and the speller finite state transducer.

8. The client device of claim 1 , wherein the alphanumeric sequence includes at least one number and includes at least one letter.

9. The client device of claim 1 , wherein the ASR model portion of the ASR engine is an end-to-end speech recognition model.

10. The client device of claim 1 , wherein the ASR engine is trained using a set of training instances, and wherein

the alphanumeric sequence is not in the set of training instances; or

the alphanumeric sequence occurs a number of times, in the set of training instances, that is below a threshold value.

11. A client device comprising:

one or more microphones;

memory;

one or more processors that execute instructions, stored in the memory, to:

generate a text representation of audio data capturing a spoken utterance, including an alphanumeric sequence, using an automatic speech recognition (“ASR”) engine, wherein in generating the text representation of the audio data capturing the spoken utterance, including the alphanumeric sequence, using the ASR engine, comprises:

determine contextual information for the alphanumeric sequence, wherein in determining the contextual information for the alphanumeric sequence is based on a rendered system prompt that immediately preceded the spoken utterance, and wherein determining the contextual information based on the rendered system prompt one or more of the processors are to:

determine the contextual information based on at least one predicted response to the rendered system prompt;

select, based on the contextual information, one or more contextual finite state transducers for the alphanumeric sequence;

generate a set of candidate recognitions of the spoken utterance based on processing the audio data using an ASR model portion of the ASR engine;

and generate the text representation of the spoken utterance, wherein the text representation includes the alphanumeric sequence, and wherein generating the text representation is based on the generated set of candidate recognitions and the one or more selected contextual finite state transducers.

12. The system of claim 11 , wherein the ASR model is a recurrent neural network transducer (RNN-T) model, and wherein the ASR engine further comprises a beam search portion.

13. The system of claim 12 , wherein in generating the text representation of the spoken utterance based on the generated set of candidate recognitions and the one or more selected contextual finite state transducers one or more of the processors are to:

modify the beam search portion of the ASR engine using the one or more contextual finite state transducers.

14. The system of claim 13 , wherein in generating the text representation of the spoken utterance based on the generated set of candidate recognitions and the one or more selected contextual finite state transducers one or more of the processors are further to:

determine a corresponding probability measure for each candidate recognition in the set of candidate recognitions;

modify the corresponding probability measures using the beam search portion of the ASR engine modified using the one or more contextual finite state transducers;

select a candidate recognition, from the set of candidate recognitions, based on determining that the corresponding probability measure for the candidate recognition satisfies one or more conditions; and

generate the text representation of the spoken utterance based on the selected candidate recognition.

15. The system of claim 11 , wherein the audio data is captured via the one or more microphones of the client device.

16. The system of claim 11 , wherein determining the contextual information for the alphanumeric sequence is based on the audio data capturing the spoken utterance, and wherein in determining the contextual information for the alphanumeric sequence based on the audio data one or more of the processors are to:

generate the contextual information for the alphanumeric sequence based on one or more recognized terms of the audio data, the one or more recognized terms being in addition to the alphanumeric sequence.

17. The system of claim 11 , wherein in determining the contextual information for the alphanumeric sequence is based on a rendered system prompt that immediately preceded the spoken utterance, and wherein determining the contextual information based on the rendered system prompt one or more of the processors are to:

determine the contextual information based on at least one predicted response to the rendered system prompt.

18. The system of claim 11 , wherein one or more of the processors are further to generate at least a given contextual finite state transducer, of the one or more contextual finite state transducers, wherein in generating the given contextual finite state transducer one or more of the processors are to:

select an alphanumeric grammar finite state transducer corresponding to an alphanumeric sequence;

select a speller finite state transducer which maps wordpieces to constitute graphemes; and

generate an unweighted wordpiece based acceptor grammar based on the alphanumeric grammar finite state transducer and the speller finite state transducer.

19. A method implemented by one or more processors, the method comprising:

generating a text representation of audio data capturing a spoken utterance, including an alphanumeric sequence, using an automatic speech recognition (“ASR”) engine, wherein in generating the text representation of the audio data capturing the spoken utterance, including the alphanumeric sequence, using the ASR engine, comprises:

determining contextual information for the alphanumeric sequence, wherein in determining the contextual information for the alphanumeric sequence is based on a rendered system prompt that immediately preceded the spoken utterance, and wherein determining the contextual information based on the rendered system prompt one or more of the processors are to:

determine the contextual information based on at least one predicted response to the rendered system prompt;

selecting, based on the contextual information, one or more contextual finite state transducers for the alphanumeric sequence;

generating a set of candidate recognitions of the spoken utterance based on processing the audio data using an ASR model portion of the ASR engine; and

generating the text representation of the spoken utterance, wherein the text representation includes the alphanumeric sequence, and wherein generating the text representation is based on the generated set of candidate recognitions and the one or more selected contextual finite state transducers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2024
From: HAYNOR, BENJAMIN; ALEKSIC, PETAR
To: GOOGLE LLC
Reel/Frame 067159/0191 →
Continuity (2)
Continuation 17251465
Related Publication 20240233732A1 · Jul 11, 2024
References Cited (33)
US 8972243B1 · Strom et al. · 2015 [cited by applicant]
US 9886946B2 · Moreno-Mengibar et al. · 2018 [cited by applicant]
US 9971765B2 · Li et al. · 2018 [cited by applicant]
US 10032451B1 · Mamkina · 2018 [cited by applicant]
US 10140981B1 · Filimonov · 2018 [cited by applicant]
US 10176802B1 · Ladhak et al. · 2019 [cited by applicant]
US 11232799B1 · Birthare · 2022 [cited by examiner]
US 20110153324A1 · Ballinger et al. · 2011 [cited by applicant]
US 20150194149A1 · Faizakof et al. · 2015 [cited by applicant]
US 20200349923A1 · Hu · 2020 [cited by examiner]
US 20200357388A1 · Zhao · 2020 [cited by examiner]
US 20210035566A1 · Ponniah · 2021 [cited by examiner]
US 20220013126A1 · Haynor et al. · 2022 [cited by applicant]
CN 107004407 · 2017 [cited by applicant]
CN 107644638 · 2018 [cited by applicant]
CN 109410949 · 2019 [cited by applicant]
EP 0425291 · 1991 [cited by applicant]
JP 2000267691 · 2000 [cited by applicant]
JP 2010066493 · 2010 [cited by applicant]
JP 2017219769 · 2017 [cited by applicant]
JP 2019020597 · 2019 [cited by applicant]
WO 2021145893 · 2021 [cited by applicant]
European Patent Office; Intention to Grant issued in Application No. 20707891.6; 55 pages; dated May 19, 2023. [cited by applicant]
Intellectual Property India, First Examination Report issued in Application 202227038620; 7 pages; dated Oct. 20, 2022. [cited by applicant]
Serrino, J. et al., “Contextual Recovery of Out-of-Lattice Named Entities in Automatic Speech Recognition;” Interspeech 2019; 5 pages; Sep. 15, 2019. [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT Ser. No. PCT/US2020/014141; 11 pages; dated Oct. 15, 2020. [cited by applicant]
Williams, I. et al., “Contextual Speech Recognition in End-to-End Neural Network Systems Using Beam Search;” Proceedings of Interspeech 2018; 5 pages; Sep. 2, 2018. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued in Application No. 202080093228.3; 15 pages; dated Sep. 19, 2024. [cited by applicant]
Coucke, Alice et al.; Snips Voice Platform: an embedded Spoken Language Understanding system for private-by-design voice interfaces; 29 pages; dated 2018. [cited by applicant]
Zhang, G. et al.; Speech Recognition Decoding Acceleration Method Based on Heterogeneous Computing; Internet New Media Technology, vol. 8/No. 3; 6 pages; dated May 15, 2019. [cited by applicant]
Rao, K. et al., “Exploring Architectures, Data and Units for Streaming End-to-End Speech Recongition with RNN-Transducer,” in Proceedings of IEEE Automatic Speech Recogition and Understanding; pp. 193-199; dated Dec. 20… [cited by applicant]
Gulic, M. et al., “A digital and spelling speech recognition system for the Croatian language”; Proceedings of the 34th International Convention MIPRO; pp. 1673-1678; dated May 2011. [cited by applicant]
Korean Patent Office; Notice of Submission of Opinions issued in Application No. 10-2022-7027865; 16 pages; dated Feb. 27, 2025. [cited by applicant]