IP Library › Granted Patent US 12,658,181
Granted Patent B2
US 12,658,181 · App. 18/285,345 · Granted Jun 16, 2026

Interactive decoding of words from phoneme score distributions

Inventors: Ioannis Alexandros Assael (London, GB); Brendan Shillingford (London, GB); Misha Man Ray Denil (London, GB)
Assignee: GDM Holding LLC
G10L15/16G06V40/20G10L15/183G10L15/187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,658,181
App. No.
18/285,345
Granted
Jun 16, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for interactive decoding of a word sequence.

Claims (66)

1 . A method for performing automatic speech recognition and performed by one or more computers, the method comprising:

processing an input comprising one or more of a video of a speaker speaking an utterance or audio of the utterance using a speech recognition neural network to generate, for each time step in a sequence of time steps of the utterance, a respective phoneme score distribution for the time step that assigns a respective score to each of a plurality of phoneme tokens, the phoneme tokens comprising (i) a plurality of phonemes and (ii) a blank symbol that indicates that no phoneme is spoken at the time step;

generating, using the respective phoneme score distributions, a word sequence of words that represents a decoded word sequence of the utterance, the generating comprising:

initializing fringe data that specifies a plurality of states, wherein each state identifies (i) a candidate sequence of phonemes and (ii) a corresponding candidate sequence of words that is represented by the phonemes; and

generating the word sequence by adding a new word to the word sequence at each of a plurality of update iterations, wherein performing a particular update iteration comprises:

updating the fringe data using the respective phoneme score distributions by, for each state in the fringe data, extending, using the respective phoneme score distributions that are generated by the speech recognition neural network and that each assign a respective score to each of the plurality of phoneme tokens, the candidate sequence of words identified by the state to include a respective additional candidate word after the last word in the word sequence as of the particular update iteration;

providing, for presentation on a user device, one or more of the additional candidate words specified by the states specified in the fringe data;

receiving, from the user device, a user selection of one of the additional candidate words;

responsive to the user selection of the additional candidate word:

updating the word sequence by adding the additional candidate word selected by the user after the last word of the word sequence; and

removing, from the fringe data, any state that does not identify a candidate sequence of words that ends in the additional candidate word selected by the user.

2 . The method of claim 1 , wherein obtaining, for each time step in a sequence of time steps, a respective phoneme score distribution comprises:

processing a video of a speaker using a visual speech recognition neural network that is configured to process the video to generate the respective phoneme score distributions.

3 . The method of claim 1 , wherein obtaining, for each time step in a sequence of time steps, a respective phoneme score distribution comprises:

processing an audio input representing an utterance using an audio speech recognition neural network that is configured to process the audio input to generate the respective phoneme score distributions.

4 . The method of claim 1 , wherein generating the word sequence by adding a new word to the word sequence at each of a plurality of update iterations comprises performing update iterations until a user input is received.

5 . The method of claim 1 , wherein generating the word sequence by adding a new word to the word sequence at each of a plurality of update iterations comprises performing update iterations until the candidate sequences of phonemes identified by all of the states in the fringe data have been generated using all of the respective phoneme distributions at all of the time steps.

6 . The method of claim 1 , wherein updating the fringe data using the respective phoneme score distributions comprises:

after removing from the fringe data any states that do not identify a candidate sequence of words that ends in the last word in the word sequence as of the particular update iteration, repeatedly performing operations until each state specified by the fringe data identifies an additional candidate word after the last word in the word sequence as of the particular update iteration, the operations comprising:

for each particular state in the fringe data that does not yet end in an additional candidate word after the last word in the word sequence as of the particular update iteration:

generating, from the particular state, a plurality of new candidate states; and

generating a respective ranking score for each new candidate state using the respective score distributions; and

updating the fringe data to only specify a predetermined number of states with the highest ranking scores.

7 . The method of claim 6 , wherein providing, for presentation on a user device, the additional candidate words specified by the states specified in the fringe data comprises:

providing the additional candidate words for presentation in an order according to the ranking scores for the corresponding states.

8 . The method of claim 6 , wherein generating, from the particular state, a plurality of new candidate states comprises, for each phoneme token:

generating a new state that identifies a candidate phoneme sequence that includes the phonemes in the candidate sequence identified by the particular state followed by the phoneme token.

9 . The method of claim 8 , wherein each state further identifies (iii) a time step in the sequence of time steps, and wherein generating a respective ranking score for each new candidate state using the respective probability distribution comprises:

generating the respective ranking score for each new state generated from a given existing state from the respective score assigned to the phoneme token that was added to the given existing state at the time step immediately following the time step identified by the given existing state, wherein the new state identifies the time step immediately following the time step identified by the existing state.

10 . The method of claim 9 , wherein the respective ranking score for each new state is further based on a finite state transducer (FST) path weight assigned to the new state by a decoder FST based on a relation between the corresponding candidate phoneme sequence for the new state and the corresponding candidate word sequence for the new state.

11 . One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

processing an input comprising one or more of a video of a speaker speaking an utterance or audio of the utterance using a speech recognition neural network to generate, for each time step in a sequence of time steps of the utterance, a respective phoneme score distribution for the time step that assigns a respective score to each of a plurality of phoneme tokens, the phoneme tokens comprising (i) a plurality of phonemes and (ii) a blank symbol that indicates that no phoneme is spoken at the time step;

generating, using the respective phoneme score distributions, a word sequence of words that represents a decoded word sequence of the utterance, the generating comprising:

initializing fringe data that specifies a plurality of states, wherein each state identifies (i) a candidate sequence of phonemes and (ii) a corresponding candidate sequence of words that is represented by the phonemes; and

generating the word sequence by adding a new word to the word sequence at each of a plurality of update iterations, wherein performing a particular update iteration comprises:

updating the fringe data using the respective phoneme score distributions by, for each state in the fringe data, extending, using the respective phoneme score distributions that are generated by the speech recognition neural network and that each assign a respective score to each of the plurality of phoneme tokens, the candidate sequence of words identified by the state to include a respective additional candidate word after the last word in the word sequence as of the particular update iteration;

providing, for presentation on a user device, one or more of the additional candidate words specified by the states specified in the fringe data;

receiving, from the user device, a user selection of one of the additional candidate words;

responsive to the user selection of the additional candidate word:

updating the word sequence by adding the additional candidate word selected by the user after the last word of the word sequence; and

removing, from the fringe data, any state that does not identify a candidate sequence of words that ends in the additional candidate word selected by the user.

12 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations comprising:

processing an input comprising one or more of a video of a speaker speaking an utterance or audio of the utterance using a speech recognition neural network to generate, for each time step in a sequence of time steps of the utterance, a respective phoneme score distribution for the time step that assigns a respective score to each of a plurality of phoneme tokens, the phoneme tokens comprising (i) a plurality of phonemes and (ii) a blank symbol that indicates that no phoneme is spoken at the time step;

generating, using the respective phoneme score distributions, a word sequence of words that represents a decoded word sequence of the utterance, the generating comprising:

initializing fringe data that specifies a plurality of states, wherein each state identifies (i) a candidate sequence of phonemes and (ii) a corresponding candidate sequence of words that is represented by the phonemes; and

generating the word sequence by adding a new word to the word sequence at each of a plurality of update iterations, wherein performing a particular update iteration comprises:

updating the fringe data using the respective phoneme score distributions by, for each state in the fringe data, extending, using the respective phoneme score distributions that are generated by the speech recognition neural network and that each assign a respective score to each of the plurality of phoneme tokens, the candidate sequence of words identified by the state to include a respective additional candidate word after the last word in the word sequence as of the particular update iteration;

providing, for presentation on a user device, one or more of the additional candidate words specified by the states specified in the fringe data;

receiving, from the user device, a user selection of one of the additional candidate words;

responsive to the user selection of the additional candidate word:

updating the word sequence by adding the additional candidate word selected by the user after the last word of the word sequence; and

removing, from the fringe data, any state that does not identify a candidate sequence of words that ends in the additional candidate word selected by the user.

13 . The system of claim 12 , wherein obtaining, for each time step in a sequence of time steps, a respective phoneme score distribution comprises:

processing a video of a speaker using a visual speech recognition neural network that is configured to process the video to generate the respective phoneme score distributions.

14 . The system of claim 12 , wherein obtaining, for each time step in a sequence of time steps, a respective phoneme score distribution comprises:

processing an audio input representing an utterance using an audio speech recognition neural network that is configured to process the audio input to generate the respective phoneme score distributions.

15 . The system of claim 12 , wherein generating the word sequence by adding a new word to the word sequence at each of a plurality of update iterations comprises performing update iterations until a user input is received.

16 . The system of claim 12 , wherein generating the word sequence by adding a new word to the word sequence at each of a plurality of update iterations comprises performing update iterations until the candidate sequences of phonemes identified by all of the states in the fringe data have been generated using all of the respective phoneme distributions at all of the time steps.

17 . The system of claim 12 , wherein updating the fringe data using the respective phoneme score distributions comprises:

after removing from the fringe data any states that do not identify a candidate sequence of words that ends in the last word in the word sequence as of the particular update iteration, repeatedly performing operations until each state specified by the fringe data identifies an additional candidate word after the last word in the word sequence as of the particular update iteration, the operations comprising:

for each particular state in the fringe data that does not yet end in an additional candidate word after the last word in the word sequence as of the particular update iteration:

generating, from the particular state, a plurality of new candidate states; and

generating a respective ranking score for each new candidate state using the respective score distributions; and

updating the fringe data to only specify a predetermined number of states with the highest ranking scores.

18 . The system of claim 17 , wherein providing, for presentation on a user device, the additional candidate words specified by the states specified in the fringe data comprises:

providing the additional candidate words for presentation in an order according to the ranking scores for the corresponding states.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2023
From: ASSAEL, IOANNIS ALEXANDROS; SHILLINGFORD, BRENDAN; DENIL, MISHA MAN RAY
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 065164/0136 →
Priority Claims (1)
GR 20210100234 · Apr 7, 2021 · national
Continuity (1)
Related Publication 20240185842A1 · Jun 6, 2024
References Cited (31)
US 20020082831A1 · Hwang · 2002 [cited by examiner]
US 20110218802A1 · Bouganim · 2011 [cited by examiner]
US 20180210874A1 · Fuxman · 2018 [cited by examiner]
US 20190139540A1 · Kanda · 2019 [cited by examiner]
US 20200082808A1 · Li · 2020 [cited by examiner]
WO WO2019219968 · 2019 [cited by applicant]
WO WO2019219968A1 · 2019 [cited by examiner]
Ogata, J., & Goto, M. (Sep. 2005). Speech repair: quick error correction just by using selection operation for speech input interfaces. In Interspeech (pp. 133-136). (Year: 2005). [cited by examiner]
Abadi et al., “Tensorflow: A system for large-scale machine learning” in USENIX Symposium on Operating Systems Design and Implementation, Nov. 2016, 265-283. [cited by applicant]
Afouras et al., “Deep audio-visual speech recognition” CoRR, Submitted on Dec. 2018, arXiv:1809.02108v2, 13 pages. [cited by applicant]
Afouras et al., “LRS3-TED: a large-scale dataset for visual speech recognition” CoRR, Submitted on Oct. 2018, arXiv:1809.00496v2, 2 pages. [cited by applicant]
Allauzen et al., “OpenFst: A general and efficient weighted finite-state transducer library” International Conference on Implementation and Application of Automata, Springer, 2007, 11-23. [cited by applicant]
Assael et al., “LipNet: End-to-end sentence-level lipreading” CoRR, Submitted on Dec. 2016, arXiv:1611.01599v2, 13 pages. [cited by applicant]
Chung et al., “Lip reading in the wild” in Asian Conference on Computer Vision, 2016, 17 pages. [cited by applicant]
Fernandez-Lopez et al., “Survey on automatic lip-reading in the era of deep learning” Image and Vision Computing, 2018, 27 pages. [cited by applicant]
Graves et al., “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks” Proceedings of the 23rd international conference on Machine learning, Jun. 2006, 8 pages. [cited by applicant]
Graves et al., “Towards end-to-end speech recognition with recurrent neural networks” in international conference on machine learning, 2014, 9 pages. [cited by applicant]
Harwath et al., “Choosing useful word alternates for automatic speech recognition correction interfaces” Fifteenth Annual Conference of the International Speech Communication Association, 2014, 949-953. [cited by applicant]
Hcupnet.ahrq.gov [online], “Health Care Utilization Project Network, Hospital inpatient national statistics” available on or before May 12, 2019, via Internet Archive: Wayback Machine URL <https://web.archive.org/web/20… [cited by applicant]
Hinton et al., “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups” IEEE Signal Processing Magazine, vol. 29, No. 6, Oct. 2012, 82-97. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/EP2022/059331, mailed on Oct. 19, 2023, 7 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/EP2022/059331, mailed on Aug. 11, 2022, 13 pages. [cited by applicant]
Mohri et al., “Weighted finite-state transducers in speech recognition” Computer Speech & Language 16.1, Jan. 2002, 69-88. [cited by applicant]
Ney et al., “On structuring probabilistic dependences in stochastic language modelling.” Computer Speech & Language 8.1, Jan. 1994, 1-38. [cited by applicant]
Ogata et al., “Speech Repair Quick Error Correction Just By Using Selection Operation for Speech Input Interfaces” Interspeech and Eurospeech, Sep. 2005, 4 pages. [cited by applicant]
Petridis et al., “Deep complementary bottleneck features for visual speech recognition,” In International Conference on Acoustics, Speech, and Signal Processing, IEEE, 2016, 2304-2308. [cited by applicant]
Shillingford et al., “Large-scale visual speech recognition” CoRR, Submitted on Oct. 2018, arXiv:1807.05162v3, 21 pages. [cited by applicant]
Stafylakis et al., “Pushing the boundaries of audiovisual word recognition using residual networks and LSTMs” CoRR, Submitted on Nov. 2018, arXiv:1811.01194v1, 13 pages. [cited by applicant]
Toselli et al., “Chapter 6: Interactive machine translation” Multimodal interactive pattern recognition and applications, 2011, 131-137. [cited by applicant]
Wand et al., “Lipreading with long short-term memory” CoRR, Submitted on Jan. 2016, arXiv:1601.08188v1, 5 pages. [cited by applicant]
Zhou et al., “A review of recent advances in visual speech decoding,” Image and vision computing, vol. 32, No. 9, 2014, 590-605. [cited by applicant]