IP Library › Granted Patent US 12,380,882
Granted Patent B2
US 12,380,882 · App. 16/906,775 · Granted Aug 5, 2025

Speech processing system and method

Inventors: Thomas William John Ash (Cambridge, GB); Anthony John Robinson (Cambridge, GB)
Assignee: THE CHANCELLOR, MASTERS, AND SCHOLARS OF THE UNIVERSITY OF CAMBRIDGE
G10L15/193G09B19/06G10L15/187G10L25/51G10L25/87G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,882
App. No.
16/906,775
Granted
Aug 5, 2025
Kind
B2
Abstract

A speech processing system includes an input for receiving an input utterance spoken by a user and a word alignment unit configured to align different sequences of acoustic speech models with the input utterance spoken by the user. Each different sequence of acoustic speech models corresponds to a different possible utterance that a user might make. The system identifies any parts of a read prompt text that the user skipped; any parts of the read prompt text that the user repeated; and any speech sounds that the user inserted between words of the read prompt text. The information from the word alignment unit can be used to assess the proficiency and/or fluency of the user's speech.

Claims (32)

1. A speech processing system comprising:

an input for receiving an input utterance spoken by a user;

a speech recognition system that recognizes the input utterance spoken by the user and that outputs a recognition result comprising a sequence of recognized textual words, and an additional sequence of recognized sub-word units corresponding to the input utterance;

an acoustic model store that stores acoustic speech models;

a word alignment unit configured to receive the sequence of recognized textual words and the sequence of recognized sub-word units output by the speech recognition system and to align a sequence of said acoustic speech models corresponding to the received sequence of recognized textual words and the received sequence of recognized sub-word units with a sequence of acoustic feature vectors representing the input utterance spoken by the user and to output an alignment result identifying:

a time alignment between each word of the received sequence of recognized textual words and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each recognized textual word is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user;

a time alignment between each sub-word unit in the sequence of recognized sub-word units and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each sub-word unit is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user;

a speech scoring feature determining unit configured to use:

i) the time alignment between each word of the received sequence of recognized textual words and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each recognized textual word is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user; and

ii) the time alignment between each sub-word unit in the sequence of recognized sub-word units and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each sub-word unit is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user;

output from the word alignment unit to determine a plurality of speech scoring feature values for the input utterance; and

a scoring unit operable to use the plurality of speech scoring feature values for the input utterance determined by the speech scoring feature determining unit, to generate a score representing language ability of the user.

2. The speech processing system of claim 1 , wherein the word alignment unit is configured to output a sequence of sub-word units corresponding to a dictionary pronunciation of the recognized input utterance.

3. The speech processing system of claim 1 , wherein the word alignment unit is configured to output a sequence of sub-word units corresponding to a dictionary pronunciation of the matching possible utterance.

4. A speech processing system according to claim 3 , further comprising a sub-word alignment unit configured to receive the sequence of sub-word units corresponding to the dictionary pronunciation and configured to align the sequence of sub-word units corresponding to the dictionary pronunciation received from the word alignment unit with the input utterance spoken by the user whilst allowing for sub-word units to be inserted between words and for sub-word units of a word to be replaced by other sub-word units to determine where the input utterance spoken by the user differs from the dictionary pronunciation and to output a sequence of sub-word units corresponding to an actual pronunciation of the input utterance spoken by the user.

5. A speech processing system according to claim 4 , wherein the sub-word alignment unit is configured to use the sequence of sub-word units corresponding to the dictionary pronunciation of the recognized input utterance to generate a network having a plurality of paths allowing for sub-word units to be inserted between recognized words and for sub-word units of a recognized word to be replaced by other sub-word units and wherein the sub-word alignment unit is configured to align acoustic speech models for the different paths defined by the network with the input utterance spoken by the user.

6. A speech processing system according to claim 5 , wherein the sub-word alignment unit is configured to maintain a score representing the closeness of the match between the acoustic speech models for the different paths defined by the second network and input utterance spoken by the user.

7. A speech processing system according to claim 4 , further comprising a speech scoring feature determining unit configured to receive and to determine a measure of similarity between the sequence of sub-word units output by the word alignment unit and the sequence of sub-word units output by the sub-word alignment unit.

8. A speech processing system according to claim 1 , further comprising a free align unit configured to align acoustic speech models with the input utterance spoken by the user and to output an alignment result including a sequence of sub-word units that matches with the input utterance spoken by the user.

9. A speech processing system according to claim 1 , wherein the score represents the fluency and/or proficiency of the user's spoken utterance.

10. A speech processing method comprising:

receiving an input utterance spoken by a user;

using a speech recognition system to recognize the input utterance spoken by the user and to output a recognition result comprising a sequence of recognized textual words and an additional sequence of sub-word units corresponding to the input utterance; and

receiving the sequence of recognized textual words and the sequence of recognized sub-word units output by the speech recognition system and aligning a sequence of acoustic speech models corresponding to the received sequence of recognized textual words and the sequence of recognized sub-word units with a sequence of acoustic feature vectors representing the input utterance spoken by the user; and

outputting an alignment result identifying:

a time alignment between each word of the received sequence of recognized textual words and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each recognized textual word is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user; and

a time alignment between each sub-word unit in the sequence of recognized sub-word units and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each sub-word unit is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user;

using;

i) the time alignment between each word of the received sequence of recognized textual words and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each recognized textual word is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user; and

ii) the time alignment between each sub-word unit in the sequence of recognized sub-word units and the sequence of acoustic feature vectors representing the input utterance spoken by the user that identifies where each sub-word unit is located within the sequence of acoustic feature vectors representing the input utterance spoken by the user;

to determine a plurality of speech scoring feature values for the input utterance; and

using the determined plurality of speech scoring feature values for the input utterance to generate a score representing language ability of the user.

Assignments (2)
CHANGE OF ADDRESS Recorded Aug 4, 2020
From: THE CHANCELLOR, MASTERS, AND SCHOLARS OF THE UNIVERSITY OF CAMBRIDGE
To: THE CHANCELLOR, MASTERS, AND SCHOLARS OF THE UNIVERSITY OF CAMBRIDGE
Reel/Frame 053400/0400 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2020
From: ASH, THOMAS WILLIAM JOHN; ROBIINSON, ANTHONY JOHN
To: THE CHANCELLOR, MASTERS, AND SCHOLARS OF THE UNIVERSITY OF CAMBRIDGE
Reel/Frame 052992/0975 →
Priority Claims (1)
GB 1519494 · Nov 4, 2015 · national
Continuity (2)
Continuation 15773946
Related Publication 20200320987A1 · Oct 8, 2020
References Cited (45)
US 4805219A · Baker et al. · 1989 [cited by applicant]
US 5819223A · Takagi · 1998 [cited by examiner]
US 6389394B1 · Fanty · 2002 [cited by applicant]
US 6801891B2 · Garner et al. · 2004 [cited by applicant]
US 7054812B2 · Charlesworth et al. · 2006 [cited by applicant]
US 7062441B1 · Townshend · 2006 [cited by applicant]
US 7219059B2 · Gupta et al. · 2007 [cited by applicant]
US 7299188B2 · Gupta et al. · 2007 [cited by applicant]
US 7302389B2 · Gupta et al. · 2007 [cited by applicant]
US 7337116B2 · Charlesworth et al. · 2008 [cited by applicant]
US 7590533B2 · Hwang · 2009 [cited by applicant]
US 7668718B2 · Kahn et al. · 2010 [cited by applicant]
US 8457959B2 · Kaiser · 2013 [cited by applicant]
US 8494850B2 · Chelba et al. · 2013 [cited by applicant]
US 8959014B2 · Xu et al. · 2015 [cited by applicant]
US 9123339B1 · Shaw · 2015 [cited by examiner]
US 9336771B2 · Chelba · 2016 [cited by applicant]
US 9424834B2 · Simmons et al. · 2016 [cited by applicant]
US 20020022960A1 · Charlesworth et al. · 2002 [cited by applicant]
US 20020120447A1 · Charlesworth et al. · 2002 [cited by applicant]
US 20020120448A1 · Garner et al. · 2002 [cited by applicant]
US 20040006461A1 · Gupta et al. · 2004 [cited by applicant]
US 20040006468A1 · Gupta et al. · 2004 [cited by applicant]
US 20040230430A1 · Gupta et al. · 2004 [cited by applicant]
US 20040230431A1 · Gupta et al. · 2004 [cited by applicant]
US 20050203738A1 · Hwang · 2005 [cited by examiner]
US 20060110711A1 · Julia et al. · 2006 [cited by applicant]
US 20060149558A1 · Kahn et al. · 2006 [cited by applicant]
US 20080040119A1 · Ichikawa et al. · 2008 [cited by applicant]
US 20080140401A1 · Abrrash et al. · 2008 [cited by applicant]
US 20080221893A1 · Kaiser · 2008 [cited by applicant]
US 20100145698A1 · Chen et al. · 2010 [cited by applicant]
US 20100145707A1 · Ljolje et al. · 2010 [cited by applicant]
US 20110307241A1 · Waibel · 2011 [cited by examiner]
US 20120065976A1 · Deng · 2012 [cited by examiner]
US 20120078630A1 · Hagen et al. · 2012 [cited by applicant]
US 20130006612A1 · Xu et al. · 2013 [cited by applicant]
US 20130006623A1 · Chelba et al. · 2013 [cited by applicant]
US 20130325464A1 · Huang et al. · 2013 [cited by applicant]
US 20140141392A1 · Yoon · 2014 [cited by examiner]
US 20150348541A1 · Epstein et al. · 2015 [cited by applicant]
US 20150371633A1 · Chelba · 2015 [cited by applicant]
International Search Report and Written Opinion for PCT/GB2016/053456, mailed Mar. 6, 2017. [cited by applicant]
Search Report for British Patent Application No. 1519494.7, mailed May 20, 2016. [cited by applicant]
Search Report for British Patent Application No. 1519494.7, mailed Dec. 13, 2016. [cited by applicant]