IP Library Granted Patent US 7,467,089
Granted Patent B2
US 7,467,089 · App. 11/005,633 · Granted Dec 16, 2008

Combined speech and handwriting recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,467,089
App. No.
11/005,633
Granted
Dec 16, 2008
Kind
B2
Abstract

The invention relates to the combination of speech recognition with handwriting and/or character recognition. This includes the innovation of selecting one or more best-scoring recognition candidates as a function of recognition of both handwritten and spoken representations of a sequence of one or more words to be recognized. It also includes the innovation of using character or handwriting recognition of one or more letters to alphabetically filter speech recognition of one or more words. It also includes the innovations of using speech recognition of one or more letter-identifying words to alphabetically filter handwriting recognition, and of using speech recognition to correct handwriting recognition of one or more words.

Claims (82)

1. A method of word recognition comprising:

receiving a given spoken representation of a given sequence of one or more words to be recognized;

in association with said given spoken representation, receiving a filtering input consisting of handwriting input;

using handwriting recognition to define a filter representing one or more filtering sequences of characters selected by said recognition as most likely corresponding to said filtering input; and

using a combination of said filter and speech recognition performed on said given spoken representation to select one or more recognition candidates, each consisting of a sequence of one or more words, selected as a function of the closeness of their match against the given spoken representation and whether or not they match one of the one or more filtering character sequences associated with said filter, where a recognition candidate matches a given filtering sequence if the candidate starts with the same sequence of one or more characters as said filtering sequence, independently of whether the candidate contains more characters than the given filtering sequence.

2. A method as in claim 1 wherein:

a plurality of selected recognition candidates are displayed on a touch screen by handwriting recognition of said handwritten representation;

the user can select to respond to said display of recognition candidates by either:

selecting a given one of the recognition candidates in said choice lists as corresponding to said handwritten representation by touching a portion of the screen associated with the given candidate; or

selecting, by input on said touch screen, to enter said filtering input either by said speech recognition or by handwriting one or more characters on said touch screen.

3. A method as in claim 1 wherein said filtering input consists of connected-character handwriting.

4. A method as in claim 3 wherein:

said filter represents a plurality of sequences of one or more characters; and

said selection of recognition candidates selects a plurality of best scoring recognition candidates, different ones of which can match different sequences of characters represented by said filter.

5. A method as in claim 4 wherein said plurality of character sequences represented by one filter and used in said selection of recognition candidates can be of different character length.

6. A method as in claim 3 wherein:

said filter represents only one sequence of one or more characters which is used for filtering; and

said selection of recognition candidates selects a plurality of best scoring recognition candidates, all of which match said one character sequence.

7. A method as in claim 1 wherein said handwriting filtering input consists of one or more unconnected separate character drawings.

8. A method as in claim 7 wherein:

said filter represents a plurality of sequences of one or more characters; and

said selection of recognition candidates selects a plurality of best scoring recognition candidates, different ones of which can match different sequences of characters represented by said filter.

9. A method as in claim 7 wherein:

said filter represents only one sequence of one or more characters which is used for filtering; and

said selection of recognition candidates selects a plurality of best scoring recognition candidates, all of which match said one character sequence.

10. A method as in claim 1 :

further including:

receiving a plurality of said spoken representation, of which said given spoken representation is one, each representing a sequence of one or more words to be recognized;

using speech recognition to output a corresponding sequence of one or more words into a sequential body of text for each of said plurality of spoken representations;

responding to user input with a pointing device that touches a user selected sequence of one or more words in said body of text by selecting the touched sequence as a sequence to be corrected;

treating the sequence of words to be corrected as said given spoken representation; and then performing said:

receiving of said handwritten filtering input;

using of said handwriting recognition to define said filter; and

using of said combination of the filter and speech recognition to select one or more recognition candidates for said given spoken representation.

11. A method as in claim 1 wherein:

said receiving of handwriting filtering input includes receiving a succession of said filtering inputs in association with said spoken representation, in which each successive filtering input represents one or more additional initial characters of said spoken representation;

said using of handwriting recognition to define a filter includes performing said recognition after the receipt of each given one of said succession of filtering inputs to produce one or more incremental filter sequences, each representing one or more characters recognized as likely to correspond to the given filtering input;

said incremental filter sequences are added to said filter, concatenated to the end of filtering sequences produced by the addition of any previous incremental filtering sequences associated with the given spoken representation, to create a current filter after the receipt of each of said succession of filter inputs that represents one or more initial spellings for the spoken representation indicated by the one or more filtering inputs that have currently been received; and

said using of a combination of said filter and speech recognition to select one or more recognition candidates is performed separately in response to each of said current filters.

12. A method as in claim 11 wherein:

a plurality of said selected recognition candidates are displayed on a touch screen in response to each of a plurality of said selections made in response to said successive filtering inputs;

the user can select to respond to one of said displays of said recognition candidates by either:

selecting a given one of the recognition candidates in said choice lists as corresponding to said spoken representation by touching a portion of the screen associated with the given candidate; or

selecting to enter one of said successive filtering inputs by handwriting one or more characters on said touch screen.

13. A method as in claim 1 wherein:

a plurality of selected recognition candidates are displayed on a touch screen by speech recognition of said spoken representation;

the user can select to respond to said display of recognition candidates by either:

selecting a given one of the recognition candidates in said choice lists as corresponding to said spoken representation by touching a portion of the screen associated with the given candidate; or

selecting to enter said filtering input by handwriting one or more characters on said touch screen, which causes said handwriting recognition to define said filter and the combination of said filter and said speech recognition to select said set of one or more filtered recognition candidates.

14. A method as in claim 1 wherein:

a plurality of selected recognition candidates are displayed on a touch screen by speech recognition of said spoken representation;

the user can select to respond to said display of recognition candidates by either:

selecting a given one of the recognition candidates in said choice lists as corresponding to said spoken representation by touching a portion of the screen associated with the given candidate; or

selecting, by input on said touch screen, to enter said filtering input either by performing said handwriting of one or more characters on said touch screen or by speech recognition.

15. A method of word recognition comprising:

receiving a handwritten representation of a given sequence of one or more words to be recognized;

in association with said handwritten representation, receiving a filtering input consisting of one or more utterances representing a sequence of one or more letter identifying words;

using speech recognition to define a filter representing one or more filtering sequences of characters selected by said recognition as most likely corresponding to said filtering input; and

using a combination of said filter and handwriting recognition performed on said handwritten representation to select one or more recognition candidates, each consisting of a sequence of one or more words, selected as a function of the closeness of their match against the handwritten representation and whether or not they match one of the one or more filtering character sequences associated with said filter, where a recognition candidate matches a given filtering sequence if the candidate starts with the same sequence of one or more characters as said filtering sequence, independently of whether the candidate contains more characters than given filtering sequence.

16. A method as in claim 15 wherein:

said receiving of spoken filtering input includes receiving a succession of said filtering inputs in association with said handwritten representation, in which each successive filtering input represents one or more additional initial characters of said handwritten representation;

said using of spoken recognition to define a filter includes performing said recognition after the receipt of each given one of said succession of filtering inputs to produce one or more incremental filter sequences, each representing one or more characters recognized as likely to correspond to the given filtering input;

said incremental filter sequences are added to said filter, concatenated to the end of filtering sequences produced by the addition of any previous incremental filtering sequences associated with the given handwritten representation, to create a current filter after the receipt of each of said succession of filter inputs that represents one or more initial spellings for the handwritten representation indicated by the one or more filtering inputs that have currently been received; and

said using of a combination of said filter and handwriting recognition to select one or more recognition candidates is performed separately in response to each of said current filters.

17. A method as in claim 15 wherein:

said filter represents only one sequence of one or more characters which is used for filtering; and

said selection of recognition candidates selects a plurality of best scoring recognition candidates, all of which match said one character sequence.

18. A method as in claim 15 further including providing a user interface which enables a user to select whether the filtering input is recognized with discrete or continuous recognition.

19. A method as in claim 15 further including providing a user interface which enables a user to select whether the filtering input is recognized in a mode which favors the recognition of letter names or of non-letter-name-letter-identifying words.

20. A method as in claim 15 wherein:

the filtering input is a sequence of continuously spoken letter identifying words; and

the speech recognition is continuous speech recognition.

21. A method as in claim 15 wherein:

the filtering input is a sequence of discretely spoken letter identifying words; and

the speech recognition is discrete speech recognition.

22. A method as in claim 15 wherein:

said filter represents a plurality of sequences of characters; and

said selection of recognition candidates selects a plurality of best scoring recognition candidates, different ones of which can match different sequences of characters represented by said filter.

23. A method as in claim 22 wherein said plurality of character sequences represented by one filter and used in said selection of recognition candidates can be of different character length.

24. A method as in claim 23 wherein:

the filtering input is a sequence of continuously spoken letter names; and

the speech recognition is continuous speech recognition.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
MERGER Recorded Sep 13, 2012
From: VOICE SIGNAL TECHNOLOGIES, INC.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028952/0277 →