IP Library › Granted Patent US 8,775,183
Granted Patent B2
US 8,775,183 · App. 12/483,919 · Granted Jul 8, 2014

Application of user-specified transformations to automatic speech recognition results

Inventors: Jonathan E. Hamaker (Issaquah, WA); Keith C. Herold (Seattle, WA)
Assignee: Microsoft Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,775,183
App. No.
12/483,919
Granted
Jul 8, 2014
Kind
B2
Abstract

Textual transcription of speech is generated and formatted according to user-specified transformation and behavior requirements for a speech recognition system having input grammars and transformations. An apparatus may include a speech recognition platform configured to receive a user-specified transformation requirement, recognize speech in speech data into recognized speech according to a set of recognition grammars; and apply transformations to the recognized speech according to the user-specified transformation requirement. The apparatus may further be configured to receive a user-specified behavior requirement and transform the recognized speech according to the behavior requirement. Other embodiments are described and claimed.

Claims (42)

1. A computer-implemented method comprising:

receiving, from a client device, a user-specified behavior requirement in a speech recognition system wherein the user-specified behavior requirement changes a setting to an input speech recognition grammar specification prior to recognizing speech to select or de-select a built-in transformation from a set of built-in transformations of the speech recognition system to apply to recognized speech, and wherein the user-specified behavior requirement includes a semantic tag that specifies at least one of: a proper name, a phone number, an address, an email address, an internet protocol address, and a web address;

encoding the received user-specified behavior requirement into an input speech recognition grammar specification for the speech recognition system;

receiving a request to recognize speech from an application;

recognizing the speech to form a recognized word sequence;

transforming the recognized word sequence according to the user-specified behavior requirement according to a semantic tag associated with the word sequence; and

providing the transformed sequence to the application.

2. The method of claim 1 , comprising:

receiving a user-specified transformation requirement in the speech recognition system wherein the user-specified transformation requirement includes a semantic tag; and

transforming the recognized word sequence according to at least one of the user-specified behavior requirement or the user-specified transformation requirement according to a semantic tag associated with the word sequence.

3. The method of claim 2 , wherein transforming the recognized word sequence comprises:

determining which transformations to apply to a sub-segment of the recognized word sequence according to a semantic tag associated with the sub-segment.

4. The method of claim 3 , wherein the semantic tag further specifies at least one of: a date, and a time.

5. The method of claim 1 wherein the user-specified behavior requirement comprises at least one of: a grammar, an application program interface (API) call, an extensible markup language (XML) element, an XML schema, or speech recognition application parameters.

6. The method of claim 1 , wherein the user-specified behavior requirement specifies a first transformation for a first portion of the input speech recognition grammar specification, and a second transformation for a second portion of the input speech recognition grammar specification.

7. A system for implementing formatted recognized speech, the system comprising:

a speech recognition platform configured to:

receive a user-specified behavior requirement, wherein the user-specified behavior requirement changes a setting to an input speech recognition grammar specification to select or de-select a built-in transformation from a set of built-in transformations of the speech recognition system;

receive a user-specified transformation requirement, wherein the user-specified transformation requirement specifies prior to recognizing speech to the speech recognition platform transformations to apply to recognized speech and wherein the user-specified transformation requirement includes a semantic tag that specifies at least one of: a proper name, a phone number, an address, an email address, an internet protocol address, and a web address;

encode the user-specified transformation requirement into the input speech recognition grammar specification for the speech recognition platform;

recognize speech in speech data into recognized speech according to a set of recognition grammars including the input speech recognition grammar specification; and

apply transformations to the recognized speech according to the user-specified transformation requirement.

8. The system of claim 7 , further comprising:

a client application configured to:

create a user-specified transformation requirement;

encode the created user-specified transformation requirement into the speech recognition grammar specification for the speech recognition platform; and

request speech recognition using the user-specified transformation requirement from the speech recognition platform.

9. The system of claim 7 , wherein the user-specified behavior requirement includes a semantic tag that specifies at least one of: a proper name, a phone number, an address, an email address, an internet protocol address, and a web address; and

the speech recognition platform further configured to:

encode the user-specified behavior requirement into the speech recognition grammar specification for the speech recognition platform; and

transform the recognized word sequence according to the user-specified behavior requirement.

10. The system of claim 7 , wherein at least one of the user-specified transformation requirement or the user-specified behavior requirement comprises at least one of: a grammar, an application program interface (API) call, an extensible markup language (XML) element, an XML schema, or speech recognition application parameters.

11. The system of claim 9 , the speech recognition platform further configured to: determine which transformations to apply to a sub-segment of the recognized speech according to a semantic tag associated with the sub-segment.

12. The system of claim 7 , wherein the user-specified behavior requirement specifies a first transformation for a first portion of the input speech recognition grammar specification, and a second transformation for a second portion of the input speech recognition grammar specification.

13. An article comprising a storage medium containing instructions that if executed enable a system to:

receive a request to recognize speech from an application;

recognize the speech to form a recognized word sequence;

transform the recognized word sequence according to a user-specified behavior requirement encoded into an input speech recognition grammar specification, wherein the user-specified behavior requirement changes a setting prior to recognizing speech to the input speech recognition grammar to select or de-select a built-in transformation from a set of built-in transformations of a speech recognition system to apply to recognized speech, and wherein the user-specified behavior requirement includes a semantic tag that specifies at least one of: a proper name, a phone number, an address, an email address, an internet protocol address, and a web address; and

provide the transformed sequence to the application.

14. The article of claim 13 , further comprising instructions that if executed enable the system to transform the recognized word sequence according to at least one of the user-specified behavior requirement or a user-specified transformation requirement, wherein the user-specified transformation requirement includes a semantic tag that specifies at least one of: a proper name, a phone number, an address, an email address, an internet protocol address, and a web address.

15. The article of claim 13 , wherein at least one of the user-specified behavior requirement or a user-specified transformation requirement includes a semantic tag, and further comprising instructions that if executed enable the system to: determine which transformations to apply to a sub-segment of the recognized word sequence according to a semantic tag associated with the sub-segment.

16. The article of claim 13 , wherein the user-specified behavior requirement specifies a first transformation for a first portion of the input speech recognition grammar specification, and a second transformation for a second portion of the input speech recognition grammar specification.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034564/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2009
From: HAMAKER, JONATHAN E.; HEROLD, KEITH C.
To: MICROSOFT CORPORATION
Reel/Frame 022820/0951 →
Continuity (1)
Related Publication 20100318356A1 · Dec 16, 2010