IP Library › Granted Patent US 11,514,893
Granted Patent B2
US 11,514,893 · App. 16/818,414 · Granted Nov 29, 2022

Voice context-aware content manipulation

Inventors: Erez Kikin-Gil (Bellevue, WA); Emily Tran (Seattle, WA); Benjamin David Smith (Woodinville, WA); Alan Liu (Seattle, WA); Erik Thomas Oveson (Renton, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/1815G06N20/00G10L15/183G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,514,893
App. No.
16/818,414
Granted
Nov 29, 2022
Kind
B2
Abstract

Techniques performed by a data processing system for processing voice content received from a user herein include receiving a first audio input from a user comprising spoken content, analyzing the first audio input using one or more natural language processing models to produce a first textual output comprising a textual representation of the first audio input, analyzing the first textual output using one or more machine learning models to determine first context information of the first textual output, and processing the first textual output in the application based on the first context information.

Claims (45)

1. A data processing system comprising:

a processor; and

a computer-readable medium storing executable instructions for causing the processor to perform operations of:

receiving a first audio input from a user comprising spoken content;

analyzing the first audio input using one or more natural language processing models to produce a first textual output comprising a textual representation of the first audio input;

analyzing the first textual output using one or more machine learning models to determine first context information of the first textual output, wherein the first context information identifies a type of application being used by the user and includes information indicative of how the user was interacting with the application when the first audio input was received; and

processing the first textual output in the application based on the first context information.

2. The data processing system of claim 1 , wherein the spoken content comprises a command, textual content, or both; and wherein the first context information of the first textual output provides an indication of whether the first textual output includes the command and an indication of how the user intended to apply the command to content in the application.

3. The data processing system of claim 2 , wherein the instructions to process the first textual output in the application based on the first context information further include instructions configured to cause the processor to perform operations of rendering first textual content to a document in the application responsive to the first context information indicating that the first textual output includes the textual content.

4. The data processing system of claim 2 , wherein the instructions to process the first textual output in the application based on the first context information further include instructions configured to cause the processor to perform operations of executing a command on the contents of a document in the application responsive to the first context information indicating that the first textual output includes the command.

5. The data processing system of claim 1 , wherein the instructions to analyze the first textual output further include instructions configured to cause the processor to perform an operation of disambiguating between textual input and command input included in the first textual output based on the output of the one or more machine learning models.

6. The data processing system of claim 1 , further including instructions configured to cause the processor to perform operations of:

receiving usage information from the application indicative of user interactions with the application prior to receiving the first audio input, while receiving the first audio input, or after receiving the first audio input; and

wherein analyzing the first textual output further comprises analyzing the first textual output and the application usage information using one or more machine learning models to determine the first context information of the first textual output.

7. The data processing system of claim 6 , wherein the instructions to analyze the first textual output further include instructions configured to cause the processor to perform operations of:

disambiguating command scope based on the usage information from the application.

8. A method performed by a data processing system for processing voice content received from a user, the method comprising:

receiving a first audio input from a user comprising spoken content;

analyzing the first audio input using one or more natural language processing models to produce a first textual output comprising a textual representation of the first audio input;

analyzing the first textual output using one or more machine learning models to determine first context information of the first textual output, wherein the first context information identifies a type of application being used by the user and includes information indicative of how the user was interacting with the application when the first audio input was received; and

processing the first textual output in the application based on the first context information.

9. The method of claim 8 , wherein the spoken content comprises a command, textual content, or both, and wherein the first context information of the first textual output provides an indication of whether the first textual output includes the command and an indication of how the user intended to apply the command to content in the application.

10. The method of claim 9 , wherein processing the first textual output in the application based on the first context information further comprises:

rendering first textual content to a document in the application responsive to the first context information indicating that the first textual output includes the textual content.

11. The method of claim 9 , wherein processing the first textual output in the application based on the first context information further comprises:

executing a command on the contents of a document in the application responsive to the first context information indicating that the first textual output includes the command.

12. The method of claim 8 , wherein analyzing the first textual output using one or more machine learning models to determine the first context information of the first textual output further comprises:

disambiguating between textual input and command input included in the first textual output based on the output of the one or more machine learning models.

13. The method of claim 8 , further comprising:

receiving usage information from the application indicative of user interactions with the application prior to receiving the first audio input, while receiving the first audio input, or after receiving the first audio input; and

wherein analyzing the first textual output further comprises analyzing the first textual output and the application usage information using one or more machine learning models to determine the first context information of the first textual output.

14. The method of claim 13 , wherein analyzing the first textual output using one or more machine learning models to determine the first context information of the first textual output further comprises:

disambiguating command scope based on the usage information from the application.

15. A memory device storing instructions that, when executed on a processor of a data processing system, cause the data processing system to process voice content received from a user, by:

receiving a first audio input from a user comprising spoken content;

analyzing the first audio input using one or more natural language processing models to produce a first textual output comprising a textual representation of the first audio input;

analyzing the first textual output using one or more machine learning models to determine first context information of the first textual output, wherein the first context information identifies a type of application being used by the user and includes information indicative of how the user was interacting with the application when the first audio input was received; and

processing the first textual output in the application based on the first context information.

16. The memory device of claim 15 , wherein the spoken content comprises a command, textual content, or both; and wherein the first context information of the first textual output provides an indication of whether the first textual output includes the command and an indication of how the user intended to apply the command to content in the application.

17. The memory device of claim 16 , wherein the instructions to process the first textual output in the application based on the first context information further include instructions configured to cause the processor to perform operations of rendering first textual content to a document in the application responsive to the first context information indicating that the first textual output includes the textual content.

18. The memory device of claim 16 , wherein the instructions to process the first textual output in the application based on the first context information further include instructions configured to cause the processor to perform operations of executing a command on the contents of a document in the application responsive to the first context information indicating that the first textual output includes the command.

19. The memory device of claim 15 , wherein the instructions to analyze the first textual output using one or more machine learning models to determine the first context information of the first textual output further include instructions configured to cause the processor to perform operations of disambiguating between textual input and command input included in the first textual output based on the output of the one or more machine learning models.

20. The memory device of claim 15 , further including instructions configured to cause the processor to perform operations of:

receiving usage information from the application indicative of user interactions with the application prior to receiving the first audio input, while receiving the first audio input, or after receiving the first audio input; and

wherein the instructions configured to cause the processor to analyze the first textual output further include instructions to cause the processor to perform the operations of analyzing the first textual output and the application usage information using one or more machine learning models to determine the first context information of the first textual output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2020
From: KIKIN-GIL, EREZ; TRAN, EMILY; SMITH, BENJAMIN DAVID; LIU, ALAN; OVESON, ERIK THOMAS
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 052110/0505 →
Continuity (2)
Provisional Application 62967352 · Jan 29, 2020
Related Publication 20210233522A1 · Jul 29, 2021