IP Library Granted Patent US 10,403,275
Granted Patent B1
US 10,403,275 · App. 15/449,825 · Granted Sep 3, 2019

Speech control for complex commands

Inventors: Timothy Earl Gill (Denver, CO); Alex Nathan Capecelatro (Los Angeles, CA)
Assignee: Josh.ai LLC
G10L15/22G06F17/271G06F17/2755G06F17/2775G10L15/04G10L15/05G10L15/26G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,403,275
App. No.
15/449,825
Granted
Sep 3, 2019
Kind
B1
Abstract

Audio content associated with a verbal utterance is received. An operational meaning of a received verbal utterance comprising a compound input is recognized and determined, at least in part by: determining that a first subset of the received verbal utterance is associated with a first recognized input; determining whether a meaning of a remaining portion of the received verbal utterance other than the first subset is recognized as being associated with a second recognized input; and based at least in part on a determination that the meaning of the remaining portion of the received verbal utterance is recognized as being associated with said second recognized input, concluding that the verbal utterance comprises a compound input comprising the first and second recognized inputs.

Claims (42)

1. A system, comprising:

a communication interface configured to receive audio content associated with a verbal utterance;

a processor coupled to the communication interface and configured to recognize and determine an operational meaning of a received verbal utterance, at least in part by:

determining that a first subset of the received verbal utterance is associated with a first recognized input, comprising at least in part:

placing a breakpoint between every word of the received verbal utterance at least in part to separate sentences between breakpoints; and

removing a given breakpoint when it is at least in part determined that a candidate sentence associated with the given breakpoint is of low score quantified by a quality metric that is decreased for sentences that make less sense;

determining whether a meaning of a remaining portion of the received verbal utterance other than the first subset is recognized as being associated with a second recognized input; and

based at least in part on a determination that the meaning of the remaining portion of the received verbal utterance is recognized as being associated with said second recognized input, concluding that the received verbal utterance comprises a compound input comprising the first and second recognized inputs.

2. The system of claim 1 , wherein the communication interface is coupled with at least one of the following: microphone, audio port, audio recording device, and audio card.

3. The system of claim 1 , wherein determining that the first subset of the received verbal utterance is associated with the first recognized input comprises determining the first recognized input is a candidate for being associated.

4. The system of claim 1 , wherein determining that the first subset of the received verbal utterance is associated with the first recognized input comprises determining the first recognized input is of a prescribed confidence for being associated.

5. The system of claim 1 , wherein determining whether the meaning of the remaining portion of the received verbal utterance other than the first subset is recognized as being associated with the second recognized input includes determining the second recognized input includes a second verb and second object.

6. The system of claim 1 , wherein determining whether the meaning of the remaining portion of the received verbal utterance other than the first subset is recognized as being associated with the second recognized input includes determining the second recognized input includes a second object to be associated with a first verb of the first recognized input.

7. The system of claim 1 , further comprising using a speech-to-text engine on the audio content to process the verbal utterance.

8. The system of claim 1 , wherein determining that the first subset of the received verbal utterance is associated with the first recognized input is based at least in part on a breakpoint analysis.

9. The system of claim 8 , wherein the breakpoint analysis comprises using rule matching on a sentence fragment prior to a proposed sentence boundary.

10. The system of claim 8 , wherein the breakpoint analysis comprises using separator words to propose sentence boundaries.

11. The system of claim 8 , wherein the breakpoint analysis comprises using separator words to propose sentence boundaries.

12. The system of claim 1 , wherein a sorting rule matching is used to determine whether a given input is at least one of the following: a question, a statement, and a command.

13. The system of claim 12 , further comprising:

inserting appropriate punctuation into a written representation of the received verbal utterance based at least in part on the sorting rule matching;

capitalizing sentences associated with a written representation of the received verbal utterance based at least in part on the sorting rule matching; and

modifying proper nouns of a written representation of the received verbal utterance based at least in part on the sorting rule matching.

14. The system of claim 1 , further comprising determining a natural response to the compound input based on at least one of the following: simple deletion, simple combination, simple reordering, abstraction and anaphora resolution.

15. The system of claim 14 , wherein simple deletion comprises using a rule that determines importance with candidate responses and deletes less important responses.

16. The system of claim 14 , wherein simple combination comprises using at least one of a rule and a special-case text processing to combine a plurality of candidate responses into a more natural response.

17. The system of claim 14 , wherein abstraction comprises determining the compound input have equal weighting of importance and cannot be combined.

18. The system of claim 14 , wherein anaphora resolution comprises tracking a referent to shorten a response.

19. A method, comprising:

receiving audio content associated with a verbal utterance;

recognizing and determining an operational meaning of a received verbal utterance comprising a compound input, at least in part by:

determining that a first subset of the received verbal utterance is associated with a first recognized input, comprising at least in part:

placing a breakpoint between every word of the received verbal utterance at least in part to separate sentences between breakpoints; and

removing a given breakpoint when it is at least in part determined that a candidate sentence associated with the given breakpoint is of low score quantified by a quality metric that is decreased for sentences that make less sense;

determining whether a meaning of a remaining portion of the received verbal utterance other than the first subset is recognized as being associated with a second recognized input; and

based at least in part on a determination that the meaning of the remaining portion of the received verbal utterance is recognized as being associated with said second recognized input, concluding that the received verbal utterance comprises a compound input comprising the first and second recognized inputs.

20. A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

determining that a first subset of a received verbal utterance is associated with a first recognized input, comprising at least in part:

placing a breakpoint between every word of the received verbal utterance at least in part to separate sentences between breakpoints; and

removing a given breakpoint when it is at least in part determined that a candidate sentence associated with the given breakpoint is of low score quantified by a quality metric that is decreased for sentences that make less sense;

determining whether a meaning of a remaining portion of the received verbal utterance other than the first subset is recognized as being associated with a second recognized input; and

based at least in part on a determination that the meaning of the remaining portion of the received verbal utterance is recognized as being associated with said second recognized input, concluding that the received verbal utterance comprises a compound input comprising the first and second recognized inputs.

Assignments (3)
CHANGE OF NAME Recorded Dec 11, 2020
From: JOSH.AI LLC
To: JOSH.AI, INC.
Reel/Frame 054700/0547 →
CHANGE OF NAME Recorded Apr 12, 2019
From: J* LLC
To: JOSH.AI LLC
Reel/Frame 048879/0543 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2017
From: GILL, TIMOTHY EARL; CAPECELATRO, ALEX NATHAN
To: J* LLC
Reel/Frame 041938/0085 →
Continuity (1)
Provisional Application 62367999 · Jul 28, 2016