IP Library › Granted Patent US 10,942,703
Granted Patent B2
US 10,942,703 · App. 16/249,301 · Granted Mar 9, 2021

Proactive assistance based on dialog communication between devices

Inventors: Mathieu Jean Martel (Paris, FR); Thomas Deniau (Paris, FR)
Assignee: Apple Inc.
G06F3/167G06F3/165G06F16/3329G06F40/279G06F40/30G10L15/22G10L15/26G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,942,703
App. No.
16/249,301
Granted
Mar 9, 2021
Kind
B2
Abstract

Systems and processes for proactive assistance based on dialog communication between devices are provided. In one example process, while voice communication between an electronic device and a second electronic device is established, a stream of audio data associated with the second electronic device can be received. In response to detecting a user input, a text representation of speech contained in a portion of the stream of audio data can be generated. The process can determine whether the text representation contains information corresponding to one of a plurality of types of information. In response to determining that the text representation contains information corresponding to one of a plurality of types of information, one or more tasks based on the information can be performed.

Claims (68)

1. A non-transitory computer-readable medium storing instructions for providing proactive assistance based on dialog communication between devices, the instructions, when executed by one or more processors, cause the one or more processors to:

while voice communication is established between an electronic device and a second electronic device:

receive a stream of audio data associated with the second electronic device;

identify, based on at least one sentence boundary, a plurality of portions of the stream of audio data;

store the plurality of portions of the stream of audio data;

detect a user input;

in response to detecting the user input, generate a text representation of speech contained in a first portion of the plurality of portions of the stored audio data;

determine whether the text representation contains information corresponding to one of a plurality of types of information;

in response to determining that the text representation contains information corresponding to one of a plurality of types of information, determine whether the information is complete;

in response to determining that the information is not complete:

generate a text representation of speech contained in a second portion of the plurality of portions of the stored audio data; and

obtain second information from the second portion of the plurality of portions of the stored audio data;

perform one or more tasks based on at least the information and the second information.

2. The computer-readable medium of claim 1 , wherein the second portion of stored audio data corresponds to audio that is less recent that the first portion of the stored audio data.

3. The computer-readable medium of claim 1 , wherein performing one or more tasks based on at least the information and the second information comprises performing one or more tasks based on both the information and the second information.

4. The computer-readable medium of claim 1 , wherein determining whether the information is complete comprises:

determining whether the information is missing at least one parameter.

5. The computer-readable medium of claim 1 , wherein the information includes a telephone number, wherein the one or more tasks include displaying the telephone number, and wherein the instructions further cause the one or more processors to:

in response to detecting a user selection of the displayed telephone number, initiate a voice call based on the telephone number.

6. The computer-readable medium of claim 1 , wherein the information includes a telephone number, wherein the one or more tasks include displaying the telephone number, and wherein the instructions further cause the one or more processors to:

in response to detecting a user selection of the displayed telephone number, store the telephone number in association with an address book of the electronic device.

7. The computer-readable medium of claim 1 , wherein the information includes an email address, wherein the one or more tasks include displaying the email address, and wherein the

instructions further cause the one or more processors to:

in response to detecting a user selection of the displayed email address, initiate a composition of an email message, wherein a recipient of the email message is based on the email address.

8. The computer-readable medium of claim 1 , wherein the information includes a location, and wherein the one or more tasks include displaying a map indicating the location.

9. The computer-readable medium of claim 1 , wherein the one or more tasks are performed after dialog communication has ended.

10. The computer-readable medium of claim 1 , wherein the one or more tasks includes a plurality of tasks, and performing one or more tasks based on at least the information and the second information comprises:

performing at least a first task of the plurality of tasks while dialog communication is established; and

performing at least a second task of the plurality of tasks after dialog communication has ended.

11. The computer-readable medium of claim 1 , wherein performing one or more tasks based on at least the information and the second information comprises providing at least one sound output.

12. The computer-readable medium of claim 1 , wherein performing one or more tasks based on at least the information and the second information comprises providing at least one haptic output.

13. The computer-readable medium of claim 1 , wherein identifying, based on at least one sentence boundary, a plurality of portions of the stream of audio data comprises:

detecting an audio amplitude of the stream of audio data; and

identifying a first sentence boundary based on the audio amplitude decreasing below a first threshold level.

14. The computer-readable medium of claim 13 , wherein identifying a first sentence boundary based on the audio amplitude decreasing below a first threshold level comprises:

within a predetermined time interval:

detecting the audio amplitude decrease from above the first threshold level; and

detecting the audio amplitude decrease to below a second threshold level.

15. An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

while voice communication is established between the electronic device and a second electronic device:

receiving a stream of audio data associated with the second electronic device;

identifying, based on at least one sentence boundary, a plurality of portions of the stream of audio data;

storing the plurality of portions of the stream of audio data;

detecting a user input;

in response to detecting the user input, generating a text representation of speech contained in a first portion of the plurality of portions of the stored audio data;

determining whether the text representation contains information corresponding to one of a plurality of types of information;

in response to determining that the text representation contains information corresponding to one of a plurality of types of information, determining whether the information is complete;

in response to determining that the information is not complete:

generating a text representation of speech contained in a second portion of the plurality of portions of the stored audio data; and

obtaining second information from the second portion of the plurality of portions of the stored audio data;

performing one or more tasks based on at least the information and the second information.

16. A method, comprising:

at an electronic device with one or more processors and memory:

while voice communication is established between the electronic device and a second electronic device:

receiving a stream of audio data associated with the second electronic device;

identifying, based on at least one sentence boundary, a plurality of portions of the stream of audio data;

storing the plurality of portions of the stream of audio data;

detecting a user input;

in response to detecting the user input, generating a text representation of speech contained in a first portion of the plurality of portions of the stored audio data;

determining whether the text representation contains information corresponding to one of a plurality of types of information;

in response to determining that the text representation contains information corresponding to one of a plurality of types of information, determining whether the information is complete;

in response to determining that the information is not complete:

generating a text representation of speech contained in a second portion of the plurality of portions of the stored audio data; and

obtaining second information from the second portion of the plurality of portions of the stored audio data;

performing one or more tasks based on at least the information and the second information.

Continuity (3)
Continuation 15169348 · May 31, 2016
Provisional Application 62387547 · Dec 23, 2015
Related Publication 20190220245A1 · Jul 18, 2019
Cited By (20)
US 12,197,817 US 12,200,297 US 12,211,502 US 12,216,894 US 12,219,314 US 12,236,952 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,477,470 US 12,608,171 US 12,619,452 US 12,633,289 US 12,748,568