IP Library › Granted Patent US 10,991,367
Granted Patent B2
US 10,991,367 · App. 15/857,012 · Granted Apr 27, 2021

Voice activated assistant activation prevention system

Inventor: Norihiro Edwin Aoki (San Jose, CA)
Assignee: PAYPAL, INC.
G10L15/22G10L15/30G10L25/51H04R3/005G10L2015/223H04R1/406H04R27/00H04R2227/003H04R2430/23
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,991,367
App. No.
15/857,012
Granted
Apr 27, 2021
Kind
B2
Abstract

A voice activated assistant activation prevention system includes a database storing audio source information describing a relative location of an audio source. The system is configured to monitor, using at least one microphone, for an activation command that is associated with a recording of a subsequent instruction command and a sending that subsequent instruction command through a network. When the system detects, using first audio information received by the at least one microphone, a first instance of the activation command, it determines a source location of the first instance of the activation command. If the system matches the source location of the first instance of the activation command to the relative location of the audio source described by the audio source information in the database, the system may prevent the sending though the network of an instruction command that immediately follows the first instance of the activation command.

Claims (86)

1. A voice activated assistant activation prevention system, comprising:

a database storing audio source information comprising a respective relative physical location for each of a plurality of machine-based audio sources that are co-located in a common physical area; and

one or more hardware processors coupled to a non-transitory memory and configured to read instructions to cause the system to perform operations comprising:

monitoring, using at least one microphone, for an activation command that is associated with a recording of a subsequent instruction command and a sending of the subsequent instruction command through a network;

detecting, using first audio information received by the at least one microphone, a first instance of the activation command in the common physical area;

determining a source physical location of the first instance of the activation command in the common physical area;

determining that the source physical location of the first instance of the activation command matches the relative physical location(s) in the common physical area of at least one of the plurality of machine-based audio sources stored in the database;

determining, in response to the matching, that an instruction command that immediately follows the first instance of the activation command includes a purchase instruction;

determining that the instruction command that immediately follows the first instance of the activation command includes the purchase instruction and is unauthorized, in response to determining that the source physical location of the first instance of the activation command in the common physical area matches the relative physical location(s) in the common physical area of the at least one of the plurality of machine based audio sources; and

preventing the sending of the unauthorized purchase instruction through the network.

2. The system of claim 1 , wherein the operations further comprise:

detecting, using second audio information received by the at least one microphone, audio from at least one of the plurality of machine-based audio sources;

determining, based at least partially on the second audio information, the respective relative physical location(s) of the at least one of the plurality of machine-based audio sources; and

storing the respective relative physical location(s) of the at least one of the plurality of machine-based audio sources in the database.

3. The system of claim 2 , wherein the determining the respective relative physical location(s) of at least one of the plurality of machine-based audio sources further comprises:

determining a respective direction of each of at the least one of the plurality of machine-based audio sources from the at least one microphone;

determining a respective distance of each of at the least one of the plurality of machine-based audio sources from the at least one microphone; and

storing each respective direction(s) and respective distance(s) in the database.

4. The system of claim 1 , wherein the operations further comprise:

receiving, through the network via a graphical user interface displayed on a user device, a designation of the relative physical location(s) of at least one of the plurality of machine-based audio sources; and

storing the relative physical location(s) of the at least one of the plurality of machine-based audio sources in the database.

5. The system of claim 1 , wherein the operations further comprise:

detecting, using second audio information received by the at least one microphone, a second instance of the activation command;

determining a source physical location of the second instance of the activation command;

determining that the source physical location of the second instance of the activation command does not match the relative physical location(s) of at least one of the plurality of machine-based audio sources stored in the database; and

allowing, in response to determining that the source physical location of the activation command does not match the relative physical location(s) of the at least one of the plurality of machine-based audio sources, the sending through the network of an instruction command that follows the second instance of the activation command.

6. A method for preventing activation of a voice activated assistant device, comprising:

detecting, by a voice activated assistant device using first audio information received by at least one microphone, a first instance of an activation command;

determining, by the voice activated assistant device, a source physical location of the first instance of the activation command;

determining, by the voice activated assistant device, that the source physical location of the first instance of the activation command matches a relative physical location in a common physical area of at least one of a plurality of machine-based audio sources that were previously stored in a database, wherein the plurality of machine-based audio sources are co-located in the common physical area;

determining, by the voice activated assistant device in response to determining that the source physical location of the first instance of the activation command matches the relative physical location in the common physical area of the at least one of the plurality of machine-based audio sources that were previously stored in the database, that an instruction command that immediately follows the first instance of the activation command includes a purchase instruction;

determining that the instruction command that immediately follows the first instance of the activation command includes the purchase instruction and is unauthorized; and

preventing the sending of the unauthorized purchase instruction through a network.

7. The method of claim 6 , further comprising:

detecting, by the voice activated assistant device using second audio information received by the at least one microphone, audio from at least one of the plurality of machine-based audio sources;

identifying, by the voice activated assistant device based at least partially on the second audio information, the relative physical location(s) of the at least one of the plurality of machine-based audio sources; and

storing, by the voice activated assistant device, the respective relative physical location(s) of the at least one of the plurality of machine-based audio sources in the database.

8. The method of claim 7 , wherein the determining the respective relative physical location(s) of at least one of the plurality of machine-based audio sources further comprises:

identifying, by the voice activated assistant device, a respective direction of the at least one of the plurality of machine-based audio sources from the at least one microphone;

identifying, by the voice activated assistant device, a respective distance of the at least one of the plurality of machine-based audio sources from the at least one microphone; and

storing, by the voice activated assistant device, each respective direction and each respective distance in the database.

9. The method of claim 6 , further comprising receiving, by the voice activated assistant device through the network via a graphical user interface displayed on a user device, a designation of the relative physical location(s) of at least one of the plurality of machine-based audio sources; and

storing, by the voice activated assistant device, the relative physical location(s) of the at least one of the plurality of machine-based audio sources in the database.

10. The method of claim 6 , further comprising:

detecting, by the voice activated assistant device using second audio information received by the at least one microphone, a second instance of the activation command;

determining, by the voice activated assistant device, a source physical location of the second instance of the activation command;

determining, by the voice activated assistant device, that the source physical location of the second instance of the activation command does not correspond to the relative physical location(s) of at least one of the plurality of machine-based audio sources described by the audio source information that was previously stored in the database; and

allowing, by the voice activated assistant device in response to determining that the source physical location of the activation command does not correspond to the relative physical location(s) of the at least one of the plurality of machine-based audio sources, the sending through the network of an instruction command that immediately follows the second instance of the activation command.

11. The method of claim 6 , wherein the determining that the instruction command that immediately follows the first instance of the activation command includes the purchase instruction further includes determining a priority of the instruction command that immediately follows the first instance of the activation command.

12. A non-transitory machine-readable medium having stored thereon machine-readable instructions executable to cause a machine to perform operations comprising:

detecting, using at least one microphone, a first instance of an activation command that is associated with a recording of a subsequent instruction command and a sending that subsequent instruction command through a network;

determining a source physical location of the first instance of the activation command;

matching the source physical location of the first instance of the activation command to relative physical location(s) in a common physical area of at least one of a plurality of machine-based audio sources that are identified in a database and that are co-located in the common physical area;

determining, in response to the matching, that an instruction command that immediately follows the first instance of the activation command includes a purchase instruction and is unauthorized; and

preventing, in response to determining that the instruction command that immediately follows the first instance of the activation command includes the purchase instruction, the sending of the unauthorized purchase instruction through the network.

13. The non-transitory machine-readable medium of claim 12 , wherein the operations further comprise:

detecting, using second audio information received by the at least one microphone, audio from at least one of the plurality of machine-based audio sources;

determining, based at least partially on the second audio information, the respective relative physical location(s) of the at least one of the plurality of machine-based audio sources; and

storing the respective relative physical location(s) of the at least one of the plurality of machine-based audio sources in the database.

14. The non-transitory machine-readable medium of claim 13 , wherein the determining the relative physical location of the audio source further comprises:

determining a respective direction of each of at least one of the plurality of machine-based audio sources from the at least one microphone;

determining a respective distance of each of at least one of the plurality of machine-based audio sources from the at least one microphone; and

storing each respective direction and each respective distance in the database.

15. The non-transitory machine-readable medium of claim 12 , wherein the operations further comprise:

receiving, through the network via a graphical user interface displayed on a user device, a designation of the relative physical location(s) of at least one of the plurality of machine-based audio sources; and

storing the relative physical location(s) of the at least one of the plurality of machine-based audio sources in the database.

16. The non-transitory machine-readable medium of claim 12 , wherein the operations further comprise:

detecting, using the at least one microphone, a second instance of the activation command;

determining a source physical location of the second instance of the activation command;

determining that the source physical location of the second instance of the activation command does not match the relative physical location(s) of at least one of the plurality of machine-based audio sources identified in the database; and

allowing, in response to determining that the source physical location of the activation command does not match the relative physical location(s) of the at least one of the plurality of machine-based audio sources, the sending through the network of an instruction command that immediately follows the second instance of the activation command.

17. The non-transitory machine-readable medium of claim 12 , wherein the preventing the recording of the purchase instruction includes:

providing, using at least one speaker and in response to determining that the instruction command that immediately follows the first instance of the activation command includes the purchase instruction, an audio request to confirm the purchase instruction; and

preventing, in response to not receiving a confirmation to the audio request, the recording of the purchase instruction.

18. The system of claim 1 , wherein the operations further comprise:

storing audio content that is to be transmitted to the at least one of the plurality of machine-based audio sources;

determining, by analyzing the stored audio content prior to transmitting the stored audio content to the at least one of the plurality of machine-based audio sources, that the stored audio content includes a second instance of the activation command; and

in response to determining that the stored audio content includes the second instance of the activation command, muting the at least one of the plurality of machine-based audio sources while the at least one of the plurality of machine-based audio sources outputs the second instance of the activation command.

19. The method of claim 6 , further comprising:

storing, by the voice activated assistant device, audio content that is to be transmitted to the at least one of the plurality of machine-based audio sources;

determining, by the voice activated assistant device analyzing the stored audio content prior to transmitting the stored audio content to the at least one of the plurality of machine-based audio sources, that the stored audio content includes a second instance of the activation command; and

in response to determining that the stored audio content includes the second instance of the activation command, muting, by the voice activated assistant device, the at least one of the plurality of machine-based audio sources while the at least one of the plurality of machine-based audio sources outputs the second instance of the activation command.

20. The non-transitory machine-readable medium of claim 12 , wherein the operations further comprise:

storing audio content that is to be transmitted to the at least one of the plurality of machine-based audio sources;

determining, by analyzing the stored audio content prior to transmitting the stored audio content to the at least one of the plurality of machine-based audio sources, that the stored audio content includes a second instance of the activation command; and

in response to determining that the stored audio content includes the second instance of the activation command, muting the at least one of the plurality of machine-based audio sources while the at least one of the plurality of machine-based audio sources outputs the second instance of the activation command.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE CORRESPONDENCE ADDRESS PREVIOUSLY RECORDED ON REEL 044586 FRAME 0092. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 16, 2018
From: AOKI, NORIHIRO EDWIN
To: PAYPAL, INC.
Reel/Frame 045612/0755 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2018
From: AOKI, NORIHIRO EDWIN
To: PAYPAL, INC,
Reel/Frame 044586/0092 →
Continuity (1)
Related Publication 20190206395A1 · Jul 4, 2019
Cited By (1)
US 12,243,437