IP Library Granted Patent US 10,541,998
Granted Patent B2
US 10,541,998 · App. 15/863,042 · Granted Jan 21, 2020

Authentication of packetized audio signals

Inventors: Gaurav Bhaya (Sunnyvale, CA); Robert Stets (Mountain View, CA)
Assignee: GOOGLE LLC
H04L63/0861G06F21/32G06F21/34G10L15/1822G10L17/02G10L17/24G10L25/51H04L67/02H04L67/142H04L67/20H04W4/025G06F2221/2111G06F2221/2115G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,541,998
App. No.
15/863,042
Granted
Jan 21, 2020
Kind
B2
Abstract

The present disclosure is generally directed a data processing system for authenticating packetized audio signals in a voice activated computer network environment. The data processing system can improve the efficiency and effectiveness of auditory data packet transmission over one or more computer networks by, for example, disabling malicious transmissions prior to their transmission across the network. The present solution can also improve computational efficiency by disabling remote computer processes possibly affected by or caused by the malicious audio signal transmissions. By disabling the transmission of malicious audio signals, the system can reduce bandwidth utilization by not transmitting the data packets carrying the malicious audio signal across the networks.

Claims (84)

1. A system to authenticate packetized audio signals in a voice activated computer network environment, comprising:

a data processing system comprising at least one processor and memory;

a natural language processor component executed by the data processing system to receive, via an interface of the data processing system, data packets comprising an input audio signal detected by a sensor of a client device;

the natural language processor component to parse the input audio signal to identify a request and a trigger keyword corresponding to the request;

a direct action application programming interface of the data processing system to generate, based on the trigger keyword, a first action data structure responsive to the request;

a network security appliance to authenticate the input audio signal detected by the sensor of the client device based on a characteristic of the input audio signal and the first action data structure generated by the direct action application programming interface;

the direct action application programming interface to identify an account based on the input audio signal authenticated by the network security appliance, and transmit, to a third-party provider device, the first action data structure responsive to the network security appliance authenticating the input audio signal based on the characteristic of the input audio signal and the first action data structure, receipt of the first action data structure by the third-party provider device causing the third-party provider device to execute the first action data structure;

receive second data packets comprising a second input audio signal detected by the sensor of the client device;

generate an alarm condition responsive to a second characteristic of the second input audio signal not matching a parameter associated with the first action data structure; and

transmit, responsive to generation of the alarm condition, an instruction to the third-party provider device to terminate a communication session associated with, or execution of, the first action data structure.

2. The system of claim 1 , comprising:

the network security appliance to authenticate the input audio signal based on the characteristic of the input audio signal, the characteristic including at least one a voiceprint.

3. The system of claim 1 , comprising the network security appliance to:

measure a spectral component of the input audio signal to identify a voiceprint;

compare the spectral component with a stored voiceprint generated during a setup phase;

and authenticate the input audio signal based on the comparison of the spectral component with the stored voiceprint.

4. The system of claim 1 , comprising the network security appliance to:

determine the characteristic of the input audio signal based on a location of the client device; and

authenticate the input audio signal based on the location of the client device.

5. The system of claim 1 , comprising the network security appliance to:

determine the characteristic of the input audio signal with a physical authentication device; and

authenticate the input audio signal based on a response from the physical authentication device.

6. The system of claim 1 , comprising the network security appliance to:

determine the characteristic of the input audio signal based on a number of voices detected in the input audio signal; and

authenticate the input audio signal based on the number of voices detected.

7. The system of claim 1 , comprising the network security appliance to:

receive a third input audio signal detected by the sensor of the client device;

detect, based on a characteristic of the third input audio signal, a second alarm condition; and

prevent, responsive to the second alarm condition, execution of a third action data structure based on the third input audio signal.

8. The system of claim 1 , comprising the data processing system to:

perform a search based on the first action data structure and the characteristic of the input audio signal authenticated by the network security appliance.

9. The system of claim 1 , comprising the data processing system to:

identify the account based on the characteristic of the input audio signal;

identify a preference associated with the account; and

perform a search based on the preference.

10. The system of claim 1 , comprising the data processing system to:

identify the account based on the characteristic of the input audio signal;

identify a first preference associated with the account;

perform a first search based on the first preference;

identify a second account based on the characteristic of the second input audio signal detected by the sensor of the client device;

identify a second preference associated with the second account; and

perform a second search based on the second preference.

11. A method of authenticating packetized audio signals in a voice activated computer network environment, comprising:

receiving, by a data processing system comprising at least one processor and memory that executes a natural language processor component, via an interface of the data processing system, data packets comprising an input audio signal detected by a sensor of a client device;

parsing, by the natural language processor component, the input audio signal to identify a request and a trigger keyword corresponding to the request;

generating, by a direct action application programming interface of the data processing system, based on the trigger keyword, a first action data structure responsive to the request;

authenticating, by a network security appliance, the input audio signal detected by the sensor of the client device based on a characteristic of the input audio signal and the first action data structure generated by the direct action application programming interface; and

identifying, by the data processing system, an account based on the input audio signal authenticated by the network security appliance;

transmitting, by the data processing system, to a third-party provider device, the first action data structure responsive to the network security appliance authenticating the input audio signal based on the characteristic of the input audio signal and the first action data structure, receipt of the first action data structure by the third-party provider device causing the third-party provider device to execute the first action data structure;

receiving, by the data processing system, second data packets comprising a second input audio signal detected by the sensor of the client device;

generating, by the data processing system, an alarm condition responsive to a second characteristic of the second input audio signal not matching a parameter associated with the first action data structure; and

transmitting, by the data processing system, responsive to generation of the alarm condition, an instruction to the third-party provider device to terminate a communication session associated with, or execution of, the first action data structure.

12. The method of claim 11 , comprising:

authenticating, by the network security appliance, the input audio signal based on the characteristic of the input audio signal, the characteristic including at least one a voiceprint.

13. The method of claim 11 , comprising:

measuring a spectral component of the input audio signal to identify a voiceprint;

comparing the spectral component with a stored voiceprint generated during a setup phase; and

authenticating the input audio signal based on the comparing of the spectral component with the stored voiceprint.

14. The method of claim 11 , comprising:

determining the characteristic of the input audio signal based on a location of the client device; and

authenticating the input audio signal based on the location of the client device.

15. The method of claim 11 , comprising:

determining the characteristic of the input audio signal with a physical authentication device; and

authenticating the input audio signal based on a response from the physical authentication device.

16. The method of claim 11 , comprising:

determining the characteristic of the input audio signal based on a number of voices detected in the input audio signal; and

authenticating the input audio signal based on the number of voices detected.

17. The method of claim 11 , comprising:

receive a third input audio signal detected by the sensor of the client device;

detect, based on a characteristic of the third input audio signal, a second alarm condition; and

prevent, responsive to the second alarm condition, execution of a third action data structure based on the third input audio signal.

18. The method of claim 11 , comprising:

performing a search based on the first action data structure and the characteristic of the input audio signal authenticated by the network security appliance.

19. The method of claim 11 , comprising:

identifying the account based on the characteristic of the input audio signal;

identifying a preference associated with the account; and

performing a search based on the preference.

20. The method of claim 11 , comprising:

identifying the account based on the characteristic of the input audio signal;

identifying a first preference associated with the account;

performing a first search based on the first preference;

identifying a second account based on the characteristic of the second input audio signal detected by the sensor of the client device;

identifying a second preference associated with the second account; and

performing a second search based on the second preference.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2019
From: BHAYA, GAURAV; STETS, ROBERT
To: GOOGLE INC.
Reel/Frame 050709/0392 →
CHANGE OF NAME Recorded Oct 14, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 050724/0387 →
Continuity (2)
Continuation 15395729 · Dec 30, 2016
Related Publication 20180191713A1 · Jul 5, 2018