IP Library Granted Patent US 10,917,404
Granted Patent B2
US 10,917,404 · App. 16/725,371 · Granted Feb 9, 2021

Authentication of packetized audio signals

Inventors: Gaurav Bhaya (Sunnyvale, CA); Robert Stets (Mountain View, CA)
Assignee: GOOGLE LLC
H04L63/0861G06F21/32G06F21/34G10L15/1822G10L17/02G10L17/24G10L25/51H04L67/02H04L67/142H04L67/20H04W4/025G06F2221/2111G06F2221/2115G10L2015/088H04W12/00503
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,917,404
App. No.
16/725,371
Granted
Feb 9, 2021
Kind
B2
Abstract

The present disclosure is generally directed a data processing system for authenticating packetized audio signals in a voice activated computer network environment. The data processing system can improve the efficiency and effectiveness of auditory data packet transmission over one or more computer networks by, for example, disabling malicious transmissions prior to their transmission across the network. The present solution can also improve computational efficiency by disabling remote computer processes possibly affected by or caused by the malicious audio signal transmissions. By disabling the transmission of malicious audio signals, the system can reduce bandwidth utilization by not transmitting the data packets carrying the malicious audio signal across the networks.

Claims (66)

1. A system to authenticate packetized audio signals in voice-activated computer network environments, comprising:

a data processing system comprising one or more processors and memory;

a natural language processor component executed by the data processing system to parse a first data packet comprising a first input audio signal acquired via a sensor of a client device to determine that a content provider is to handle a request indicated in the first input audio signal;

a conversational application programming interface executed by the data processing system to establish a communication session between the client device and the content provider determined to handle the request;

a content selector component executed by the data processing system to select a content item to provide to the client device for authentication of the first input audio signal; and

a network security appliance executed by the data processing system to:

receive, from the client device, a second data packet including a second input audio signal, the second input audio signal corresponding a response to the content item provided to the client device;

compare, for the authentication of the first audio packet, a characteristic of the second input audio signal of the second audio packet to a characteristic associated with the first input audio signal of the first audio packet; and

generate, in accordance with the comparison for the authentication of the first audio packet, a condition indicating one of a continuation or a termination of the communication session established between the content provider and the client device.

2. The system of claim 1 , comprising the network security appliance to:

determine, based on the comparison, that the characteristic of the second input audio signal of the second audio packet does not match the characteristic associated with the first input audio signal of the first audio packet;

generate, responsive to the determination, the condition indicating termination of the communication session; and

transmit, responsive to the generation of the condition indicating the termination, an instruction to the content provider to disable the communication session.

3. The system of claim 1 , comprising the network security appliance to:

determine, based on the comparison, that the characteristic of the second input audio signal of the second audio packet matches the characteristic associated with the first input audio signal of the first audio packet;

generate, responsive to the determination, the condition indicating continuation of the communication session; and

transmit, responsive to the generation of the condition indicating the continuation, an instruction to the content provider to maintain the communication session.

4. The system of claim 1 , comprising the network security appliance to:

identify an action data structure to handle the request indicated by the first input audio signal of the first data packet, the action data structure having a parameter to cause the content provider to perform an action; and

determine that the parameter of the action data structure does not match the characteristic of the first input audio signal of the first data packet; and

the content selector component to select, responsive to the determination, the content item for the authentication of the first input audio signal.

5. The system of claim 1 , comprising the network security appliance to:

identify a location of a second client device, the second client device associated with the client device; and

determine that a distance between a location of the client device and the location of the second client device is greater than a threshold distance, and

the content selector component to select, responsive to the determination, the content item for the authentication of the first input audio signal.

6. The system of claim 1 , comprising the network security appliance to:

identify an amount of computational resources to be consumed to complete the request indicated in the first input audio signal of the first data packet; and

determine that the amount of computation resources to be consumed is greater than a threshold amount, and

the content selector component to select, responsive to the determination, the content item for the authentication of the first input audio signal.

7. The system of claim 1 , comprising the network security appliance to:

identify a restriction specified by the content provider for the characteristic associated with the first input audio signal; and

the content selector component to select, responsive to the identification, the content item for the authentication of the first input audio signal.

8. The system of claim 1 , comprising the network security appliance to:

identify, from a plurality of characteristics of the first input audio signal, the characteristic of the first input audio signal based on the request indicating in the first input audio signal; and

identify, from a plurality of characteristic of the second input audio signal, the characteristic of the second input audio signal to compare with the characteristic of the first input audio signal.

9. The system of claim 1 , comprising the network security appliance to:

determine the characteristic of the first input audio signal including at least one of a voiceprint, a keyword, a number of voices detected, an identification of the client device, and a location of a source of the first input audio signal; and

determine the characteristic of the second input audio signal including at least one of a voiceprint, a keyword, a number of voices detected, an identification of the client device, and a location of a source of the second input audio signal.

10. The system of claim 1 , comprising the network security appliance to compare, for the authentication of the first audio packet, a hash of the characteristic of the second input audio signal to a hash of the characteristic associated with the first input audio signal.

11. The system of claim 1 , comprising a direct action application programming interface executed by the data processing system to determine, from a plurality of content providers, the content provider to handle the request indicated in the first input audio signal.

12. The system of claim 1 , comprising a direct action application programming interface executed by the data processing system to generate an action data structure to establish the communication session between the content provider and the client device, the action data structure having a parameter to cause the content provider to perform the action corresponding to the request.

13. The system of claim 1 , comprising the content selector component to select the content item for the authentication of the first input audio signal, the content item including at least one an output audio signal prompting for input audio or a visual notification prompting for an interaction.

14. The system of claim 1 , comprising the natural language processor component to parse the first input audio signal of the first data packet to identify the request, a trigger keyword corresponding to the request, the request corresponding to an action to be performed by the content provider in accordance with the trigger keyword.

15. A method of authenticating packetized audio signals in voice-activated computer network environments, comprising:

parsing, by a data processing system having one or more processors, a first data packet comprising a first input audio signal acquired via a sensor of a client device to determine that a content provider is to handle a request indicated in the first input audio signal;

establishing, by the data processing system, a communication session between the client device and the content provider determined to handle the request;

selecting, by the data processing system, a content item to provide to the client device for authentication of the first input audio signal;

receiving, by the data processing system, from the client device, a second data packet including a second input audio signal, the second input audio signal corresponding a response to the content item provided to the client device;

compare, by the data processing system, for the authentication of the first audio packet, a characteristic of the second input audio signal of the second audio packet to a characteristic associated with the first input audio signal of the first audio packet; and

generating, by the data processing system, in accordance with comparing for the authentication of the first audio packet, a condition indicating one of a continuation or a termination of the communication session established between the content provider and the client device.

16. The method of claim 15 , comprising

determining, by the data processing system, based on comparing, that the characteristic of the second input audio signal of the second audio packet does not match the characteristic associated with the first input audio signal of the first audio packet;

generating, by the data processing system, responsive to determining, the condition indicating termination of the communication session; and

transmitting, by the data processing system, responsive to generating the condition indicating the termination, an instruction to the content provider to disable the communication session.

17. The method of claim 15 , comprising

determining, by the data processing system, based on comparing, that the characteristic of the second input audio signal of the second audio packet matches the characteristic associated with the first input audio signal of the first audio packet;

generating, by the data processing system, responsive to determining, the condition indicating continuation of the communication session; and

transmitting, by the data processing system, responsive to the generation of the condition indicating the continuation, an instruction to the content provider to maintain the communication session.

18. The method of claim 15 , comprising

identifying, by the data processing system, an action data structure to handle the request indicated by the first input audio signal of the first data packet, the action data structure having a parameter to cause the content provider to perform an action; and

determining, by the data processing system, that the parameter of the action data structure does not match the characteristic of the first input audio signal of the first data packet; and

selecting, by the data processing system, responsive to determining, the content item for the authentication of the first input audio signal.

19. The method of claim 15 , comprising

determining, by the data processing system, the characteristic of the first input audio signal including at least one of a voiceprint, a keyword, a number of voices detected, an identification of the client device, and a location of a source of the first input audio signal; and

determining, by the data processing system, the characteristic of the second input audio signal including at least one of a voiceprint, a keyword, a number of voices detected, an identification of the client device, and a location of a source of the second input audio signal.

20. The method of claim 15 , comprising selecting, by the data processing system, the content item for the authentication of the first input audio signal, the content item including at least one an output audio signal prompting for input audio or a visual notification prompting for an interaction.

Assignments (2)
CHANGE OF NAME Recorded Jan 5, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 054899/0037 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2019
From: BHAYA, GAURAV; STETS, ROBERT
To: GOOGLE INC.
Reel/Frame 051358/0962 →
Continuity (2)
Continuation 15395729 · Dec 30, 2016
Related Publication 20200137053A1 · Apr 30, 2020