IP Library Granted Patent US 11,880,442
Granted Patent B2
US 11,880,442 · App. 17/543,371 · Granted Jan 23, 2024

Authentication of audio-based input signals

Inventors: Ken Krieger (Jackson, WY); Andrew Joseph Alexander Gildfind (London, GB); Nicholas Salvatore Arini (Southampton, GB); Simon Michael Rowe (London, GB); Raimundo Mirisola (Zug, CH); Gaurav Bhaya (Sunnyvale, CA); Robert Stets (Mountain View, CA)
Assignee: GOOGLE LLC
G06F21/32G06F21/316G06F21/34G06F21/35G06V40/172G10L17/00G10L17/24H04L63/0861H04L63/107H04N21/4223H04N21/42203H04N21/44218
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,880,442
App. No.
17/543,371
Granted
Jan 23, 2024
Kind
B2
Abstract

The present disclosure is generally directed a data processing system for authenticating packetized audio signals in a voice activated computer network environment. The data processing system can improve the efficiency and effectiveness of auditory data packet transmission over one or more computer networks by, for example, disabling malicious transmissions prior to their transmission across the network. The present solution can also improve computational efficiency by disabling remote computer processes possibly affected by or caused by the malicious audio signal transmissions. By disabling the transmission of malicious audio signals, the system can reduce bandwidth utilization by not transmitting the data packets carrying the malicious audio signal across the networks.

Claims (38)

1. A method implemented by one or more processors, the method comprising:

receiving an audio-based input, detected by a microphone of a computing device, that includes a voice input of a user;

processing, using facial recognition, video-based input that is detected at a camera of the computing, device;

determining, based on processing of the audio-based input and the processing of the video-based input, that a probability, that the audio-based input was generated by a particular registered user of the computing device, satisfies a threshold;

determining that a distance, between the computing device and the particular registered user, satisfies a distance threshold; and

in response to determining that the probability satisfies the threshold and that the distance satisfies the distance threshold:

selecting a content item based on an action identified by the voice input and based on a profile of the particular registered user; and

causing the content item to be rendered in response to the audio-based input and to be rendered at the computing device or an additional computing device that is separate from, but associated with, the computing device.

2. The method of claim 1 , further comprising:

determining the distance, between the computing device and the particular registered user, based on a separation distance between the computing device and the additional computing device.

3. The method of claim 2 , wherein the additional computing device is a smartphone of the user.

4. The method of claim 1 , wherein causing the content item to be rendered comprises causing the content item to be rendered at the additional computing device.

5. The method of claim 1 , further comprising:

processing the voice input using a natural language processor (NLP) component to determine the action identified by the voice input.

6. The method of claim 5 , wherein processing the voice input using the NLP component to determine the action identified by the voice input comprises:

converting the voice input into text; and

processing the text to determine the action identified by the voice input.

7. The method of claim 1 , wherein the processing of the audio-based input, in determining that the probability satisfies the threshold, occurs locally at the computing device.

8. A computing device, comprising:

a microphone;

memory storing instructions;

one or more processors, executing the instructions, to:

receive an audio-based input, detected by the microphone, that includes a voice input of a user;

process, using facial recognition, video-based input that is detected at a camera of the computing device;

determine, based on processing of the audio-based input and the processing of the video-based input, that a probability, that the audio-based input was generated by a particular registered user of the computing device, satisfies a threshold;

determine that a distance, between the computing device and the particular registered user, satisfies a distance threshold; and

in response to determining that the probability satisfies the threshold and that the distance satisfies the distance threshold:

select a content item based on an action identified by the voice input and based on a profile of the particular registered user; and

cause the content item to be rendered in response to the audio-based input and to be rendered at the computing device or an additional computing device that is separate from, but associated with, the computing device.

9. The computing device of claim 8 , wherein in executing the instructions one or more of the processors are further to:

determine the distance, between the computing device and the particular registered user, based on a separation distance between the computing device and the additional computing device.

10. The computing device of claim 8 , wherein the additional computing device is a smartphone of the user.

11. The computing device of claim 8 , wherein in causing the content item to be rendered one or more of the processors are to cause the content item to be rendered at the additional computing device.

12. The computing device of claim 8 , wherein in executing the instructions one or more of the processors are further to:

process the voice input using a natural language processor (NLP) component to determine the action identified by the voice input.

13. The computing device of claim 12 , wherein in processing the voice input using the NLP component to determine the action identified by the voice input, one or more of the processors are to:

convert the voice input into text; and

process the text to determine the action identified by the voice input.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2022
From: KRIEGER, KEN; GILDFIND, ANDREW JOSEPH ALEXANDER; ARINI, NICHOLAS SALVATORE; ROWE, SIMON MICHAEL; MIRISOLA, RAIMUNDO; BHAYA, GAURAV; STETS, ROBERT
To: GOOGLE INC.
Reel/Frame 060150/0894 →
CHANGE OF NAME Recorded Jun 9, 2022
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 060328/0009 →
Continuity (6)
Continuation 15862963 · Jan 5, 2018
Continuation 15638316 · Jun 29, 2017
Continuation In Part 15395729 · Dec 30, 2016
Continuation In Part 14933937 · Nov 5, 2015
Continuation In Part 13843559 · Mar 15, 2013
Related Publication 20220237273A1 · Jul 28, 2022