IP Library Granted Patent US 11,664,017
Granted Patent B2
US 11,664,017 · App. 17/047,472 · Granted May 30, 2023

Systems and methods for identifying and providing information about semantic entities in audio signals

Inventors: Tim Wantland (Bellevue, WA); Brandon Barbello (Mountain View, CA)
Assignee: GOOGLE LLC
G10L15/1815G06F3/017G06F3/0481G06F3/167G06F16/685G06N20/00G10L15/22G10L15/30H04R1/08H04R1/1016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,664,017
App. No.
17/047,472
Granted
May 30, 2023
Kind
B2
Abstract

Systems and methods for determining identifying semantic entities in audio signals are provided. A method can include obtaining, by a computing device comprising one or more processors and one or more memory devices, an audio signal concurrently heard by a user. The method can further include analyzing, by a machine-learned model stored on the computing device, at least a portion of the audio signal in a background of the computing device to determine one or more semantic entities. The method can further include displaying the one or more semantic entities on a display screen of the computing device.

Claims (44)

1. A method for identifying semantic entities within an audio signal, comprising:

obtaining, by a computing system comprising one or more processors and one or more memory devices, an audio signal output by an output device;

receiving, by the computing system, a request from a user device to display one or more semantic entities;

determining, by the computing system, a selected portion of the audio signal for analysis based at least in part on a predetermined time period preceding receipt of the request from the user to display the one or more semantic entities;

analyzing, by a machine-learned model stored on the computing system, the selected portion of the audio signal to determine the one or more semantic entities; and

in response to receiving the request, displaying the one or more semantic entities on a display screen of the user device.

2. The method of claim 1 , further comprising:

receiving, by the computing system, a user selection of a selected semantic entity;

wherein the selected semantic entity comprises one of the one or more semantic entities displayed on display screen of the user device.

3. The method of claim 2 , further comprising:

determining, by the computing system, one or more supplemental information options associated with the selected semantic entity; and

displaying the one or more supplemental information options associated with the selected semantic entity on the display screen of the user device.

4. The method of claim 3 , wherein the one or more supplemental information options are determined based at least in part on a context of the audio signal or a context of the selected semantic entity.

5. The method of claim 4 , wherein the one or more supplemental information options are determined based at least in part on a classification of a type of the audio signal.

6. The method of claim 1 , wherein the user device comprises the output device.

7. The method of claim 1 , wherein the user device is associated with a user account, and wherein the output device is associated with the user account.

8. The method of claim 1 , wherein the selected portion is analyzed using the machine-learned model responsive to receiving the request.

9. The method of claim 1 , wherein the selected portion is pre-processed using the machine-learned model prior to receiving the request.

10. The method of claim 1 , comprising:

analyzing, using the machine-learned model, a rolling buffer of the audio signal to obtain data descriptive of the one or more semantic entities.

11. A computing system for identifying semantic entities within an audio signal, comprising:

one or more processors; and

one or more non-transitory computer-readable media storing instructions that are executable by the one or more processors to cause the computing system to perform operations, the operations comprising:

obtaining an audio signal output by an output device;

receiving a request from a user device to display one or more semantic entities;

determining a selected portion of the audio signal for analysis based at least in part on a predetermined time period preceding receipt of the request from the user to display the one or more semantic entities;

analyzing, by a machine-learned model, the selected portion of the audio signal to determine the one or more semantic entities; and

generating data descriptive of the one or more semantic entities for output by the user device.

12. The computing system of claim 11 , wherein the operations comprise:

determining one or more supplemental information options based at least in part on a context of the audio signal obtained by the computing system or a context of a selected semantic entity of the one or more semantic entities.

13. The computing system of claim 12 , wherein the one or more supplemental information options are determined based at least in part on a classification of a type of the audio signal.

14. The computing system of claim 11 , wherein the user device comprises the output device.

15. The computing system of claim 11 , wherein the user device is associated with a user account, and wherein the output device is associated with the user account.

16. The computing system of claim 11 , wherein the selected portion is analyzed using the machine-learned model responsive to receiving the request.

17. The computing system of claim 11 , wherein the selected portion is pre-processed using the machine-learned model prior to receiving the request.

18. The computing system of claim 11 , wherein the operations comprise:

analyzing, using the machine-learned model, a rolling buffer of the audio signal to obtain data descriptive of the one or more semantic entities.

19. The computing system of claim 11 , wherein the computing system is a server computing system external to the user device.

20. One or more non-transitory computer-readable media storing instructions that are executable by one or more processors to cause a computing system to perform operations, the operations comprising:

obtaining an audio signal output by an output device;

receiving a request from a user device to display one or more semantic entities;

determining a selected portion of the audio signal for analysis based at least in part on a predetermined time period preceding receipt of the request from the user to display the one or more semantic entities;

analyzing, by a machine-learned model, the selected portion of the audio signal to determine the one or more semantic entities; and

generating data descriptive of the one or more semantic entities for output by the user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2020
From: WANTLAND, TIM; BARBELLO, BRANDON
To: GOOGLE LLC
Reel/Frame 054049/0944 →
Continuity (1)
Related Publication 20210142792A1 · May 13, 2021