IP Library › Granted Patent US 10,381,022
Granted Patent B1
US 10,381,022 · App. 15/041,379 · Granted Aug 13, 2019

Audio classifier

Inventors: Sourish Chaudhuri (San Francisco, CA); Achal D. Dave (San Diego, CA); Bryan Andrew Seybold (San Francisco, CA)
Assignee: Google LLC
G10L25/57G06F16/638G06K9/00744G10L15/01G10L15/04G10L15/063G10L17/005G10L17/04G10L17/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,381,022
App. No.
15/041,379
Granted
Aug 13, 2019
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for audio classifiers. In one aspect, a method includes obtaining a plurality of video frames from a plurality of videos, wherein each of the plurality of video frames is associated with one or more image labels of a plurality of image labels determined based on image recognition; obtaining a plurality of audio segments corresponding to the plurality of video frames, wherein each audio segment has a specified duration relative to the corresponding video frame; and generating an audio classifier trained using the plurality of audio segment and the associated image labels as input, wherein the audio classifier is trained such that the one or more groups of audio segments are determined to be associated with respective one or more audio labels.

Claims (70)

1. A computer-implemented method for generating an audio classifier that assigns audio labels to input audio segments, the method comprising:

obtaining a plurality of video frames from a plurality of videos, wherein each of the plurality of video frames is unlabeled;

generating, for each of the plurality of video frames, one or more image labels of a plurality of image labels based on image recognition and associating the one or more image labels with each of the plurality of video frames;

obtaining a plurality of unlabeled audio segments corresponding to the plurality of video frames, wherein each unlabeled audio segment has a specified duration relative to the corresponding video frame;

generating the audio classifier, wherein the audio classifier is trained using the plurality of unlabeled audio segments and the associated image labels generated from the plurality of video frames as input training data, and wherein the audio classifier is trained to assign one or more audio labels to unlabeled audio segments input to the audio classifier;

receiving a query, the query including one or more query terms;

using the one or more query terms to identify one or more audio resources responsive to the query using the audio classifier, wherein each audio resource is associated with one or more audio labels generated by the audio classifier based on the query terms; and

providing search results identifying the one or more audio resources.

2. The method of claim 1 , wherein obtaining the plurality of video frames from a plurality of videos comprising:

scoring each of a collection of video frames from the plurality of videos, wherein the each video frame is scored for one or more of the plurality of image labels;

determining a score of the each video frame satisfies a threshold; and

selecting the plurality of video frames in response to determining a score of the each video frame satisfies the threshold.

3. The method of claim 2 , wherein scoring the video frames, from the plurality of videos, for the image label comprising:

selecting video frames periodically from the plurality of video; and

scoring the selected video frames for the image label.

4. The method of claim 1 , wherein the corresponding video frame occurs during the specified duration.

5. The method of claim 1 , wherein the one or more audio labels are determined by using an image classifier.

6. The method of claim 1 , wherein obtaining a plurality of video frames from a plurality of videos further comprising:

identifying an object on video frames from the plurality of videos; and

determining the image label associated with the object.

7. The method of claim 1 , further comprising:

evaluating the audio classifier using sample videos having known an audio label for the sample videos.

8. A computer-implemented method for providing search results identifying audio resources, the method comprising:

receiving a query, the query including one or more query terms;

using the one or more query terms to identify one or more audio resources responsive to the query, wherein each audio resource is associated with one or more audio labels generated by an audio classifier, wherein the audio classifier is trained using a plurality of unlabeled audio segments extracted from a plurality of video frames that are obtained from a plurality of videos and corresponding image labels respectively generated for the plurality of video frames as input training data, and wherein the audio classifier is trained to assign one or more audio labels to unlabeled audio segments input to the audio classifier; and

providing search results identifying the one or more audio resources.

9. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for generating an audio classifier that assigns audio labels to input audio segments, the operations comprising:

obtaining a plurality of video frames from a plurality of videos, wherein each of the plurality of video frames is unlabeled;

generating, for each of the plurality of video frames, one or more image labels of a plurality of image labels based on image recognition and associating the one or more image labels with each of the plurality of video frames;

obtaining a plurality of unlabeled audio segments corresponding to the plurality of video frames, wherein each unlabeled audio segment has a specified duration relative to the corresponding video frame;

generating the audio classifier, wherein the audio classifier is trained using the plurality of unlabeled audio segments and the associated image labels generated from the plurality of video frames as input training data, and wherein the audio classifier is trained to assign one or more audio labels to unlabeled audio segments input to the audio classifier;

receiving a query, the query including one or more query terms;

using the one or more query terms to identify one or more audio resources responsive to the query using the audio classifier, wherein each audio resource is associated with one or more audio labels generated by the audio classifier based on the query terms; and

providing search results identifying the one or more audio resources.

10. The system of claim 9 , wherein obtaining the plurality of video frames from a plurality of videos comprising:

scoring each of a collection of video frames from the plurality of videos, wherein the each video frame is scored for one or more of the plurality of image labels;

determining a score of the each video frame satisfies a threshold; and

selecting the plurality of video frames in response to determining a score of the each video frame satisfies the threshold.

11. The system of claim 10 , wherein scoring the video frames, from the plurality of videos, for the image label comprising:

selecting video frames periodically from the plurality of video; and

scoring the selected video frames for the image label.

12. The system of claim 9 , wherein the corresponding video frame occurs during the specified duration.

13. The system of claim 9 , wherein the one or more audio labels are determined by using an image classifier.

14. The system of claim 9 , wherein obtaining a plurality of video frames from a plurality of videos further comprising:

identifying an object on video frames from the plurality of videos; and

determining the image label associated with the object.

15. The system of claim 9 , further comprising:

evaluating the audio classifier using sample videos having known an audio label for the sample videos.

16. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for providing search results identifying audio resources, the operations comprising:

receiving a query, the query including one or more query terms;

using the one or more query terms to identify one or more audio resources responsive to the query, wherein each audio resource is associated with one or more audio labels generated by an audio classifier, wherein the audio classifier is trained using a plurality of unlabeled audio segments extracted from a plurality of video frames that are obtained from a plurality of videos and corresponding image labels respectively generated for the plurality of video frames as input training data, and wherein the audio classifier is trained to assign one or more audio labels to unlabeled audio segments input to the audio classifier; and

providing search results identifying the one or more audio resources.

17. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations for generating an audio classifier that assigns audio labels to input audio segments, the operations comprising:

obtaining a plurality of video frames from a plurality of videos, wherein each of the plurality of video frames is unlabeled;

generating, for each of the plurality of video frames, one or more image labels of a plurality of image labels based on image recognition and associating the one or more image labels with each of the plurality of video frames;

obtaining a plurality of unlabeled audio segments corresponding to the plurality of video frames, wherein each audio segment has a specified duration relative to the corresponding video frame;

generating the audio classifier, wherein the audio classifier is trained using the plurality of unlabeled audio segments and the associated image labels generated from the plurality of video frames as input training data, and wherein the audio classifier is trained to assign one or more audio labels to unlabeled audio segments input to the audio classifier;

receiving a query, the query including one or more query terms;

using the one or more query terms to identify one or more audio resources responsive to the query using the audio classifier, wherein each audio resource is associated with one or more audio labels generated by the audio classifier based on the query terms; and

providing search results identifying the one or more audio resources.

18. The non-transitory computer-readable medium of claim 17 , wherein the corresponding video frame occurs during the specified duration.

19. The non-transitory computer-readable medium of claim 17 , wherein the image label is determined by using an image classifier.

20. The non-transitory computer-readable medium of claim 17 , further comprising:

evaluating the audio classifier using sample videos having known an audio label for the sample videos.

21. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations for providing search results identifying audio resources, the operations comprising:

receiving a query, the query including one or more query terms;

using the one or more query terms to identify one or more audio resources responsive to the query, wherein each audio resource is associated with one or more audio labels generated by an audio classifier, wherein the audio classifier is trained using a plurality of unlabeled audio segments extracted from a plurality of video frames that are obtained from a plurality of videos and corresponding image labels respectively generated for the plurality of video frames as input training data, and wherein the audio classifier is trained to assign one or more audio labels to unlabeled audio segments input to the audio classifier; and

providing search results identifying the one or more audio resources.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2016
From: CHAUDHURI, SOURISH; DAVE, ACHAL D.; SEYBOLD, BRYAN ANDREW
To: GOOGLE INC.
Reel/Frame 038293/0144 →
Continuity (1)
Provisional Application 62387297 · Dec 23, 2015
Cited By (2)
US 12,430,895 US 12,494,228