IP Library Granted Patent US 8,948,894
Granted Patent B2
US 8,948,894 · App. 13/186,603 · Granted Feb 3, 2015

Method of selectively inserting an audio clip into a primary audio stream

Inventor: Dinkar N. Bhat (Princeton, NJ)
Assignee: Google Technology Holdings LLC
G11B27/038G11B27/036
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,948,894
App. No.
13/186,603
Granted
Feb 3, 2015
Kind
B2
Abstract

A method of processing audio signals in which an audio signature of a primary audio stream is compared to a set of audio signatures corresponding to a series of separate audio clips is disclosed. One of the audio clips from the series of separate audio clips is selected for insertion into the primary audio stream such that the selected audio clip has an audio signature that most closely matches the audio signature of the primary audio stream. The matching and selecting steps are performed by at least one signal processing electronic device.

Claims (24)

1. A method of processing audio signals, comprising:

computing Mel Frequency Cepstral Coefficients (MFCC) feature vectors of a primary audio stream and of a series of separate audio clips;

forming ranked vectors from the MFCC feature vectors of the primary audio stream and the series of separate audio clips, wherein the ranked vectors respectively rank coefficients on an ordinal scale within each of the MFCC feature vectors;

comparing the ranked vector of the primary audio stream with the ranked vectors of the series of separate audio clips; and

selecting an audio clip for insertion into the primary audio stream from the series of separate audio clips based on results of said comparing step such that a selected audio clip has an audio signature that most closely matches an audio signature of the primary audio stream;

said computing, forming, comparing, and selecting steps being performed by at least one signal processing electronic device.

2. A method according to claim 1 , further comprising

extracting audio segments from the primary audio stream during a time period closely before an audio clip is to be inserted into the primary audio stream;

said computing step computing the MFCC feature vector of the primary audio stream from the audio segments.

3. A method according to claim 2 , wherein in said selecting the audio clip having an audio signature that most closely matches the audio signature of the primary audio stream is the audio clip having one of the plurality of ranked MFCC feature vectors coefficients most similar in order to the ranked MFCC feature vector coefficients of the primary audio stream.

4. A method according to claim 2 , wherein, said step of computing the MFCC feature vector of the primary audio stream includes:

computing Fast Fourier Transform (FFT) for each of the extracted audio segments; applying mel filters and Log transform to the computed FFT for each of the extracted audio segments to produce a list of mel log-amplitudes for each of the extracted audio segments; and

computing Discrete Cosine Transform (DCT) of each of the lists of mel log-amplitudes to produce MFCC feature vectors for each of the extracted audio segments.

5. A method according to claim 4 , further comprising discarding a first listed coefficient of each of the MFCC vectors; and

after said discarding computing an average of each of the MFCC vectors to produce the MFCC feature vector of the primary audio stream.

6. A method according to claim 4 , wherein the MFCC feature vector of the primary audio stream is determined in substantially real-time by the at least one signal processing electronic device and the MFCC feature vector of each of the audio clips is pre-calculated and stored in memory accessible by the at least one signal processing electronic device.

7. A method according to claim 1 , further comprising splicing the selected audio clip into the primary audio stream.

8. A method according to claim 1 , wherein the at least one signal processing electronic device is selected from a group consisting of consumer premises equipment, a set-top box, a digital television, a personal computer, a desktop computer, a laptop computer, a pad or tablet computer, a media player, and a smart phone.

9. An electronic device for inserting a signal corresponding to an audio clip into a signal corresponding to a primary audio stream, comprising:

at least one signal processing module for computing Mel Frequency Cepstral Coefficients (MFCC) feature vectors of the primary audio stream and of a series of separate audio clips; forming ranked vectors from the MFCC feature vectors of the primary audio stream and the series of separate audio clips, wherein the ranked vectors respectively rank coefficients on an ordinal scale within each of the MFCC feature vectors; comparing the ranked vector of the primary audio stream with the ranked vectors of the series of separate audio clips; and selecting an audio clip for insertion into the primary audio stream from the series of separate audio clips based on results of said comparing step such that a selected audio clip has an audio signature that most closely matches an audio signature of the primary audio stream.

10. An electronic device according to claim 9 , wherein said at least one signal processing module includes a module for defining a size of sliding window of audio segments to be extracted sequentially from the primary audio stream.

11. An electronic device according to claim 9 , wherein said at least one signal processing module includes a decoder for receiving said primary audio stream and for outputting a pulse code modulated audio signal.

12. An electronic device according to claim 9 , further comprising an output router for outputting the primary audio stream with a best matching audio clip inserted therein.

13. An electronic device according to claim 9 , wherein the electronic device is selected from a group consisting of consumer premises equipment, a set-top box, a digital television, a personal computer, a desktop computer, a laptop computer, a pad or tablet computer, a media player, and a smart phone.

Assignments (5)
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE INCORRECT PATENT NO. 8577046 AND REPLACE WITH CORRECT PATENT NO. 8577045 PREVIOUSLY RECORDED ON REEL 034286 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Dec 3, 2014
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 034538/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2014
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 034286/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2013
From: GENERAL INSTRUMENT CORPORATION
To: GENERAL INSTRUMENT HOLDINGS, INC.
Reel/Frame 030764/0575 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2013
From: GENERAL INSTRUMENT HOLDINGS, INC.
To: MOTOROLA MOBILITY LLC
Reel/Frame 030866/0113 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2011
From: BHAT, DINKAR N.
To: GENERAL INSTRUMENT CORPORATION
Reel/Frame 026620/0353 →
Continuity (1)
Related Publication 20130024016A1 · Jan 24, 2013