IP Library Granted Patent US 11,132,997
Granted Patent B1
US 11,132,997 · App. 16/140,538 · Granted Sep 28, 2021

Robust audio identification with interference cancellation

Inventors: Jose Pio Pereira (Cupertino, CA); Sunil Suresh Kulkarni (San Jose, CA); Mihailo M. Stojancic (San Jose, CA); Shashank Merchant (Sunnyvale, CA); Peter Wendt (San Jose, CA)
Assignee: Roku, Inc.
G10L15/20G10L15/02G10L15/063G10L15/10G10L15/142G10L21/0232G10L25/81G10L2015/025G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,132,997
App. No.
16/140,538
Granted
Sep 28, 2021
Kind
B1
Abstract

Audio distortion compensation methods to improve accuracy and efficiency of audio content identification are described. The method is also applicable to speech recognition. Methods to detect the interference from speakers and sources, and distortion to audio from environment and devices are discussed. Additional methods to detect distortion to the content after performing search and correlation are illustrated. The causes of actual distortion at each client are measured and registered and learnt to generate rules for determining likely distortion and interference sources. The learnt rules are applied at the client, and likely distortions that are detected are compensated or heavily distorted sections are ignored at audio level or signature and feature level based on compute resources available. Further methods to subtract the likely distortions in the query at both audio level and after processing at signature and feature level are described.

Claims (44)

1. A method of audio speech recognition, comprising:

receiving an audio signal;

generating a spectrogram of the audio signal and a plurality of transform coefficients;

processing, using at least one of a hidden Markov model (HMM) or a pattern matching model, the plurality of transform coefficients to detect at least one of a first plurality of speech phonemes or human speech;

determining, based on the at least one of the first plurality of speech phonemes or human speech, one or more components of the spectrogram and one or more of the plurality of transform coefficients associated with likely interfering speech from an interference source;

subtracting, from the spectrogram, the one or more components associated with the likely interfering speech to form a clean audio signal; and

processing, using the at least one of the HMM or the pattern matching model, the clean audio signal to detect a second plurality of speech phonemes.

2. The method of claim 1 , wherein processing the clean audio signal comprises computing the HMM on the clean audio signal to detect the second plurality of speech phonemes.

3. The method of claim 2 , further comprising performing, using the HMM, speech recognition on the second plurality of speech phonemes.

4. The method of claim 1 , wherein processing the clean audio signal comprises computing a state based model on the clean audio signal to detect the second plurality of speech phonemes.

5. The method of claim 1 , wherein processing the clean audio signal comprises computing the pattern matching model on the clean audio signal to detect the second plurality of speech phonemes.

6. The method of claim 1 , wherein the spectrogram and the plurality of transform coefficients include a fundamental frequency and a plurality of formants.

7. The method of claim 1 , further comprising synthesizing the likely interfering speech from the spectrogram and the plurality of transform coefficients of the interference source.

8. The method of claim 1 , further comprising:

identifying, from the audio signal and using a plurality of microphones, a plurality of audio sources; and

separating the plurality of audio sources from audio signal.

9. The method of claim 1 , further comprising segmenting the audio signal into a plurality of separate sources for (i) speech and (ii) music.

10. A non-transitory machine-readable medium having instructions embodied thereon, which, when executed by one or more processors of a machine, cause the machine to perform operations comprising:

receiving an audio signal;

generating a spectrogram of the audio signal;

processing, using at least one of a hidden Markov model (HMM) or a pattern matching model, the plurality of transform coefficients to detect at least one of a first plurality of speech phonemes or human speech;

determining, based on the at least one of the first plurality of speech phonemes or human speech, one or more components of the spectrogram and one or more of the plurality of transform coefficients associated with likely interfering speech from an interference source;

subtracting, from the spectrogram, the one or more components associated with the likely interfering speech to form a clean audio signal; and

processing, using the at least one of the HMM or the pattern matching model, the clean audio signal to detect a second plurality of speech phonemes.

11. The non-transitory machine-readable medium of claim 10 , wherein processing the clean audio signal comprises computing the HMM on the clean audio signal to detect the second plurality of speech phonemes.

12. The non-transitory machine-readable medium of claim 11 , further comprising performing, using the HMM, speech recognition on the second plurality of speech phonemes.

13. The non-transitory machine-readable medium of claim 10 , wherein processing the clean audio signal comprises computing a state based model on the clean audio signal to detect the second plurality of speech phonemes.

14. The non-transitory machine-readable medium of claim 10 , wherein processing the clean audio signal comprises computing the pattern matching model on the clean audio signal to detect the second plurality of speech phonemes.

15. The non-transitory machine-readable medium of claim 10 , wherein the spectrogram and the plurality of transform coefficients include a fundamental frequency and a plurality of formants.

16. The non-transitory machine-readable medium of claim 10 , further comprising synthesizing the likely interfering speech from the spectrogram and the plurality of transform coefficients of the interference source.

17. The non-transitory machine-readable medium of claim 10 , further comprising:

identifying, from the audio signal and using a plurality of microphones, a plurality of audio sources; and

separating the plurality of audio sources from audio signal.

18. A speech recognition system, comprising:

a receiver configured to receive an audio signal;

a memory that stores instructions; and

one or more processors configured by the instructions to perform operations comprising:

generating a spectrogram of the audio signal;

processing, using at least one of a hidden Markov model (HMM) or a pattern matching model, the plurality of transform coefficients to detect at least one of a first plurality of speech phonemes or human speech;

determining, based on the at least one of the first plurality of speech phonemes or human speech, one or more components of the spectrogram and one or more of the plurality of transform coefficients associated with likely interfering speech from an interference source;

subtracting, from the spectrogram, the one or more components associated with the likely interfering speech to form a clean audio signal; and

processing, using the at least one of the HMM or the pattern matching model, the clean audio signal to detect a second plurality of speech phonemes.

19. The speech recognition system of claim 18 , wherein processing the clean audio signal comprises computing the HMM on the clean audio signal to detect the second plurality of speech phonemes.

20. The speech recognition system of claim 18 , wherein processing the clean audio signal comprises computing the pattern matching model on the clean audio signal to detect the second plurality of speech phonemes.

Assignments (9)
SECURITY INTEREST Recorded Sep 18, 2024
From: ROKU, INC.
To: CITIBANK, N.A.
Reel/Frame 068982/0377 →
RELEASE (REEL 053473 / FRAME 0001) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063603/0001 →
RELEASE (REEL 054066 / FRAME 0064) Recorded May 11, 2023
From: CITIBANK, N.A.
To: A. C. NIELSEN COMPANY, LLC; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE MEDIA SERVICES, LLC; THE NIELSEN COMPANY (US), LLC; NETRATINGS, LLC
Reel/Frame 063605/0001 →
TERMINATION AND RELEASE OF INTELLECTUAL PROPERTY SECURITY AGREEMENT (REEL/FRAME 056982/0194) Recorded Feb 22, 2023
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: ROKU, INC.; ROKU DX HOLDINGS, INC.
Reel/Frame 062826/0664 →
PATENT SECURITY AGREEMENT SUPPLEMENT Recorded Jun 29, 2021
From: ROKU, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 056982/0194 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 30, 2021
From: GRACENOTE, INC.
To: ROKU, INC.
Reel/Frame 056103/0786 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PATENTS LISTED ON SCHEDULE 1 RECORDED ON 6-9-2020 PREVIOUSLY RECORDED ON REEL 053473 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE SUPPLEMENTAL IP SECURITY AGREEMENT. Recorded Oct 7, 2020
From: A.C. NIELSEN (ARGENTINA) S.A.; A.C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A
Reel/Frame 054066/0064 →
SUPPLEMENTAL SECURITY AGREEMENT Recorded Jun 9, 2020
From: A. C. NIELSEN COMPANY, LLC; ACN HOLDINGS INC.; ACNIELSEN CORPORATION; ACNIELSEN ERATINGS.COM; AFFINNOVA, INC.; ART HOLDING, L.L.C.; ATHENIAN LEASING CORPORATION; CZT/ACN TRADEMARKS, L.L.C.; EXELATE, INC.; GRACENOTE, INC.; GRACENOTE DIGITAL VENTURES, LLC; GRACENOTE MEDIA SERVICES, LLC; NETRATINGS, LLC; NIELSEN AUDIO, INC.; NIELSEN CONSUMER INSIGHTS, INC.; NIELSEN CONSUMER NEUROSCIENCE, INC.; NIELSEN FINANCE CO.; NIELSEN FINANCE LLC; NIELSEN INTERNATIONAL HOLDINGS, INC.; NIELSEN MOBILE, LLC; NIELSEN UK FINANCE I, LLC; NMR INVESTING I, INC.; TCG DIVESTITURE INC.; TNC (US) HOLDINGS, INC.; THE NIELSEN COMPANY (US), LLC; VIZU CORPORATION; VNU MARKETING INFORMATION, INC.; NMR LICENSING ASSOCIATES, L.P.; NIELSEN HOLDING AND FINANCE B.V.; THE NIELSEN COMPANY B.V.; VNU INTERNATIONAL B.V.
To: CITIBANK, N.A.
Reel/Frame 053473/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2019
From: PEREIRA, JOSE PIO; KULKARNI, SUNIL SURESH; STOJANCIC, MIHAILO M.; MERCHANT, SHASHANK; WENDT, PETER
To: GRACENOTE, INC.
Reel/Frame 050182/0630 →
Continuity (4)
Continuation 15456859 · Mar 13, 2017
Provisional Application 62306719 · Mar 11, 2016
Provisional Application 62306707 · Mar 11, 2016
Provisional Application 62306700 · Mar 11, 2016