IP Library Granted Patent US 11,687,585
Granted Patent B2
US 11,687,585 · App. 17/081,097 · Granted Jun 27, 2023

Systems and methods for identifying a media asset from an ambiguous audio indicator

Inventors: Lucas Waye (Cambridge, MA); Theresa Tokesky (Boston, MA); Michael A. Montalto (South Hamilton, MA); Kanagasabai Sivanadian (Natick, MA)
Assignee: Rovi Guides, Inc.
G06F16/634G06F16/435G06F16/487G06F16/90335G06F16/9537H04N21/47202G06Q50/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,687,585
App. No.
17/081,097
Granted
Jun 27, 2023
Kind
B2
Abstract

Systems and methods are disclosed herein for identifying a media asset in response to an ambiguous input. The media guidance application may detect a portion of music provided by a user, e.g., a melody from user humming. The media guidance application may retrieve information about the user's location for a predetermined time period prior to detecting the portion of music. The media guidance application may then determine content accessible by the user at the location, e.g., a commercial played at a display screen at a train station when the user was waiting for the train, to identify the media asset corresponding to the user humming.

Claims (108)

1. A method for identifying a media asset, the method comprising:

capturing, from an environment where a user is present, an audio recording;

retrieving information about a user's location;

identifying a source of media available at the user's location;

determining content accessible from the source of media;

comparing the audio recording to the content to identify the media asset having a characteristic associated with the audio recording by:

determining a plurality of music tunes transmitted with the content;

comparing the frequency domain representation of the plurality of tones with each music tune of the plurality of music tunes to generate a respective similarity metric; and

in response to determining that the respective similarity metric is greater than a similarity threshold, identifying a respective music tune from the media asset corresponds to the audio recording provided by the user;

upon determining a match between the audio recording and the content, generating for display a recommendation of the media asset;

identifying, from the audio recording, a tune having audio characteristics that match with vocal characteristics of the user, wherein the tune includes a plurality of tones;

in response to identifying the tune, generating a frequency domain representation of the plurality of tones; and

transmitting a query based on the generated frequency domain representation to a music database.

2. The method of claim 1 , further comprising:

receiving a result indicating a failure to find a match in the music database based on the generated frequency domain representation.

3. The method of claim 1 , wherein identifying, from the audio recording, the tune having audio characteristics that match with vocal characteristics of the user, comprises:

extracting a set of mono signals from the audio recording;

for each mono signal from the set of mono signals:

generating a set of audio characteristics corresponding to the mono signal, wherein the set of audio characteristics includes any of mel-frequency cepstral coefficients (MFCC) and perceptual linear prediction relative spectra (PLP-RASTA);

retrieving, from a profile of the user, a set of vocal characteristics;

comparing each characteristic of the set of audio characteristics with a corresponding characteristic from the set of vocal characteristics that has a same type;

determining whether the set of audio characteristics and the set of vocal characteristics overlap for more than a similarity threshold; and

in response to determining that the set of audio characteristics and the set of vocal characteristics overlap for more than the similarity threshold, identifying the mono signal as a vocal signal from the user.

4. The method of claim 1 , wherein the determining the plurality of music tunes transmitted with the content of the first media asset comprises:

retrieving metadata associated with the content;

in response to determining that the metadata includes information relating to a theme song:

transmitting a query to the music database based on a title of the theme song; and

in response to the query, obtaining an audio asset of the theme song and generating a tune for the audio asset; and

in response to determining that the metadata includes no information relating to any theme song, performing audio analysis of the content to generate the plurality of music tunes.

5. A method for identifying a media asset, the method comprising:

capturing, from an environment where a user is present, an audio recording;

retrieving information about a user's location including retrieving, from a profile of the user, a location history of the user and an application usage history corresponding to the user;

identifying a source of media available at the user's location;

determining content accessible from the source of media;

comparing the audio recording to the content to identify the media asset having a characteristic associated with the audio recording;

upon determining a match between the audio recording and the content, generating for display a recommendation of the media asset;

searching the application usage history for one or more application usage records corresponding to when the user was present at the user's location; and

in response to identifying the one or more application usage records:

determining an application type, an application usage status and an application usage time duration for each of the one or more application usage records;

searching, an application usage table, for a distraction score corresponding to each application type;

computing a distraction metric based on the distraction score corresponding to each application type, the respective application usage status and the respective application usage time duration; and

in response to determining that the distraction metric is higher than a distraction threshold, determining that the user was not exposed to the retrieved content.

6. The method of claim 5 , wherein the retrieving, from the profile of the user, the location history of the user and the application usage history comprises:

transmitting a query to a device of the user for a GPS log and an application usage log;

in response to the query, obtaining the GPS log from the device of the user;

in response to receiving a notification that the GPS log is unavailable, transmitting, to a server, a query for a record of social media activities relating to the user;

in response to the query for the record of social media activities, searching the record of social media activities relating to the user for a first subset of social media activities, each social media activity from the subset identifying a location; and

storing the first subset of social media activities and corresponding locations as part of the location history.

7. The method of claim 6 , further comprising:

in response to receiving a notification that the application usage history is unavailable from the device of the user:

searching the record of social media activities relating to the user for a second subset of social media activities, each social media activity from the second subset indicates that the user is using a respective application; and

storing the second subset of social media activities and information relating to respective applications as part of the application usage history.

8. A system for identifying a media asset, the system comprising:

a storage device; and

control circuitry configured to:

capture, from an environment where a user is present, an audio recording;

retrieve information about a user's location;

identify a source of media available at the user's location;

determine content accessible from the source of media;

compare the audio recording to the content to identify the media asset having a characteristic associated with the audio recording by:

determining a plurality of music tunes transmitted with the content;

comparing the frequency domain representation of the plurality of tones with each music tune of the plurality of music tunes to generate a respective similarity metric; and

in response to determining that the respective similarity metric is greater than a similarity threshold, identifying a respective music tune from the media asset corresponds to the audio recording provided by the user;

upon determining a match between the audio recording and the content, generate for display a recommendation of the media asset;

identify, from the audio recording, a tune having audio characteristics that match with vocal characteristics of the user, wherein the tune includes a plurality of tones;

in response to identifying the tune, generate a frequency domain representation of the plurality of tones; and

transmit a query based on the generated frequency domain representation to a music database.

9. The system of claim 8 , wherein the control circuitry is further configured to:

receive a result indicating a failure to find a match in the music database based on the generated frequency domain representation.

10. The system of claim 8 , wherein the control circuitry, when identifying, from the audio recording, the tune having audio characteristics that match with vocal characteristics of the user, further configured to:

extract a set of mono signals from the audio recording;

for each mono signal from the set of mono signals:

generate a set of audio characteristics corresponding to the mono signal, wherein the set of audio characteristics includes any of mel-frequency cepstral coefficients (MFCC) and perceptual linear prediction relative spectra (PLP-RASTA);

retrieve, from a profile of the user, a set of vocal characteristics;

compare each characteristic of the set of audio characteristics with a corresponding characteristic from the set of vocal characteristics that has a same type;

determine whether the set of audio characteristics and the set of vocal characteristics overlap for more than a similarity threshold; and

in response to determining that the set of audio characteristics and the set of vocal characteristics overlap for more than the similarity threshold, identify the mono signal as a vocal signal from the user.

11. The system of claim 8 , wherein the control circuitry, when determining the plurality of music tunes transmitted with the content of the first media asset, further configured to:

retrieve metadata associated with the content;

in response to determining that the metadata includes information relating to a theme song:

transmit a query to the music database based on a title of the theme song; and

in response to the query, obtain an audio asset of the theme song and generating a tune for the audio asset; and

in response to determining that the metadata includes no information relating to any theme song, perform audio analysis of the content to generate the plurality of music tunes.

12. A system for identifying a media asset, the system comprising:

a storage device; and

control circuitry configured to:

capture, from an environment where a user is present, an audio recording;

retrieve information about a user's location including retrieving, from a profile of the user, a location history of the user and an application usage history corresponding to the user;

identify a source of media available at the user's location;

determine content accessible from the source of media;

compare the audio recording to the content to identify the media asset having a characteristic associated with the audio recording; and

upon determining a match between the audio recording and the content, generate for display a recommendation of the media asset;

search the application usage history for one or more application usage records corresponding to when the user was present at the user's location; and

in response to identifying the one or more application usage records:

determine an application type, an application usage status and an application usage time duration for each of the one or more application usage records;

search, an application usage table, for a distraction score corresponding to each application type;

compute a distraction metric based on the distraction score corresponding to each application type, the respective application usage status and the respective application usage time duration; and

in response to determining that the distraction metric is higher than a distraction threshold, determine that the user was not exposed to the retrieved content.

13. The system of claim 12 , wherein the retrieving, from the profile of the user, the location history of the user and the application usage history comprises:

transmit a query to a device of the user for a GPS log and an application usage log;

in response to the query, obtain the GPS log from the device of the user;

in response to receiving a notification that the GPS log is unavailable, transmit, to a server, a query for a record of social media activities relating to the user;

in response to the query for the record of social media activities, search the record of social media activities relating to the user for a first subset of social media activities, each social media activity from the subset identifying a location; and

store the first subset of social media activities and corresponding locations as part of the location history.

14. The system of claim 13 , wherein the control circuitry is further configured to:

in response to receive a notification that the application usage history is unavailable from the device of the user:

search the record of social media activities relating to the user for a second subset of social media activities, each social media activity from the second subset indicates that the user is using a respective application; and

store the second subset of social media activities and information relating to respective applications as part of the application usage history.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0129 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2020
From: WAYE, LUCAS; TOKESKY, THERESA; MONTALTO, MICHAEL A.; SIVANADIAN, KANAGASABAI
To: ROVI GUIDES, INC.
Reel/Frame 054180/0078 →
Continuity (2)
Continuation 15947345 · Apr 6, 2018
Related Publication 20210224316A1 · Jul 22, 2021