IP Library › Granted Patent US 12,499,099
Granted Patent B2
US 12,499,099 · App. 18/422,624 · Granted Dec 16, 2025

Systems and methods for interpreting natural language search queries using training data

Inventors: Jeffry Copps Robert Jose (Chennai, IN); Ajay Kumar Mishra (Karnataka, IN)
Assignee: Adeia Guides Inc.
G06F16/2228G06F16/243G06F16/2455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,099
App. No.
18/422,624
Granted
Dec 16, 2025
Kind
B2
Abstract

A frequency of occurrence for each term in a training data set is determined in relation to the entire training data set. A relational data structure is generated that associates each term in the training data with its respective frequency. Any term that has a frequency below a threshold frequency is then added to a list of relevant words. When a natural language search query is received, a plurality of terms in the natural language search query are identified and compared with the list of relevant words. If any term of the natural language search query is included in the relevant words list, that term is identified as a keyword. The natural language search query is then interpreted based on any identified keywords.

Claims (88)

1 . A method comprising:

determining a frequency of a first term based at least in part on a number of occurrences of the first term within a totality of training data, wherein the training data comprises prior search queries and a plurality of metadata each corresponding to a content item of a plurality of content items;

associating the first term with a part of speech based at least in part on analyzing a usage of the first term within the totality of training data;

based at least in part on (1) determining that the frequency of the first term within the training data is below a threshold frequency and (2) associating the first term with the part of speech, adding the first term to a relevancy list;

receiving a natural language search query comprising a plurality of terms, wherein the plurality of terms comprises the first term;

determining, for each term of the plurality of terms, whether the term is included in the relevancy list;

in response to determining that the first term of the plurality of terms is included in the relevancy list, identifying the term as a keyword of the natural language search query;

interpreting the natural language search query based at least in part on the identified keyword; and

performing a search for the natural language search query based at least in part on the identified keyword; and

generating a selectable icon identifying a content item of the plurality of content items, wherein the metadata corresponding to the content item comprises the identified keyword.

2 . The method of claim 1 , wherein the frequency of the first term is stored in a relation data structure that associates each term identified in the totality of training data with its respective frequency, the method comprising:

creating a data structure comprising at least a token field and a corresponding value field;

adding to the data structure, for each term identified in the totality of training data, the respective term as a token; and

setting a value corresponding to the token to the frequency of the respective term.

3 . The method of claim 1 , wherein the natural language search query is received as audio data, the method further comprising transcribing the natural language search query into a plurality of words.

4 . The method of claim 1 , wherein the receiving the natural language search query further comprises:

analyzing individual words of the natural language search query to determine whether a first word and second word should be identified as a single term.

5 . The method of claim 4 , wherein the analyzing individual words of the natural language search query comprises:

splitting the natural language search query into a plurality of words comprising the first word and the second word;

analyzing the first word of the plurality of words;

determining, based at least in part on analyzing the first word, whether the first word can be part of a phrase;

in response to determining that the first word can be part of a phrase, analyzing the first word together with the second word that immediately follows the first word;

determining, based at least in part on analyzing the first word together with the second word, whether the first word and the second word form a phrase together;

in response to determining that the first word and the second word form a phrase together, identifying the first word and the second word as the single term.

6 . The method of claim 1 , wherein the determining the frequency of the first term based at least in part on the number of occurrences of the first term within the totality of training data comprises:

counting the total number of words contained in the totality of training data;

determining the number of occurrences of the first term in the totality of training data; and

calculating a percentage of the number of occurrences of the first term in the totality of training data with respect to the total number of words contained in the totality of training data.

7 . A system comprising:

a user interface configured to receive a natural language search query;

control circuitry, communicatively coupled to the user interface, configured to:

determine a frequency of a first term based at least in part on a number of occurrences of the first term within a totality of training data, wherein the training data comprises prior search queries and a plurality of metadata each corresponding to a content item of a plurality of content items;

associate the first term with a part of speech based at least in part on analyzing a usage of the first term within the totality of training data;

based at least in part on (1) determining that the frequency of the first term is below a threshold frequency and (2) associating the first term with the part of speech, add the first term to a relevancy list;

receive the natural language search query comprising a plurality of terms, wherein the plurality of terms comprises the first term;

determine, for each term of the plurality of terms, whether the term is included in the relevancy list;

in response to determining that the first term of the plurality of terms is included in the relevancy list, identify the term as a keyword of the natural language search query;

interpret the natural language search query based at least in part on the identified keyword; and

perform a search for the natural language search query based at least in part on the identified keyword; and

generate a selectable icon identifying a content item of the plurality of content items, wherein the metadata corresponding to the content item comprises the identified keyword.

8 . The system of claim 7 , wherein the frequency of the first term is stored in a relation data structure that associates each term identified in the totality of training data with its respective frequency, wherein the control circuitry is further configured to:

create a data structure comprising at least a token field and a corresponding value field;

add to the data structure, for each term identified in the totality of training data, the respective term as a token; and

set a value corresponding to the token to the frequency of the respective term.

9 . The system of claim 7 , wherein the natural language search query is received as audio data, and wherein the control circuitry is further configured to transcribe the natural language search query into a plurality of words.

10 . The system of claim 7 , wherein the control circuitry configured to receive the natural language search query is further configured to:

analyze individual words of the natural language search query to determine whether a first word and second word should be identified as a single term.

11 . The system of claim 10 , wherein the control circuitry is configured to analyze individual words of the natural language search query by:

splitting the natural language search query into a plurality of words comprising the first word and the second word;

analyzing the first word of the plurality of words;

determining, based at least in part on analyzing the first word, whether the first word can be part of a phrase;

in response to determining that the first word can be part of a phrase, analyzing the first word together with the second word that immediately follows the first word;

determining, based at least in part on analyzing the first word together with the second word, whether the first word and the second word form a phrase together;

in response to determining that the first word and the second word form a phrase together, identifying the first word and the second word as the single term.

12 . The system of claim 7 , wherein the control circuitry is configured to determine the frequency of the first term based at least in part on the number of occurrences of the first term within the totality of training data by:

counting the total number of words contained in the totality of training data;

determining the number of occurrences of the first term in the totality of training data; and

calculating a percentage of the number of occurrences of the first term in the totality of training data with respect to the total number of words contained in the totality of training data.

13 . A non-transitory computer-readable medium having non-transitory computer-readable instructions encoded thereon, wherein the non-transitory computer-readable instructions, when executed by control circuitry, cause the control circuitry to:

determine a frequency of a first term based at least in part on a number of occurrences of the first term within a totality of training data, wherein the training data comprises prior search queries and a plurality of metadata each corresponding to a content item of a plurality of content items;

associate the first term with a part of speech based at least in part on analyzing a usage of the first term within the totality of training data;

based at least in part on (1) determining that the frequency of the first term is below a threshold frequency and (2) associating the first term with the part of speech, add the first term to a relevancy list;

receive a natural language search query comprising a plurality of terms, wherein the plurality of terms comprises the first term;

determine, for each term of the plurality of terms, whether the term is included in the relevancy list;

in response to determining that the first term of the plurality of terms is included in the relevancy list, identify the term as a keyword of the natural language search query;

interpret the natural language search query based at least in part on the identified keyword; and

perform a search for the natural language search query based at least in part on the identified keyword; and

generate a selectable icon identifying a content item of the plurality of content items, wherein the metadata corresponding to the content item comprises the identified keyword.

14 . The non-transitory computer-readable medium of claim 13 , wherein the frequency of the first term is stored in a relation data structure that associates each term identified in the totality of training data with its respective frequency, further comprising instructions that when executed by the control circuitry cause the control circuitry to:

create a data structure comprising at least a token field and a corresponding value field;

add to the data structure, for each term identified in the totality of training data, the respective term as a token; and

set a value corresponding to the token to the frequency of the respective term.

15 . The non-transitory computer-readable medium of claim 13 , wherein:

the natural language search query is received as audio data; and

the non-transitory computer-readable medium further comprises instructions that when executed by the control circuitry cause the control circuitry to transcribe the natural language search query into a plurality of words.

16 . The non-transitory computer-readable medium of claim 13 , wherein the instructions that cause the control circuitry to receive the natural language search query further cause the control circuitry to:

analyze individual words of the natural language search query to determine whether a first word and second word should be identified as a single term.

17 . The non-transitory computer-readable medium of claim 16 , wherein the instructions that cause the control circuitry to analyze individual words of the natural language search query further cause the control circuitry to:

split the natural language search query into a plurality of words comprising the first word and the second word;

analyze the first word of the plurality of words;

determine, based at least in part on analyzing the first word, whether the first word can be part of a phrase;

in response to determining that the first word can be part of a phrase, analyze the first word together with the second word that immediately follows the first word;

determine, based at least in part on analyzing the first word together with the second word, whether the first word and the second word form a phrase together;

in response to determining that the first word and the second word form a phrase together, identify the first word and the second word as the single term.

18 . The non-transitory computer-readable medium of claim 13 , wherein the instructions that cause the control circuitry to determine the frequency of the first term based at least in part on the number of occurrences of the first term within the totality of training data further cause the control circuitry to:

count the total number of words contained in the totality of training data;

determine the number of occurrences of the first term in the totality of training data; and

calculate a percentage of the number of occurrences of the first term in the totality of training data with respect to the total number of words contained in the totality of training data.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2024
From: ROBERT JOSE, JEFFRY COPPS; MISHRA, AJAY KUMAR
To: ROVI GUIDES, INC.
Reel/Frame 066257/0382 →
Continuity (2)
Continuation 16807419 · Mar 3, 2020
Related Publication 20240160613A1 · May 16, 2024
References Cited (82)
US 5146405A · Church · 1992 [cited by applicant]
US 5852801A · Hon et al. · 1998 [cited by applicant]
US 6307548B1 · Flinchem et al. · 2001 [cited by applicant]
US 6523026B1 · Gillis · 2003 [cited by examiner]
US 6615172B1 · Bennett et al. · 2003 [cited by applicant]
US 6721697B1 · Duan et al. · 2004 [cited by applicant]
US 6901399B1 · Corston et al. · 2005 [cited by applicant]
US 7536408B2 · Patterson · 2009 [cited by applicant]
US 8484017B1 · Sharifi · 2013 [cited by examiner]
US 8996994B2 · Alonichau et al. · 2015 [cited by applicant]
US 10515125B1 · Lavergne · 2019 [cited by applicant]
US 10769371B1 · Barrientos et al. · 2020 [cited by applicant]
US 10860631B1 · Huang et al. · 2020 [cited by applicant]
US 10949483B2 · Tseng et al. · 2021 [cited by applicant]
US 11507572B2 · Robert Jose et al. · 2022 [cited by applicant]
US 11594213B2 · Robert Jose et al. · 2023 [cited by applicant]
US 11734512B1 · Feyisetan · 2023 [cited by applicant]
US 11914561B2 · Robert Jose et al. · 2024 [cited by applicant]
US 12062366B2 · Robert Jose et al. · 2024 [cited by applicant]
US 20030078766A1 · Appelt et al. · 2003 [cited by applicant]
US 20040024739A1 · Copperman · 2004 [cited by examiner]
US 20050137843A1 · Lux · 2005 [cited by applicant]
US 20070016862A1 · Kuzmin · 2007 [cited by applicant]
US 20070033221A1 · Copperman et al. · 2007 [cited by applicant]
US 20080104056A1 · Li et al. · 2008 [cited by applicant]
US 20080208566A1 · Alonichau · 2008 [cited by applicant]
US 20080221890A1 · Kurata et al. · 2008 [cited by applicant]
US 20090070300A1 · Bartels et al. · 2009 [cited by applicant]
US 20090109067A1 · Burstrom · 2009 [cited by applicant]
US 20090150152A1 · Wasserblat et al. · 2009 [cited by applicant]
US 20090171945A1 · Li · 2009 [cited by examiner]
US 20090216737A1 · Dexter · 2009 [cited by applicant]
US 20090287678A1 · Brown et al. · 2009 [cited by applicant]
US 20100185661A1 · Malden et al. · 2010 [cited by applicant]
US 20100306144A1 · Scholz · 2010 [cited by examiner]
US 20110161341A1 · Johnston · 2011 [cited by applicant]
US 20110213761A1 · Song · 2011 [cited by examiner]
US 20120035932A1 · Jitkoff et al. · 2012 [cited by applicant]
US 20120084312A1 · Jenson · 2012 [cited by applicant]
US 20120136649A1 · Freising et al. · 2012 [cited by applicant]
US 20130021346A1 · Terman · 2013 [cited by applicant]
US 20130262361A1 · Arroyo et al. · 2013 [cited by applicant]
US 20140280081A1 · Tropin et al. · 2014 [cited by applicant]
US 20150026176A1 · Bullock · 2015 [cited by applicant]
US 20150120723A1 · Deshmukh et al. · 2015 [cited by applicant]
US 20150269176A1 · Marantz et al. · 2015 [cited by applicant]
US 20150287096A1 · Jacobsson · 2015 [cited by applicant]
US 20160034600A1 · Joshi · 2016 [cited by applicant]
US 20160124926A1 · Fallah · 2016 [cited by applicant]
US 20160147872A1 · Agarwalla · 2016 [cited by examiner]
US 20160147893A1 · Mashiach et al. · 2016 [cited by applicant]
US 20160180438A1 · Boston et al. · 2016 [cited by applicant]
US 20160275178A1 · Liu · 2016 [cited by examiner]
US 20170017718A1 · Nakayama et al. · 2017 [cited by applicant]
US 20170017724A1 · MacGillivray · 2017 [cited by examiner]
US 20170075985A1 · Chakraborty et al. · 2017 [cited by applicant]
US 20170097967A1 · Savliwala et al. · 2017 [cited by applicant]
US 20170116260A1 · Chattopadhyay · 2017 [cited by applicant]
US 20170116332A1 · Andrade Silva et al. · 2017 [cited by applicant]
US 20170124064A1 · Lu et al. · 2017 [cited by applicant]
US 20170199928A1 · Zhao et al. · 2017 [cited by applicant]
US 20170228372A1 · Moreno et al. · 2017 [cited by applicant]
US 20170278514A1 · Mathias et al. · 2017 [cited by applicant]
US 20170344622A1 · Islam et al. · 2017 [cited by applicant]
US 20190005953A1 · Bundalo et al. · 2019 [cited by applicant]
US 20190102482A1 · Ni · 2019 [cited by applicant]
US 20190147109A1 · Offer et al. · 2019 [cited by applicant]
US 20190163781A1 · Ackermann et al. · 2019 [cited by applicant]
US 20200210891A1 · Safronov et al. · 2020 [cited by applicant]
US 20210019309A1 · Yadav et al. · 2021 [cited by applicant]
US 20210026906A1 · Reznik · 2021 [cited by applicant]
US 20210056114A1 · Price et al. · 2021 [cited by applicant]
US 20210173836A1 · Robert Jose et al. · 2021 [cited by applicant]
US 20210232613A1 · Raval Contractor et al. · 2021 [cited by applicant]
US 20210279264A1 · Robert Jose et al. · 2021 [cited by applicant]
US 20210280174A1 · Robert Jose et al. · 2021 [cited by applicant]
US 20210280175A1 · Robert Jose et al. · 2021 [cited by applicant]
US 20210280176A1 · Robert Jose et al. · 2021 [cited by applicant]
US 20220100741A1 · Robert Jose et al. · 2022 [cited by applicant]
US 20230169960A1 · Robert Jose et al. · 2023 [cited by applicant]
US 20230214382A1 · Robert Jose et al. · 2023 [cited by applicant]
International Search Report and Written Opinion, dated Feb. 18, 2021, issued in International Application No. PCT/US2020/065307, Feb. 18, 2021, 13 pages. [cited by applicant]