IP Library Granted Patent US 12,315,501
Granted Patent B2
US 12,315,501 · App. 18/423,556 · Granted May 27, 2025

Systems and methods for phonetic-based natural language understanding

Inventors: Ajay Kumar Mishra (Karnataka, IN); Jeffry Copps Robert Jose (Tamil Nadu, IN)
Assignee: ADEIA GUIDES INC.
G10L15/187G06F16/632G06F16/683G06F16/686G06N5/02G10L15/1822G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,501
App. No.
18/423,556
Granted
May 27, 2025
Kind
B2
Abstract

Systems and methods are described for modifying a phonetic search index based on a use frequency associated with phonetic representations of text terms included in metadata of a media item. A first phonetic representation of a text term of the metadata, pronounced as a word, may be generated. A second phonetic representation of the text term may be generated by concatenating a phonetic representation of each letter in the text term. A database may be queried to determine use frequencies of the first and second phonetic representations, one of which may be selected based on a comparison of the use frequencies. A phonetic search index may be modified by including an entry for the selected phonetic representation. A voice query related to the media item may be received, and a reply to the voice query may be generated for output by performing a lookup in the modified phonetic search index.

Claims (61)

1. A method comprising:

maintaining, by a media delivery service, a database of media items available for delivery, via a network, to a plurality of devices subscribed to the media delivery service;

determining that a media item will become available for delivery at a first time;

accessing metadata of the media item, the metadata comprising a text term;

generating a first phonetic representation of the text term pronounced as a word;

generating a second phonetic representation of the text term by concatenating a phonetic representation of each letter in the text term;

tracking a plurality of voice queries, received by the media delivery service, from the plurality of devices subscribed to the media delivery service, wherein the tracking is performed over a predefined period of time prior to the first time when the media item will become available for delivery to determine:

a first number of a first subset of the plurality of voice queries that matched the first phonetic representation over the predefined period of time; and

a second number of a second subset of the plurality of voice queries that matched the second phonetic representation over the predefined period of time;

after the predefined period of time, based at least in part on comparing the first number to the second number, selecting one of the first phonetic representation or the second phonetic representation;

modifying the database of media items to associate the media item with the selected phonetic representation; and

outputting an identifier of the media item based on a subsequent voice query of a device of the plurality of devices matching the selected phonetic representation.

2. The method of claim 1 , wherein the accessing the metadata of the media item is performed based in part on determining that a profile of a user indicates that the user is interested in the media item, and wherein the profile of the user comprises a plurality of user preferences and a viewing history of the user.

3. The method of claim 1 , wherein the metadata of the media item comprises at least one of: a title of the media item, a description of the media item, a scheduled broadcast time of the media item, an indication of access to the media item and an indication of what time the media item will be available to be streamed.

4. The method of claim 1 , wherein the plurality of voice queries received from the plurality of devices subscribed to the media delivery service comprises a plurality of phonetic representations of the plurality of voice queries, and wherein the plurality of phonetic representations of the plurality of voice queries comprises a plurality of phonemes.

5. The method of claim 1 , wherein the first phonetic representation and the second phonetic representation comprise a plurality of phonemes.

6. The method of claim 1 , further comprising:

training a machine learning model using labeled audio files or utterances; and

training the machine learning model to output phoneme or grapheme representations of a voice input, wherein the phoneme or the grapheme representations of the voice input are compared with the phonetic representation of the text term.

7. The method of claim 6 , wherein the machine learning model is trained to accept as input the phoneme or the grapheme representation of the voice input and output a likely intent or topic of the voice input.

8. The method of claim 1 , further comprising:

generating a hash value for the text term of a voice query of the plurality of voice queries; and

determining a variant of the text term of the voice query of the plurality of voice queries by comparing the hash value to a hash value of the variant of the text term of the voice query of the plurality of voice queries.

9. The method of claim 1 , wherein the modifying the database of media items comprises modifying a phonetic search index to include in the phonetic search index an entry for the selected phonetic representation.

10. The method of claim 1 , further comprising:

determining that a scheduled playtime or an availability time of the media item is not within the predefined period of time from a current time; and

based in part on the determining that the scheduled playtime or the availability time of the media item is not within the predefined period of time:

identifying alternative media items available to be played within the predefined period of time; and

accessing metadata associated with the alternative media items.

11. A system comprising:

a memory configured to maintain a database of media items available for delivery from a media delivery service, via a network, to a plurality of devices subscribed to the media delivery service;

an input/output (I/O) circuitry; and

a control circuitry configured to:

determine that a media item will become available for delivery at a first time;

access metadata of the media item, the metadata comprising a text term;

generate a first phonetic representation of the text term pronounced as a word;

generate a second phonetic representation of the text term by concatenating a phonetic representation of each letter in the text term;

track a plurality of voice queries, received by the media delivery service, from the plurality of devices subscribed to the media delivery service, wherein the tracking is performed over a predefined period of time prior to the first time when the media item will become available for delivery to determine:

a first number of a first subset of the plurality of voice queries that matched the first phonetic representation over the predefined period of time; and

a second number of a second subset of the plurality of voice queries that matched the second phonetic representation over the predefined period of time;

after the predefined period of time, based at least in part on comparing the first number to the second number, select one of the first phonetic representation or the second phonetic representation, wherein the selected phonetic representation is stored in the memory;

modify the database of media items stored in the memory to associate the media item with the selected phonetic representation; and

wherein the I/O circuitry is configured to:

output an identifier of the media item based on a subsequent voice query of a device of the plurality of devices matching the selected phonetic representation.

12. The system of claim 11 , wherein the control circuitry is configured to access the metadata of the media item based in part on determining that a profile of a user indicates that the user is interested in the media item, and wherein the profile of the user comprises a plurality of user preferences and a viewing history of the user.

13. The system of claim 11 , wherein the metadata of the media item comprises at least one of: a title of the media item, a description of the media item, a scheduled broadcast time of the media item, an indication of access to the media item and an indication of what time the media item will be available to be streamed.

14. The system of claim 11 , wherein the plurality of voice queries received from the plurality of devices subscribed to the media delivery service comprises a plurality of phonetic representations of the plurality of voice queries, and wherein the plurality of phonetic representations of the plurality of voice queries comprises a plurality of phonemes.

15. The system of claim 11 , wherein the first phonetic representation and the second phonetic representation comprise a plurality of phonemes.

16. The system of claim 11 , wherein the control circuitry is further configured to:

train a machine learning model using labeled audio files or utterances; and

train the machine learning model to output phoneme or grapheme representations of a voice input, wherein the phoneme or the grapheme representations of the voice input are compared with the phonetic representation of the text term.

17. The system of claim 16 , wherein the machine learning model is trained to accept as input the phoneme or the grapheme representation of the voice input and output a likely intent or topic of the voice input.

18. The system of claim 11 , wherein the control circuitry is further configured to:

generate a hash value for the text term of a voice query of the plurality of voice queries; and

determine a variant of the text term of the voice query of the plurality of voice queries by comparing the hash value to a hash value of the variant of the text term of the voice query of the plurality of voice queries.

19. The system of claim 11 , wherein the control circuitry is configured to modify the database of media items by modifying a phonetic search index to include in the phonetic search index an entry for the selected phonetic representation.

20. The system of claim 11 , wherein the control circuitry is further configured to:

determine that a scheduled playtime or an availability time of the media item is not within the predefined period of time from a current time; and

based in part on the determining that the scheduled playtime or the availability time of the media item is not within the predefined period of time:

identify alternative media items available to be played within the predefined period of time; and

access metadata associated with the alternative media items.

Assignments (2)
CHANGE OF NAME Recorded Oct 4, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069113/0348 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2024
From: MISHRA, AJAY KUMAR; ROBERT JOSE, JEFFRY COPPS
To: ROVI GUIDES, INC.
Reel/Frame 066260/0326 →
Continuity (2)
Continuation 17363651 · Jun 30, 2021
Related Publication 20240249718A1 · Jul 25, 2024
References Cited (12)
US 8731929B2 · Kennewick · 2014 [cited by examiner]
US 10977452B2 · Wang et al. · 2021 [cited by applicant]
US 11922931B2 · Mishra et al. · 2024 [cited by applicant]
US 20050015254A1 · Beaman · 2005 [cited by examiner]
US 20120078629A1 · Ikeda et al. · 2012 [cited by applicant]
US 20160335266A1 · Ogle · 2016 [cited by examiner]
US 20180260416A1 · Elkaim · 2018 [cited by examiner]
US 20200357390A1 · Bromand · 2020 [cited by examiner]
US 20220284882A1 · Peddinti et al. · 2022 [cited by applicant]
US 20230017352A1 · Mishra et al. · 2023 [cited by applicant]
Price, “End-To-End Spoken Language Understanding Without Matched Language Speech Model Pretraining Data” ICASSP 2020 IEEE International Conf. on Acoustics, Speech and Signal Processing (ICASSP), pp. 7979-7983, doi: 10.1… [cited by applicant]
Wang, et al., “Large-Scale Unsupervised Pre-Training for End-to-End Spoken Language Understanding,” ICASSP 2020—2020 IEEE International Conf. on Acoustics, Speech and Signal Processing (ICASSP) pp. 7999-8003 (2020) doi:… [cited by applicant]