IP Library Granted Patent US 12,093,312
Granted Patent B2
US 12,093,312 · App. 17/982,665 · Granted Sep 17, 2024

Systems and methods for providing search query responses having contextually relevant voice output

Inventors: Ankur Anil Aher (Maharashtra, IN); Harish Ashok Kumar (Bangalore, IN)
Assignee: Rovi Guides, Inc.
G06F16/637G06F16/90332G06F16/907G10L13/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,093,312
App. No.
17/982,665
Granted
Sep 17, 2024
Kind
B2
Abstract

Systems and methods are described for responding to a search query with a contextually relevant voice output. An illustrative method receives a search query, determines an answer to the search query, identifies a media content reference included in the search query, determines, based on the media content reference, a personality associated with the media content reference, identifies a voice profile of the personality, and generates audio output using the voice profile of the personality, the audio output including the answer to the search query.

Claims (72)

1. A computer-implemented method comprising:

receiving input;

determining that the input includes a reference to a video content item;

in response to determining that the input includes the reference to the video content item:

querying a voice profile database to identify a voice performer included in a cast of the video content item referenced in the input, wherein the voice profile database comprises a plurality of voice profiles of a plurality of voice performers, and the voice profile database comprises, for each respective voice performer of the plurality of voice performers, an indication of a video content item that the respective voice performer is included in a cast of;

determining that the voice profile database comprises a voice profile of the identified voice performer; and

in response to (a) identifying the voice performer included in the cast of the video content item referenced in the input and (b) determining that the voice profile database comprises the voice profile of the identified voice performer:

generating output in a voice of the identified voice performer using the voice profile of the identified voice performer stored at the database.

2. The method of claim 1 , wherein generating the output comprises:

determining a textual form of the output, the textual form including a plurality of words;

synthesizing, using the voice profile of the identified voice performer, audio matching at least a portion of the plurality of words; and

generating the output as an audio output based on the synthesized audio.

3. The method of claim 2 , wherein synthesizing, using the voice profile of the identified voice performer, audio matching at least a portion of the plurality of words comprises:

determining, based on the voice profile of the identified voice performer, a characteristic of the voice of the identified voice performer;

identifying a plurality of audio templates, wherein each audio template of the plurality of audio templates corresponds to a respective word of the plurality of words;

modifying each audio template of the plurality of audio templates based on the characteristic of the voice of the identified voice performer; and

generating audio corresponding to each modified audio template.

4. The method of claim 2 , further comprising:

identifying an audio recording of the voice of the identified voice performer, wherein the audio recording includes at least one of the plurality of words, and

wherein the generating of the audio output is performed based on the synthesized audio and the audio recording.

5. The method of claim 4 , further comprising:

identifying one or more words of the plurality of words that are not included in the audio recording,

wherein synthesizing the audio matching at least a portion of the plurality of words comprises synthesizing audio corresponding to the one or more words of the plurality of words that are not included in the audio recording.

6. The method of claim 1 , wherein generating the output in the voice of the identified voice performer using the voice profile of the identified voice performer comprises:

retrieving, from the voice profile database, an audio recording of the identified voice performer; and

generating the output based at least in part on the audio recording.

7. The method of claim 1 , wherein determining that the voice profile database comprises the voice profile of the identified voice performer comprises:

retrieving, from the voice profile database, an indication of a character associated with the video content item identified by the reference to the video content item; and

determining the voice performer provides a voice of the character.

8. The method of claim 1 , wherein the identified voice performer corresponds to a character in the video content item, and an indication of the character is included in a title of the video content item.

9. The method of claim 1 , further comprising:

identifying a particular keyword, phrase, tune, or jingle associated with the video content item,

wherein the generated output includes the particular keyword, phrase, tune, or jingle associated with the video content item.

10. The method of claim 1 , wherein the input is a search query, and the output is an answer to the search query.

11. A system comprising:

a voice profile database:

control circuitry configured to:

receive input;

determine that the input includes a reference to a video content item;

in response to determining that the input includes the reference to the video content item:

query the voice profile database to identify a voice performer included in a cast of the video content item referenced in the input, wherein the voice profile database comprises a plurality of voice profiles of a plurality of voice performers, and the voice profile database comprises, for each respective voice performer of the plurality of voice performers, an indication of a video content item that the respective voice performer is included in a cast of;

determine that the voice profile database comprises a voice profile of the identified voice performer; and

in response to (a) identifying the voice performer included in the cast of the video content item referenced in the input and (b) determining that the voice profile database comprises the voice profile of the identified voice performer:

generate output in a voice of the identified voice performer using the voice profile of the identified voice performer stored at the database.

12. The system of claim 11 , wherein the control circuitry is configured to generate the output by:

determining a textual form of the output, the textual form including a plurality of words;

synthesizing, using the voice profile of the identified voice performer, audio matching at least a portion of the plurality of words; and

generating the output as an audio output based on the synthesized audio.

13. The system of claim 12 , wherein the control circuitry is configured to synthesize, using the voice profile of the identified voice performer, audio matching at least a portion of the plurality of words by:

determining, based on the voice profile of the identified voice performer, a characteristic of the voice of the identified voice performer;

identifying a plurality of audio templates, wherein each audio template of the plurality of audio templates corresponds to a respective word of the plurality of words;

modifying each audio template of the plurality of audio templates based on the characteristic of the voice of the identified voice performer; and

generating audio corresponding to each modified audio template.

14. The system of claim 12 , wherein the control circuitry is further configured to:

identify an audio recording of the voice of the identified voice performer, wherein the audio recording includes at least one of the plurality of words; and

generate the audio output based on the synthesized audio and the audio recording.

15. The system of claim 14 , wherein the control circuitry is further configured to:

identify one or more words of the plurality of words that are not included in the audio recording; and

synthesize the audio matching at least a portion of the plurality of words by synthesizing audio corresponding to the one or more words of the plurality of words that are not included in the audio recording.

16. The system of claim 11 , wherein the control circuitry is configured to generate the output in the voice of the identified voice performer using the voice profile of the identified voice performer by:

retrieving, from the voice profile database, an audio recording of the identified voice performer; and

generating the output based at least in part on the audio recording.

17. The system of claim 11 , wherein the control circuitry is configured to determine that the voice profile database comprises the voice profile of the identified voice performer by:

retrieving, from the voice profile database, an indication of a character associated with the video content item identified by the reference to the video content item; and

determining the voice performer provides a voice of the character.

18. The system of claim 11 , wherein the identified voice performer corresponds to a character in the video content item, and an indication of the character is included in a title of the video content item.

19. The system of claim 11 , wherein the control circuitry is further configured to:

identify a particular keyword, phrase, tune, or jingle associated with the video content item; and

generate the output to include the identified particular keyword, phrase, tune, or jingle associated with the video content item.

20. The method of claim 1 , wherein:

the method further comprises providing a machine learning model; and

the generating of the output in the voice of the identified voice performer using the voice profile of the identified voice performer stored at the database is performed using the machine learning model.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0171 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2022
From: AHER, ANKUR ANIL; KUMAR, HARISH ASHOK
To: ROVI GUIDES, INC.
Reel/Frame 061690/0131 →
Continuity (2)
Continuation 16201352 · Nov 27, 2018
Related Publication 20230169112A1 · Jun 1, 2023