IP Library Granted Patent US 11,562,737
Granted Patent B2
US 11,562,737 · App. 16/728,476 · Granted Jan 24, 2023

Generating topic-specific language models

Inventors: David F. Houghton (Brattleboro, VT); Seth Michael Murray (Redwood City, CA); Sibley Verbeck Simon (Santa Cruz, CA)
Assignee: TIVO CORPORATION
G10L15/183G10L15/197
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,737
App. No.
16/728,476
Granted
Jan 24, 2023
Kind
B2
Abstract

Speech recognition may be improved by generating and using a topic specific language model. A topic specific language model may be created by performing an initial pass on an audio signal using a generic or basis language model. A speech recognition device may then determine topics relating to the audio signal based on the words identified in the initial pass and retrieve a corpus of text relating to those topics. Using the retrieved corpus of text, the speech recognition device may create a topic specific language model. In one example, the speech recognition device may adapt or otherwise modify the generic language model based on the retrieved corpus of text.

Claims (59)

1. A method comprising:

determining, based on a first speech recognition process associated with a first language model, a topic associated with an audio signal;

performing a plurality of searches of a corpus to identify a plurality of terms related to the topic, wherein the corpus comprises a collection of text other than a transcript of the audio signal;

in response to determining that the quantity of the plurality of terms identified by the searches as related to the topic matches or exceeds a threshold quantity:

generating, based on the plurality of terms identified in the corpus, a second language model; and

determining, based on a second speech recognition process associated with the generated second language model, the transcript of the audio signal.

2. The method of claim 1 , further comprising at least one of:

generating, based on the transcript, a closed-captioning feed for the audio signal:

causing words of the transcript to be input into a computer application; or

determining, based on the transcript, topic segments of the audio signal.

3. The method of claim 1 , wherein the quantity of the plurality of terms is based on at least one of:

a total quantity of terms needed to generate the second language model and a quantity of topics associated with the audio signal; or

a respective significance, based on the first speech recognition process, for each of a plurality of topics associated with the audio signal.

4. The method of claim 1 , wherein the second language model comprises a modification of the first language model.

5. The method of claim 1 , wherein the determining the topic comprises:

determining that a frequency of one or more terms, in the audio signal and associated with the topic, satisfies a frequency threshold.

6. The method of claim 1 , wherein the determining the plurality of terms comprises continuing to perform searches to identify terms until corresponding search results matches or exceeds a threshold quantity.

7. The method of claim 1 , wherein the one or more searches comprise at least one of: a web search; or a publication database search.

8. The method of claim 1 , further comprising: determining the second language model based on the first language model.

9. The method of claim 1 , wherein, in response to the quantity of the plurality of terms related to topic not meeting a threshold quantity:

conducting additional searches associated with the topic to determine additional plurality of terms related to the topic.

10. The method of claim 1 , wherein, performing a plurality of searches of a corpus to identify a plurality of terms related to the topic further comprises, creating the plurality of searches by assembling known keywords associated with the topic.

11. An apparatus comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the apparatus to:

determine, based on a first speech recognition process associated with a first language model, a topic associated with an audio signal;

perform a plurality of searches of a corpus to identify a plurality of terms related to the topic, wherein the corpus comprises a collection of text other than a transcript of the audio signal;

in response to determining that the quantity of the plurality of terms identified by the searches as related to the topic matches or exceeds a threshold quantity:

generate, based on the plurality of terms identified in corpus, a second language model; and

determine, based on a second speech recognition process associated with the generated second language model, the transcript of the audio signal.

12. The apparatus of claim 11 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to perform at least one of:

generate, based on the transcript, a closed-captioning feed for the audio signal; cause words of the transcript to be input into a computer application; or

determine, based on the transcript, topic segments of the audio signal.

13. The apparatus of claim 11 , wherein the quantity of the plurality of terms is based on at least one of:

a total quantity of terms needed to generate the second language model and a quantity of topics associated with the audio signal; or

a respective significance, based on the first speech recognition process, for each of a plurality of topics associated with the audio signal.

14. The apparatus of claim 11 , wherein the second language model comprises a modification of the first language model.

15. The apparatus of claim 11 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to determine the topic by determining that a frequency of one or more terms, in the audio signal and associated with the topic, satisfies a frequency threshold.

16. The apparatus of claim 11 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to determine the plurality of terms by continuing to perform searches until corresponding search results matches or exceeds a threshold quantity.

17. The apparatus of claim 11 , wherein the one or more searches comprise at least one of: a web search; or a publication database search.

18. The apparatus of claim 11 , wherein the instructions, when executed by the one or more processors, further cause the apparatus to: determine the second language model based on the first language model.

19. A non-transitory computer-readable medium storing instructions that, when executed, cause:

determining, based on a first speech recognition process associated with a first language model, a topic associated with an audio signal;

performing a plurality of searches of a corpus to identify a plurality of terms related to the topic, wherein the corpus comprises a collection of text other than a transcript of the audio signal;

in response to determining that the quantity of the plurality of terms identified by the searches as related to the topic matches or exceeds a threshold quantity:

generating, based on the plurality of terms identified in the corpus, a second language model; and

determining, based on a second speech recognition process associated with the generated second language model, the transcript of the audio signal.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the instructions, when executed, further cause at least one of:

generating, based on the transcript, a closed-captioning feed for the audio signal;

causing words of the transcript to be input into a computer application; or

determining, based on the transcript, topic segments of the audio signal.

21. The non-transitory computer-readable storage medium of claim 19 , wherein the quantity of the plurality of terms is based on at least one of:

a total quantity of terms needed to generate the second language model and a quantity of topics associated with the audio signal; or

a respective significance, based on the first speech recognition process, for each of a plurality of topics associated with the audio signal.

22. The non-transitory computer-readable storage medium of claim 19 , wherein the second language model comprises a modification of the first language model.

23. The non-transitory computer-readable storage medium of claim 19 , wherein the determining the topic comprises: determining that a frequency of one or more terms, in the audio signal and associated with the topic, satisfies a frequency threshold.

24. The non-transitory computer-readable storage medium of claim 19 , wherein the determining the plurality of terms comprises continuing to perform searches until corresponding search results matches or exceeds a threshold quantity.

25. The non-transitory computer-readable storage medium of claim 19 , wherein the one or more searches comprise at least one of: a web search; or a publication database search.

26. The non-transitory computer-readable storage medium of claim 19 , wherein the instructions, when executed, further cause: determining the second language model based on the first language model.

Assignments (6)
CHANGE OF NAME Recorded Mar 31, 2026
From: ADEIA MEDIA HOLDINGS LLC
To: ADEIA MEDIA HOLDINGS INC.
Reel/Frame 075303/0717 →
CHANGE OF NAME Recorded Oct 1, 2024
From: TIVO CORPORATION
To: TIVO LLC
Reel/Frame 069083/0230 →
CHANGE OF NAME Recorded Oct 1, 2024
From: TIVO LLC
To: ADEIA MEDIA HOLDINGS LLC
Reel/Frame 069083/0311 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2020
From: HOUGHTON, DAVID F.; MURRAY, SETH MICHAEL; SIMON, SIBLEY VERBECK
To: COMCAST INTERACTIVE MEDIA, LLC
Reel/Frame 054569/0067 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2020
From: COMCAST INTERACTIVE MEDIA, LLC
To: TIVO CORPORATION
Reel/Frame 054540/0104 →