IP Library Granted Patent US 12,562,157
Granted Patent B2
US 12,562,157 · App. 18/629,200 · Granted Feb 24, 2026

Generating topic-specific language models

Inventors: David F. Houghton (Brattleboro, VT); Seth Michael Murray (Redwood City, CA); Sibley Verbeck Simon (Santa Cruz, CA)
Assignee: Adeia Media Holdings LLC
G10L15/183G10L15/197
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,157
App. No.
18/629,200
Granted
Feb 24, 2026
Kind
B2
Abstract

Speech recognition may be improved by generating and using a topic specific language model. A topic specific language model may be created by performing an initial pass on an audio signal using a generic or basis language model. A speech recognition device may then determine topics relating to the audio signal based on the words identified in the initial pass and retrieve a corpus of text relating to those topics. Using the retrieved corpus of text, the speech recognition device may create a topic specific language model. In one example, the speech recognition device may adapt or otherwise modify the generic language model based on the retrieved corpus of text.

Claims (47)

1 . A method comprising:

determining, based on a first speech recognition process associated with a first language model, a topic associated with an audio signal;

performing a plurality of searches of a corpus to identify a plurality of terms related to the topic, wherein the corpus comprises a collection of text other than a transcript of the audio signal;

in response to determining that the quantity of the plurality of terms identified by the searches as related to the topic does not match or exceed a threshold quantity:

removing irrelevant text from the collection of text of the corpus;

continuing to perform searches to identify terms until corresponding search results match or exceed the threshold quantity;

generating, based on the plurality of terms identified in the corpus, a second language model; and

determining, based on a second speech recognition process associated with the generated second language model, the transcript of the audio signal.

2 . The method of claim 1 wherein removing irrelevant text from the collection of text of the corpus comprises cleaning the collection of text of the corpus.

3 . The method of claim 1 , wherein the quantity of the plurality of terms is based on at least one of:

a total quantity of terms needed to generate the second language model and a quantity of topics associated with the audio signal; or

a respective significance, based on the first speech recognition process, for each of a plurality of topics associated with the audio signal.

4 . The method of claim 1 , wherein the second language model comprises a modification of the first language model.

5 . The method of claim 1 , wherein the plurality of searches of a corpus comprises at least one of: a web search or a publication database search.

6 . The method of claim 1 wherein removing irrelevant text from the collection of text of the corpus comprises:

extracting the collection of text from a plurality of files of the corpus; and

generating raw text from the collection of text.

7 . The method of claim 6 wherein generating the raw text from the collection of text comprises:

extracting characters corresponding to the terms identified in the corpus; and

removing extraneous information from the collection of text, wherein the extraneous information comprises one or more of formatting or metadata.

8 . The method of claim 6 further comprising removing text which does not form a part of the content of the collection of text from the collection of text.

9 . The method of claim 8 wherein text which does not form a part of the content of the collection of text comprises one or more of a header, an HTML tag, or an HTML markup element.

10 . The method of claim 8 further comprising identifying the text which does not form a part of the content of the collection of text by cross-referencing a dictionary of extraneous text.

11 . An apparatus comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the apparatus to:

determine, based on a first speech recognition process associated with a first language model, a topic associated with an audio signal;

perform a plurality of searches of a corpus to identify a plurality of terms related to the topic, wherein the corpus comprises a collection of text other than a transcript of the audio signal;

in response to determining that the quantity of the plurality of terms identified by the searches as related to the topic does not match or exceed a threshold quantity:

remove irrelevant text from the collection of text of the corpus;

continue to perform searches to identify terms until corresponding search results match or exceed the threshold quantity;

generate, based on the plurality of terms identified in the corpus, a second language model; and

determine, based on a second speech recognition process associated with the generated second language model, the transcript of the audio signal.

12 . The apparatus of claim 11 wherein the instructions, when executed by the one or more processors, further cause the apparatus to remove irrelevant text from the collection of text of the corpus by cleaning the collection of text of the corpus.

13 . The apparatus of claim 11 , wherein the quantity of the plurality of terms is based on at least one of:

a total quantity of terms needed to generate the second language model and a quantity of topics associated with the audio signal; or a respective significance, based on the first speech recognition process, for each of a plurality of topics associated with the audio signal.

14 . The apparatus of claim 11 , wherein the second language model comprises a modification of the first language model.

15 . The apparatus of claim 11 , wherein the plurality of searches of a corpus comprises at least one of: a web search or a publication database search.

16 . The apparatus of claim 11 wherein the instructions, when executed by the one or more processors, further cause the apparatus to remove irrelevant text from the collection of text of the corpus by:

extracting the collection of text from a plurality of files of the corpus; and

generating raw text from the collection of text.

17 . The apparatus of claim 16 wherein the instructions, when executed by the one or more processors, further cause the apparatus to generate the raw text from the collection of text by:

extracting characters corresponding to the terms identified in the corpus; and

removing extraneous information from the collection of text, wherein the extraneous information comprises one or more of formatting or metadata.

18 . The apparatus of claim 16 wherein the instructions, when executed by the one or more processors, further cause the apparatus to remove text which does not form a part of the content of the collection of text from the collection of text.

19 . The apparatus of claim 18 wherein text which does not form a part of the content of the collection of text comprises one or more of a header, an HTML tag, or an HTML markup element.

20 . The apparatus of claim 18 wherein the instructions, when executed by the one or more processors, further cause the apparatus to identify the text which does not form a part of the content of the collection of text by cross-referencing a dictionary of extraneous text.

Assignments (6)
CHANGE OF NAME Recorded Mar 31, 2026
From: ADEIA MEDIA HOLDINGS LLC
To: ADEIA MEDIA HOLDINGS INC.
Reel/Frame 075303/0717 →
SECURITY INTEREST Recorded May 28, 2025
From: ADEIA INC. (F/K/A XPERI HOLDING CORPORATION); ADEIA HOLDINGS INC.; ADEIA MEDIA HOLDINGS INC.; ADEIA IMAGING LLC; ADEIA MEDIA LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA TECHNOLOGIES INC.; ADEIA GUIDES INC.; ADEIA SOLUTIONS LLC; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR INTELLECTUAL PROPERTY LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA PUBLISHING INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 071454/0343 →
CHANGE OF NAME Recorded Oct 1, 2024
From: TIVO CORPORATION
To: TIVO LLC
Reel/Frame 069083/0230 →
CHANGE OF NAME Recorded Oct 1, 2024
From: TIVO LLC
To: ADEIA MEDIA HOLDINGS LLC
Reel/Frame 069083/0311 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: HOUGHTON, DAVID F.; MURRAY, SETH MICHAEL; SIMON, SIBLEY VERBECK
To: COMCAST INTERACTIVE MEDIA, LLC
Reel/Frame 067190/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: COMCAST INTERACTIVE MEDIA, LLC
To: TIVO CORPORATION
Reel/Frame 067190/0817 →