IP Library › Granted Patent US 12,451,128
Granted Patent B2
US 12,451,128 · App. 17/780,776 · Granted Oct 21, 2025

Voice recognition method and related product

Inventors: Genshun Wan (Anhui, CN); Jianqing Gao (Anhui, CN); Zhiguo Wang (Anhui, CN)
Assignee: IFLYTEK CO., LTD.
G10L15/183G06F40/247G06F40/289G10L15/04G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,128
App. No.
17/780,776
Granted
Oct 21, 2025
Kind
B2
Abstract

A speech recognition method and related products are provided. The method includes acquiring text contents and text-associated time information transmitted by a plurality of terminals in a preset scenario and determining a shared text for the preset scenario based on the text contents and the text-associated time information, obtaining a customized language model for the preset scenario based on the shared text, and performing speech recognition for the preset scenario with the customized language model. The method provides improved speech recognition for the preset scenario due to the correlation between the customized language model and the preset scenario.

Claims (58)

1. A speech recognition method comprising:

acquiring text contents and text-associated time information transmitted by a plurality of terminals in a preset scenario and determining a shared text for the preset scenario based on the text contents and the text-associated time information; and

obtaining a customized language model for the preset scenario based on the shared text, and performing speech recognition for the preset scenario with the customized language model, comprising:

performing word segmentation and classification on the shared text to obtain one or more keywords, and updating a hot word list based on the one or more keywords to obtain a new hot word list; and

performing speech recognition with the customized language model and the new hot word list.

2. The method according to claim 1 , wherein before performing speech recognition for the preset scenario with the customized language model, the method further comprises:

performing word segmentation and classification on the shared text to obtain one or more keywords, and updating a hot word list for the preset scenario based on the one or more keywords to obtain a new hot word list.

3. The method according to claim 2 , wherein performing speech recognition for the preset scenario with the customized language model comprises:

performing speech recognition for the preset scenario with the customized language model and the new hot word list.

4. The method according to claim 1 , wherein the new hot word list takes effect when the new hot word list is generated.

5. The method according to claim 1 , wherein determining the shared text for the preset scenario based on the text contents and the text-associated time information comprises:

acquiring and recognizing a speech in the preset scenario to obtain a speech recognition result, wherein the speech recognition result comprises a sentence text and sentence-associated time information; and

comparing the text-associated time information with the sentence-associated time information, and determine a text content corresponding to text-associated time information matching the sentence-associated time information as the shared text.

6. The method according to claim 5 , further comprising:

determining the speech recognition result as the shared text.

7. The method according to claim 1 , wherein performing word segmentation and classification on the shared text to obtain the one or more keywords, and updating the hot word list for the preset scenario based on the one or more keywords to obtain the new hot word list comprises:

performing word segmentation and classification on the shared text to obtain a phrase set or a sentence set; and

determining the one or more keywords based on word frequencies of phrases and a word frequency threshold, wherein the word frequency of a phrase represents the number of occurrences of the phrase in the phrase set or the sentence set.

8. The method according to claim 7 , wherein determining the one or more keywords based on the word frequencies of the phrases and the word frequency threshold comprises:

acquiring the word frequency of each phrase in the phrase set;

determining a phrase having a word frequency greater than or equal to the word frequency threshold and transmitted by different terminals as one of the one or more keywords; and

selecting another one of the one or more keywords from phrases having word frequencies less than the word frequency threshold with a TF-IDF algorithm.

9. The method according to claim 7 , wherein before determining the one or more keywords based on the word frequencies of the phrases and the word frequency threshold, the method further comprises:

filtering the phrase set based on the hot word list.

10. The method according to claim 1 , wherein performing word segmentation and classification on the shared text to obtain the one or more keywords, and updating the hot word list based on the one or more keywords to obtain the new hot word list further comprises:

determining homonym phrases in the one or more keywords or between the one or more keywords and the words in the hot word list;

determining a sentence text including a phrase among the homonym phrases and replacing the phrase in the sentence text with a homonym of the phrase among the homonym phrases to obtain a sentence text subjected to phrase replacement; and

determining, based on language model scores of sentence texts subjected to phrase replacement, a homonym phrase in a sentence text having a highest language model score as a new word to be added to the hot word list.

11. The method according to claim 5 , wherein before obtaining the customized language model for the preset scenario based on the shared text and performing speech recognition for the preset scenario with the customized language model, the method further comprises:

segmenting the speech recognition result to obtain a segmentation time point of a paragraph; and

obtaining, after the segmentation time point, the customized language model for the preset scenario based on the shared text, and performing speech recognition for the preset scenario with the customized language model.

12. The method according to claim 11 , wherein obtaining, after the segmentation time point, the customized language model for the preset scenario based on the shared text comprises:

determining text similarities between the text contents and the speech recognition result; and

filtering out, based on the text similarities and a text similarity threshold, a text content corresponding to text similarity less than the similarity threshold.

13. The method according to claim 11 , wherein

the plurality of terminals comprise a first terminal and a second terminal; and

after the segmentation time point, the method further comprises:

acquiring text similarities between text contents of the first terminal and the second terminal as first text similarities;

determining the number of first text similarities corresponding to the first terminal that are greater than a first preset similarity threshold;

acquiring a text similarity between the text content transmitted by the first terminal and the speech recognition result transmitted by the first terminal as a second text similarity; and

filtering the shared text transmitted by the first terminal based on the number and the second text similarity.

14. The method according to claim 1 , wherein obtaining the customized language model for the preset scenario based on the shared text comprises:

acquiring an initial language model based on a shared text in a paragraph set, wherein the paragraph set is obtained after recognition of a current speech paragraph is finished; and

wherein probability interpolation is performed on the initial language model and a preset language model to obtain the customized language model.

15. The method according to claim 1 , wherein

the text content is generated by a user on the terminal and is related to the preset scenario; and

the text content comprises at least one of: a note made by the user based on the preset scenario, a mark made by the user on an electronic material related to the preset scenario, or a picture including text information taken by the user using an intelligent terminal.

16. A speech recognition apparatus comprising a processor configured to:

acquire text contents and text-associated time information transmitted by a plurality of terminals in a preset scenario and determine a shared text for the preset scenario based on the text contents and the text-associated time information; and

obtain a customized language model for the preset scenario based on the shared text, and perform speech recognition for the preset scenario with the customized language model, comprising:

performing word segmentation and classification on the shared text to obtain one or more keywords, and updating a hot word list based on the one or more keywords to obtain a new hot word list; and

performing speech recognition with the customized language model and the new hot word list.

17. A non-transient computer storage medium storing a computer program, wherein the computer program comprises program instructions that, when executed by a processor:

acquire text contents and text-associated time information transmitted by a plurality of terminals in a preset scenario and determine a shared text for the preset scenario based on the text contents and the text-associated time information; and

obtain a customized language model for the preset scenario based on the shared text, and perform speech recognition for the preset scenario with the customized language model, comprising:

performing word segmentation and classification on the shared text to obtain one or more keywords, and updating a hot word list based on the one or more keywords to obtain a new hot word list; and

performing speech recognition with the customized language model and the new hot word list.

18. The method according to claim 1 , wherein the keyword is a phrase having a word frequency greater than or equal to a word frequency threshold and transmitted by different terminals.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 27, 2022
From: WAN, GENSHUN; GAO, JIANQING; WANG, ZHIGUO
To: IFLYTEK CO., LTD.
Reel/Frame 060040/0655 →
Priority Claims (1)
CN 201911389673.X · Dec 28, 2019 · national
Continuity (1)
Related Publication 20230035947A1 · Feb 2, 2023
References Cited (35)
US 8447608B1 · Chang et al. · 2013 [cited by applicant]
US 20120143605A1 · Thorsen et al. · 2012 [cited by applicant]
US 20130346077A1 · Mengibar · 2013 [cited by examiner]
US 20150278191A1 · Levit et al. · 2015 [cited by applicant]
US 20170125013A1 · Yan · 2017 [cited by applicant]
US 20180173494A1 · Choi et al. · 2018 [cited by applicant]
US 20190251716A1 · Nelson · 2019 [cited by applicant]
US 20210065704A1 · Kim · 2021 [cited by applicant]
CN 103838756A · 2014 [cited by applicant]
CN 104464733A · 2015 [cited by applicant]
CN 105045778A · 2015 [cited by applicant]
CN 105448292A · 2016 [cited by applicant]
CN 105654945A · 2016 [cited by applicant]
CN 105719649A · 2016 [cited by applicant]
CN 106328147A · 2017 [cited by applicant]
CN 107644641A · 2018 [cited by applicant]
CN 108984529A · 2018 [cited by applicant]
CN 109272995A · 2019 [cited by applicant]
CN 110415705A · 2019 [cited by examiner]
CN 110534094A · 2019 [cited by applicant]
CN 110544477A · 2019 [cited by applicant]
CN 111161739A · 2020 [cited by applicant]
CN 112037792A · 2020 [cited by applicant]
CN 112562659A · 2021 [cited by applicant]
JP 2004233541A · 2004 [cited by applicant]
JP 201148405A · 2011 [cited by applicant]
JP 2013029652A · 2013 [cited by applicant]
KR 20180069660A · 2018 [cited by applicant]
KR 20190121721A · 2019 [cited by applicant]
Japanese Office Action issued in 2022-531437 mailed May 30, 2023, 7 pages. [cited by applicant]
Liu et al., “Scene Text Recognition with CNN Classifier and WFST-Based Word Labeling,” ICPR, 2016, pp. 1-6. [cited by applicant]
Liu et al., “Hierarchically Browse and Annotation System for News Video,” China Academic Journal Electronic Publishing House, vol. 35, No. 1, 2009, pp. 1-3. [cited by applicant]
Chinese First Office Action issued in 201911389673.X mailed Mar. 22, 2022, 11 pages. [cited by applicant]
International Search Report and Written Opinion issued in PCT/CN2020/136126 mailed Mar. 8, 2021, 15 pages. [cited by applicant]
European Search report received for PCT Patent Application No. PCT/CN2020136126, mailed on Dec. 18, 2023. [cited by applicant]