IP Library › Granted Patent US 12,744,039
Granted Patent B2
US 12,744,039 · App. 18/735,672 · Granted Sep 22, 2026

Speech recognition method and apparatus

Inventors: Xifeng Yao (Nanjing, CN); Kaiji Chen (Shenzhen, CN)
Assignee: Huawei Technologies Co., Ltd.
G10L15/22G06F40/30G10L15/1822G10L15/26G10L25/54G10L25/63G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,039
App. No.
18/735,672
Granted
Sep 22, 2026
Kind
B2
Abstract

A speech recognition method and apparatus are provided. The speech recognition method includes: obtaining a first speech text (S 510 ); obtaining modal information that matches the first speech text; and performing multimodal semantic understanding with reference to the first speech text and the modal information, to output an intention and a slot. According to the speech recognition method, an intention of a user can be accurately recognized, improving human-computer interaction efficiency and user experience.

Claims (58)

1 . A method, comprising:

obtaining a first speech text;

obtaining, based on the first speech text, first modal information that matches the first speech text, wherein a modality indicated by the first modal information is a first modality in a plurality of preset modalities; and

determining, based on the first speech text and the first modal information, a first intention and a first slot that are indicated by the first speech text when the first speech text matches the first modal information, wherein the first intention is an executable intention supported by a system handling the method, and wherein the first slot is a slot holding data needed as input for executing the first intention;

wherein the obtaining, based on the first speech text, the first modal information that matches the first speech text comprises:

obtaining a multimodal selection vector based on the first speech text; and

obtaining the first modal information based on the multimodal selection vector; and

wherein the obtaining the multimodal selection vector based on the first speech text comprises:

determining a first context category to which the first speech text belongs; and

obtaining the multimodal selection vector based on the first context category and a first mapping matrix, wherein the multimodal selection vector indicates a probability of relevance between the first context category and each of the plurality of preset modalities, wherein the first mapping matrix indicates a plurality of context categories and a plurality of multimodal selection vectors, wherein each multimodal selection vector indicates one or more modalities, and wherein categories of the plurality of context categories are in a one-to-one correspondence with multimodal selection vectors of the plurality of multimodal selection vectors.

2 . The method according to claim 1 , further comprising:

obtaining modal information of the plurality of preset modalities.

3 . The method according to claim 2 , further comprising:

obtaining the first modal information based on the multimodal selection vector and the modal information of the plurality of preset modalities.

4 . The method according to claim 3 , wherein determining the first context category to which the first speech text belongs comprises:

determining, based on the first speech text or context information of the first speech text, the first context category to which the first speech text belongs.

5 . The method according to claim 4 , wherein determining, based on the first speech text or context information of the first speech text, the first context category to which the first speech text belongs comprises:

obtaining a text feature code of the first speech text or the context information of the first speech text; and

determining, based on the text feature code and a first classification layer, the first context category to which the first speech text belongs, wherein the first classification layer is used to map the first speech text to one of a plurality of preset context categories.

6 . The method according to claim 5 , wherein the first modal information comprises a first modal feature code; and

wherein determining, based on the first speech text and the first modal information, the first intention and the first slot that are indicated by the first speech text when the first speech text matches the first modal information comprises:

determining, based on the text feature code, the first modal feature code, and a second classification layer, the first intention and the first slot that are indicated by the first speech text in the first modality, wherein the second classification layer is used to map the first speech text to one of a plurality of preset intentions.

7 . An apparatus, comprising:

at least one processor; and

at least one computer-readable storage medium storing a program that is executable by the at least one processor, the program comprising instructions for:

obtaining a first speech text;

obtaining, based on the first speech text, first modal information that matches the first speech text, wherein a modality indicated by the first modal information is a first modality in a plurality of preset modalities;

determining, based on the first speech text and the first modal information, a first intention and a first slot that are indicated by the first speech text when the first speech text matches the first modal information, wherein the first intention is an executable intention supported by the apparatus, and wherein the first slot is a slot holding data needed as input for executing the first intention;

wherein the instructions for the obtaining, based on the first speech text, the first modal information that matches the first speech text comprise instructions for:

obtaining a multimodal selection vector based on the first speech text; and

obtaining the first modal information based on the multimodal selection vector; and

wherein the instructions for the obtaining the multimodal selection vector based on the first speech text comprises:

determining a first context category to which the first speech text belongs; and

obtaining the multimodal selection vector based on the first context category and a first mapping matrix, wherein the multimodal selection vector indicates a probability of relevance between the first context category and each of the plurality of preset modalities, wherein the first mapping matrix indicates a plurality of context categories and a plurality of multimodal selection vectors, wherein each multimodal selection vector indicates one or more modalities, and wherein categories of the plurality of context categories are in a one-to-one correspondence with multimodal selection vectors of the plurality of multimodal selection vectors.

8 . The apparatus according to claim 7 , wherein the program further comprises instructions for:

obtaining modal information of the plurality of preset modalities.

9 . The apparatus according to claim 8 , wherein the program comprises instructions for:

obtaining the first modal information based on the multimodal selection vector and the modal information of the plurality of preset modalities.

10 . The apparatus according to claim 9 , wherein the program comprises instructions for:

determining, based on the first speech text or context information of the first speech text, the first context category to which the first speech text belongs.

11 . The apparatus according to claim 10 , wherein the program comprises instructions for:

obtaining a text feature code of the first speech text or the context information of the first speech text; and

determining, based on the text feature code and a first classification layer, the first context category to which the first speech text belongs, wherein the first classification layer is used to map the first speech text to one of a plurality of preset context categories.

12 . The apparatus according to claim 11 , wherein the first modal information comprises a first modal feature code; and

wherein the program comprises instructions for:

determining, based on the text feature code, the first modal feature code, and a second classification layer, the first intention and the first slot corresponding to the first intention that are indicated by the first speech text in the first modality, wherein the second classification layer is used to map the first speech text to one of a plurality of preset intentions.

13 . The apparatus according to claim 7 , wherein the program further comprises instructions for:

performing an operation related to the first intention.

14 . A non-transitory computer readable storage medium storing instructions that are executable by at least one processor, the instructions comprising instructions for:

obtaining a first speech text;

obtaining, based on the first speech text, first modal information that matches the first speech text, wherein a modality indicated by the first modal information is a first modality in a plurality of preset modalities; and

determining, based on the first speech text and the first modal information, a first intention and a first slot that are indicated by the first speech text when the first speech text matches the first modal information, wherein the first intention is an executable intention supported by a system performing the instructions, and wherein the first slot is a slot holding data needed as input for executing the first intention;

wherein the obtaining, based on the first speech text, the first modal information that matches the first speech text comprises:

obtaining a multimodal selection vector based on the first speech text; and

obtaining the first modal information based on the multimodal selection vector; and

wherein the obtaining the multimodal selection vector based on the first speech text comprises:

determining a first context category to which the first speech text belongs; and

obtaining the multimodal selection vector based on the first context category and a first mapping matrix, wherein the multimodal selection vector indicates a probability of relevance between the first context category and each of the plurality of preset modalities, wherein the first mapping matrix indicates a plurality of context categories and a plurality of multimodal selection vectors, wherein each multimodal selection vector indicates one or more modalities, and wherein categories of the plurality of context categories are in a one-to-one correspondence with multimodal selection vectors of the plurality of multimodal selection vectors.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2024
From: YAO, XIFENG; CHEN, KAIJI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 068500/0619 →
Priority Claims (1)
CN 202111659241.3 · Dec 30, 2021 · national
Continuity (2)
Continuation PCTCN2022137124 · Dec 7, 2022
Related Publication 20240331694A1 · Oct 3, 2024
References Cited (16)
US 10515625B1 · Metallinou · 2019 [cited by examiner]
US 12217749B1 · Anand · 2025 [cited by examiner]
US 20240331694A1 · Yao · 2024 [cited by examiner]
CN 104955592A · 2015 [cited by examiner]
CN 104965592A · 2015 [cited by applicant]
CN 111008532A · 2020 [cited by examiner]
CN 111063162A · 2020 [cited by examiner]
CN 111508482A · 2020 [cited by examiner]
CN 111966320A · 2020 [cited by applicant]
CN 112462940A · 2021 [cited by applicant]
CN 112613534A · 2021 [cited by examiner]
CN 112784798A · 2021 [cited by examiner]
CN 109902155B · 2021 [cited by examiner]
CN 113806470A · 2021 [cited by examiner]
Konrad Gadzicki et al: “Early vs Late Fusion in Multimodal Convolutional Neural Networks,” Jul. 6-9, 2020, total 6 pages. [cited by applicant]
Sebastian Zambanini et al: “Early versus Late Fusion in a Multiple Camera Network for Fall Detection,” Jan. 2010, total 9 pages. [cited by applicant]