IP Library › Granted Patent US 12,536,424
Granted Patent B2
US 12,536,424 · App. 18/444,050 · Granted Jan 27, 2026

Intention recognition method and apparatus, readable medium, and electronic device

Inventors: Xiaoyang Li (Beijing, CN); Zilin Yu (Beijing, CN); Xiangyang Zhang (Beijing, CN); Xiaogang Tian (Beijing, CN); Zejun Ma (Beijing, CN)
Assignee: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,424
App. No.
18/444,050
Granted
Jan 27, 2026
Kind
B2
Abstract

The present application relates to an intention recognition method and apparatus, a readable medium, and an electronic device. The method includes: by means of a preset intention recognition quantification model, performing a quantification operation on a dot product of a query vector and a key vector which correspond to each character in a target text, so as to obtain a fixed-point type target vector of a first bit; according to the fixed-point type target vector, determining, by means of a target mapping relationship, a floating-point type attention weight of a second bit corresponding to each character; and according to the floating-point type attention weight, determining a target intention corresponding to the target text, the first bit being smaller than the second bit.

Claims (83)

1 . An intention recognition method, comprising:

acquiring a target text to be recognized; and

taking the target text as an input of a preset intention recognition quantization model to output a target intention of the target text;

wherein the preset intention recognition quantization model is configured to perform a quantization operation on a dot product of a query vector and a key vector corresponding to each character in the target text to generate a fixed-point target vector having a first bit number, determine a floating-point attention weight having a second bit number corresponding to each character in the target text through looking up a target mapping relationship, which is generated and stored in a computer in advance, according to the fixed-point target vector, and generate the target intention corresponding to the target text according to the floating-point attention weight, wherein the first bit number is smaller than the second bit number;

wherein the preset intention recognition quantization model comprises at least one multi-head attention layer, and the multi-head attention layer comprises a look-up table node; and

the look-up table node is configured to receive the fixed-point target vector, take the fixed-point target vector as an input to look up the target mapping relationship to determine a storage address corresponding to the fixed-point target vector, and access the storage address to determine the floating-point attention weight having the second bit number corresponding to each character;

wherein the multi-head attention layer further comprises a quantization node, and the look-up table node is coupled with an output end of the quantization node;

the quantization node is configured to perform the quantization operation on the dot product of the query vector and the key vector corresponding to each character in the target text to generate the fixed-point target vector having the first bit number, and input the fixed-point target vector into the look-up table node;

wherein the multi-head attention layer further comprises a Matmul node and a first dequantization node coupled with an output end of the Matmul node, and an output end of the first dequantization node is coupled with an input end of the quantization node;

the Matmul node is configured to acquire a query matrix and a key matrix corresponding to a target text sequence to be recognized, wherein the query matrix consists of query vectors corresponding to respective characters in the target text, and the key matrix consists of key vectors corresponding to the respective characters in the target text;

the Matmul node is further configured to obtain a target product of the query matrix and the key matrix, so as to obtain a specified matrix formed by the dot product of the query vector and the key vector corresponding to each character in the target text, wherein the specified matrix is fixed-point data having a third bit number, and the third bit number is greater than the first bit number;

the first dequantization node is configured to dequantize the specified matrix to obtain a floating-point product matrix comprising data having the second bit number; and

the quantization node is configured to quantize the floating-point product matrix to obtain a target matrix, and the target matrix includes fixed-point target vectors corresponding to the respective characters.

2 . The method according to claim 1 , wherein the look-up table node is configured to receive the fixed-point target vector, take the fixed-point target vector as an input to look up the target mapping relationship to determine a storage address corresponding to the fixed-point target vector, and access the storage address to determine the floating-point attention weight having the second bit number corresponding to each character, including:

the look-up table node obtaining a maximum numerical value of the fixed-point target vector corresponding to each character in the target matrix, and obtaining a target difference between each numerical value of the fixed-point target vector and the maximum numerical value; for the target difference between the each numerical value and the maximum numerical value, searching for a first median value corresponding to the each numerical value from the target mapping relationship according to the target difference, so as to generate a median value vector corresponding to the fixed-point target vector, wherein the target mapping relationship includes a correspondence relationship between different differences within a preset difference range corresponding to the fixed-point target vector having the first bit number and first median values, wherein each first median value is floating-point data having the second bit number, and the floating-point attention weight having the second bit number corresponding to each character is determined according to the median value vector.

3 . The method according to claim 2 , wherein determining the floating-point attention weight having the second bit number corresponding to each character according to the median value vector comprises:

acquiring a quantization scale output by the quantization node;

calculating a second median value according to a target difference numerical value corresponding to each numerical value in the fixed-point target vector and the quantization scale, wherein a data type of the second median value is a same type as that of the first median value; and

determining the floating-point attention weight according to the median value vector and the second median value.

4 . The method according to claim 3 , wherein the determining the floating-point attention weight according to the median value vector and the second median value comprises:

acquiring a ratio of each numerical value in the median value vector to the second median value to obtain the floating-point attention weight.

5 . The method according to claim 1 , wherein the preset intention recognition quantization model is obtained in advance by:

acquiring a plurality of text sample data, wherein the text sample data comprises intention labeling data;

training a preset initial model through the plurality of text sample data to obtain a first undetermined model, wherein the preset initial model comprises at least one multi-head attention layer, and the multi-head attention layer comprises a Matmul node and a Softmax node;

adding a first dequantization node after the Matmul node, and adding a quantization node between the first dequantization node and the Softmax node, and a second dequantization node coupled with an output end of the quantization node to obtain a second undetermined model, wherein the first dequantization node is configured to dequantize fixed-point data having a third bit number output by the Matmul node to obtain a floating-point data having the second bit number, the quantization node is configured to quantize the floating-point data having the second bit number output by the first dequantization node to obtain a corresponding fixed-point data having the first bit number, and the second dequantization node is configured to dequantize the corresponding fixed-point data having the first bit number output by the quantization node to obtain the floating-point data having the second bit number corresponding to the corresponding fixed-point data having the first bit number;

training the second undetermined model through the plurality of text sample data to obtain a third undetermined model;

obtaining a quantization scale output by the quantization node from the third undetermined model;

generating the look-up table node according to the quantization scale; and

replacing the second dequantization node and the Softmax node in the third undetermined model with the look-up table node to obtain the preset intention recognition quantization model.

6 . The method according to claim 5 , wherein the generating the look-up table node according to the quantization scale comprises:

calculating a first median value corresponding to each target numerical value within a numerical value range corresponding to the fixed-point target vector having the first bit number according to the quantization scale, wherein the first median value is floating-point data having the second bit number;

generating the target mapping relationship according to a correspondence relationship between each target numerical value and the first median value; and

establishing the look-up table node including the target mapping relationship.

7 . The method according to claim 1 , wherein the target mapping relationship is a correspondence between storage addresses and data stored in the storage addresses, or the target mapping relationship is a correspondence between a list number and data stored in a list of the list number.

8 . A non-transitory computer-readable medium on which a computer program is stored, wherein the computer program, when executed by a processing device, realizes steps of an intention recognition method, comprising:

acquiring a target text to be recognized; and

taking the target text as an input of a preset intention recognition quantization model to output a target intention of the target text;

wherein the preset intention recognition quantization model is configured to perform a quantization operation on a dot product of a query vector and a key vector corresponding to each character in the target text to generate a fixed-point target vector having a first bit number, determine a floating-point attention weight having a second bit number corresponding to each character in the target text through looking up a target mapping relationship, which is generated and stored in a computer in advance, according to the fixed-point target vector, and generate the target intention corresponding to the target text according to the floating-point attention weight, wherein the first bit number is smaller than the second bit number;

wherein the preset intention recognition quantization model comprises at least one multi-head attention layer, and the multi-head attention layer comprises a look-up table node; and

the look-up table node is configured to receive the fixed-point target vector, take the fixed-point target vector as an input to look up the target mapping relationship to determine a storage address corresponding to the fixed-point target vector, and access the storage address to determine the floating-point attention weight having the second bit number corresponding to each character;

wherein the multi-head attention layer further comprises a quantization node, and the look-up table node is coupled with an output end of the quantization node;

the quantization node is configured to perform the quantization operation on the dot product of the query vector and the key vector corresponding to each character in the target text to generate the fixed-point target vector having the first bit number, and input the fixed-point target vector into the look-up table node;

wherein the multi-head attention layer further comprises a Matmul node and a first dequantization node coupled with an output end of the Matmul node, and an output end of the first dequantization node is coupled with an input end of the quantization node;

the Matmul node is configured to acquire a query matrix and a key matrix corresponding to a target text sequence to be recognized, wherein the query matrix consists of query vectors corresponding to respective characters in the target text, and the key matrix consists of key vectors corresponding to the respective characters in the target text;

the Matmul node is further configured to obtain a target product of the query matrix and the key matrix, so as to obtain a specified matrix formed by the dot product of the query vector and the key vector corresponding to each character in the target text, wherein the specified matrix is fixed-point data having a third bit number, and the third bit number is greater than the first bit number;

the first dequantization node is configured to dequantize the specified matrix to obtain a floating-point product matrix comprising data having the second bit number; and

the quantization node is configured to quantize the floating-point product matrix to obtain a target matrix, and the target matrix includes fixed-point target vectors corresponding to the respective characters.

9 . An electronic apparatus comprising:

a storage device on which one or more computer programs are stored; and

one or more processing devices for executing the one or more computer programs in the storage device to realize steps of an intention recognition method, comprising:

acquiring a target text to be recognized; and

taking the target text as an input of a preset intention recognition quantization model to output a target intention of the target text;

wherein the preset intention recognition quantization model is configured to perform a quantization operation on a dot product of a query vector and a key vector corresponding to each character in the target text to generate a fixed-point target vector having a first bit number, determine a floating-point attention weight having a second bit number corresponding to each character in the target text through looking up a target mapping relationship, which is generated and stored in a computer in advance, according to the fixed-point target vector, and generate the target intention corresponding to the target text according to the floating-point attention weight, wherein the first bit number is smaller than the second bit number;

wherein the preset intention recognition quantization model comprises at least one multi-head attention layer, and the multi-head attention layer comprises a look-up table node; and

the look-up table node is configured to receive the fixed-point target vector, take the fixed-point target vector as an input to look up the target mapping relationship to determine a storage address corresponding to the fixed-point target vector, and access the storage address to determine the floating-point attention weight having the second bit number corresponding to each character;

the multi-head attention layer further comprises a quantization node, and the look-up table node is coupled with an output end of the quantization node;

the quantization node is configured to perform the quantization operation on the dot product of the query vector and the key vector corresponding to each character in the target text to generate the fixed-point target vector having the first bit number, and input the fixed-point target vector into the look-up table node;

wherein the multi-head attention layer further comprises a Matmul node and a first dequantization node coupled with an output end of the Matmul node, and an output end of the first dequantization node is coupled with an input end of the quantization node;

the Matmul node is configured to acquire a query matrix and a key matrix corresponding to a target text sequence to be recognized, wherein the query matrix consists of query vectors corresponding to respective characters in the target text, and the key matrix consists of key vectors corresponding to the respective characters in the target text;

the Matmul node is further configured to obtain a target product of the query matrix and the key matrix, so as to obtain a specified matrix formed by the dot product of the query vector and the key vector corresponding to each character in the target text, wherein the specified matrix is fixed-point data having a third bit number, and the third bit number is greater than the first bit number;

the first dequantization node is configured to dequantize the specified matrix to obtain a floating-point product matrix comprising data having the second bit number; and

the quantization node is configured to quantize the floating-point product matrix to obtain a target matrix, and the target matrix includes fixed-point target vectors corresponding to the respective characters.

10 . The apparatus according to claim 9 , wherein the look-up table node is configured to receive the fixed-point target vector, take the fixed-point target vector as an input to look up the target mapping relationship to determine a storage address corresponding to the fixed-point target vector, and access the storage address to determine the floating-point attention weight having the second bit number corresponding to each character, including:

the look-up table node obtaining a maximum numerical value of the fixed-point target vector corresponding to each character in the target matrix, and obtaining a target difference between each numerical value of the fixed-point target vector and the maximum numerical value; for the target difference between the each numerical value and the maximum numerical value, searching for a first median value corresponding to the each numerical value from the target mapping relationship according to the target difference, so as to generate a median value vector corresponding to the fixed-point target vector, wherein the target mapping relationship includes a correspondence relationship between different differences within a preset difference range corresponding to the fixed-point target vector having the first bit number and first median values, wherein each first median value is floating-point data having the second bit number, and the floating-point attention weight having the second bit number corresponding to each character is determined according to the median value vector.

11 . The apparatus according to claim 10 , wherein determining the floating-point attention weight having the second bit number corresponding to each character according to the median value vector comprises:

acquiring a quantization scale output by the quantization node;

calculating a second median value according to a target difference numerical value corresponding to each numerical value in the fixed-point target vector and the quantization scale, wherein a data type of the second median value is a same type as that of the first median value; and

determining the floating-point attention weight according to the median value vector and the second median value.

12 . The apparatus according to claim 11 , wherein the determining the floating-point attention weight according to the median value vector and the second median value comprises:

acquiring a ratio of each numerical value in the median value vector to the second median value to obtain the floating-point attention weight.

13 . The apparatus according to claim 9 , wherein the preset intention recognition quantization model is obtained in advance by:

acquiring a plurality of text sample data, wherein the text sample data comprises intention labeling data;

training a preset initial model through the plurality of text sample data to obtain a first undetermined model, wherein the preset initial model comprises at least one multi-head attention layer, and the multi-head attention layer comprises a Matmul node and a Softmax node;

adding a first dequantization node after the Matmul node, and adding a quantization node between the first dequantization node and the Softmax node, and a second dequantization node coupled with an output end of the quantization node to obtain a second undetermined model, wherein the first dequantization node is configured to dequantize fixed-point data having a third bit number output by the Matmul node to obtain a floating-point data having the second bit number, the quantization node is configured to quantize the floating-point data having the second bit number output by the first dequantization node to obtain a corresponding fixed-point data having the first bit number, and the second dequantization node is configured to dequantize the corresponding fixed-point data having the first bit number output by the quantization node to obtain the floating-point data having the second bit number corresponding to the corresponding fixed-point data having the first bit number;

training the second undetermined model through the plurality of text sample data to obtain a third undetermined model;

obtaining a quantization scale output by the quantization node from the third undetermined model;

generating the look-up table node according to the quantization scale; and

replacing the second dequantization node and the Softmax node in the third undetermined model with the look-up table node to obtain the preset intention recognition quantization model.

14 . The apparatus according to claim 13 , wherein the generating the look-up table node according to the quantization scale comprises:

calculating a first median value corresponding to each target numerical value within a numerical value range corresponding to the fixed-point target vector having the first bit number according to the quantization scale, wherein the first median value is floating-point data having the second bit number;

generating the target mapping relationship according to a correspondence relationship between each target numerical value and the first median value; and

establishing the look-up table node including the target mapping relationship.

15 . The apparatus according to claim 9 , wherein the target mapping relationship is a correspondence between storage addresses and data stored in the storage addresses, or the target mapping relationship is a correspondence between a list number and data stored in a list of the list number.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2025
From: LI, XIAOYANG; YU, ZILIN; MA, ZEJUN
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073330/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2025
From: TIAN, XIAOGANG
To: SHANGHAI SUIXUNTONG ELECTRONIC TECHNOLOGY CO., LTD.
Reel/Frame 073330/0872 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 30, 2025
From: SHANGHAI SUIXUNTONG ELECTRONIC TECHNOLOGY CO., LTD.
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073330/0883 →
Priority Claims (1)
CN 202111402778.1 · Nov 19, 2021 · national
Continuity (2)
Continuation PCTCN2022132141 · Nov 16, 2022
Related Publication 20240185046A1 · Jun 6, 2024
References Cited (20)
US 20200210839A1 · Lo · 2020 [cited by examiner]
US 20200242302A1 · Liang et al. · 2020 [cited by applicant]
US 20210248450A1 · Tay · 2021 [cited by examiner]
US 20230028226A1 · Thorsley · 2023 [cited by examiner]
US 20230154221A1 · Gu · 2023 [cited by examiner]
US 20240104342A1 · Li · 2024 [cited by examiner]
CN 110377686A · 2019 [cited by applicant]
CN 111159346A · 2020 [cited by applicant]
CN 112287672A · 2021 [cited by applicant]
CN 112313642A · 2021 [cited by applicant]
CN 113255354A · 2021 [cited by applicant]
CN 114090740A · 2022 [cited by applicant]
WO 2020140612A · 2020 [cited by applicant]
WO 2021204017A1 · 2021 [cited by applicant]
Wenxiao Wang, Wei Chen, Yicong Luo, Yongliu Long, Zhengkai Lin, Liye Zhang, Binbin Lin, Deng Cai, & Xiaofei He. (2024). Model Compression and Efficient Inference for Large Language Models: A Survey. https://arxiv.org/ab… [cited by examiner]
Shen, Sheng, et al. “Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT.” arXiv preprint arXiv:1909.05840 (2019) (Year: 2019). [cited by examiner]
Ham, Tae Jun, et al. “A3: Accelerating attention mechanisms in neural networks with approximation.” 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 2020, pp. 328-341 (Year: 2020… [cited by examiner]
International Search Report for PCT/CN2022/132141, mailed Jan. 28, 2023, 6 pages. [cited by applicant]
Meng et al., “Query Intention Recognition Model Based on Character Level Cyclic Network”, Computer Engineering, vol. 43, No. 3, Mar. 2017, 6 pages, English translation of Abstract. [cited by applicant]
Office Action for Chinese Patent Application No. 202111402778.1, mailed Mar. 23, 2023, 9 pages. [cited by applicant]