IP Library Granted Patent US 12670363
Granted Patent B2
US 12670363 · App. 18/026,140 · Granted Jun 30, 2026

Mixture-of-experts model implementation method and system, electronic device, and storage medium

Inventors: Liang Shen (Beijing, CN); Haifeng Wang (Beijing, CN); Huachao Wu (Beijing, CN); Weibao Gong (Beijing, CN); Zhihua Wu (Beijing, CN); Dianhai Yu (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06N3/045G06N3/0495
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670363
App. No.
18/026,140
Filed
Mar 14, 2023
Granted
Jun 30, 2026
Kind
B2
Examiner
LU, HWEI-MIN
Art Unit
2142
USPC
706/15
Abstract

The present disclosure provides a mixture-of-experts (MoE) model implementation method and system, an electronic device, and a storage medium, and relates to the field of artificial intelligence (AI) such as deep learning and distributed storage. The method includes: constructing a communication group, the communication group including a tensor-parallelism communication group, the tensor-parallelism communication group including at least two computing devices, tensor-parallelism segmentation being adopted for sparse parameters of each of the computing devices in a same tensor-parallelism communication group; and training an MoE model based on the communication group. By use of the solutions of the present disclosure, normal operation of model training can be guaranteed.

Claims (24)

1 . A method of mixture-of-experts (MoE) model implementation, comprising:

constructing a communication group, the communication group comprising at least two tensor-parallelism communication groups and at least two data-parallelism communication groups, each of the at least two tensor-parallelism communication groups comprising at least two computing devices, tensor-parallelism segmentation being adopted for sparse parameters of each of the at least two computing devices in a same tensor-parallelism communication group, adopting the tensor-parallelism segmentation for dense parameters of each of the at least two computing devices in the same tensor-parallelism communication group; wherein the sparse parameters comprise: expert network parameters; and the dense parameters comprise: backbone network parameters, each of the at least two data-parallelism communication groups comprising two or more computing devices, data parallelism being adopted for each of the two or more computing devices in a same data-parallelism communication group, and for any one of the at least two tensor-parallelism communication groups, each of the at least two computing devices being comprised in a data-parallelism communication group, and a first computing device set formed by the computing devices comprised in all the at least two data-parallelism communication groups being equal to a second computing device set formed by the computing devices comprised in all the at least two tensor-parallelism communication groups; and

training an MoE model based on the communication group using speech data,

wherein any one of the at least two computing devices comprises:

a backbone network configured to perform a first predetermined processing on the speech data to obtain a first processing result of the speech data;

a gating network configured to select an expert network as a routing network from expert networks of the computing devices in the data-parallelism communication group, and send the first processing result to the selected expert network; and

the expert network configured to perform a second predetermined processing on the acquired first processing result to obtain a second processing result of the speech data, and return a return result determined according to the second processing result to the computing device corresponding to the acquired first processing result, the return result comprises: part of a content in the second processing result and the method further comprises: for any one of the at least two tensor-parallelism communication groups, combining acquired parts of the content belonging to a same second processing result to obtain a complete second processing result, the parts of the content being respectively returned by a same expert network corresponding to the computing devices in the same tensor-parallelism communication group, different computing devices in the same tensor-parallelism communication group correspond to a same backbone network and the same expert network.

2 . An electronic device, comprising:

at least one processor; and

a memory communicatively connected with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a method of mixture-of-experts (MoE) model implementation, wherein the method comprises:

constructing a communication group, the communication group comprising at least two tensor-parallelism communication groups and at least two data-parallelism communication groups, each of the at least two tensor-parallelism communication groups comprising at least two computing devices, tensor-parallelism segmentation being adopted for sparse parameters of each of the at least two computing devices in a same tensor-parallelism communication group, adopting the tensor-parallelism segmentation for dense parameters of each of the at least two computing devices in the same tensor-parallelism communication group; wherein the sparse parameters comprise: expert network parameters; and the dense parameters comprise: backbone network parameters, each of the at least two data-parallelism communication groups comprising two or more computing devices, data parallelism being adopted for each of the two or more computing devices in a same data-parallelism communication group, and for any one of the at least two tensor-parallelism communication groups, each of the at least two computing devices being comprised in a data-parallelism communication group, and a first computing device set formed by the computing devices comprised in all the at least two data-parallelism communication groups being equal to a second computing device set formed by the computing devices comprised in all the at least two tensor-parallelism communication groups; and

training an MoE model based on the communication group using speech data,

wherein any one of the at least two computing devices comprises:

a backbone network configured to perform a first predetermined processing on the speech data to obtain a first processing result of the speech data;

a gating network configured to select an expert network as a routing network from expert networks of the computing devices in the data-parallelism communication group, and send the first processing result to the selected expert network; and

the expert network configured to perform a second predetermined processing on the acquired first processing result to obtain a second processing result of the speech data, and return a return result determined according to the second processing result to the computing device corresponding to the acquired first processing result, the return result comprises: part of a content in the second processing result; and the method further comprises: for any one of the at least two tensor-parallelism communication groups, combining acquired parts of the content belonging to a same second processing result to obtain a complete second processing result, the parts of the content being respectively returned by a same expert network corresponding to the computing devices in the same tensor-parallelism communication group, different computing devices in the same tensor-parallelism communication group correspond to a same backbone network and the same expert network.

3 . A non-transitory computer readable storage medium with computer instructions stored thereon, wherein the computer instructions are used for causing a method of mixture-of-experts (MoE) model implementation, wherein the method comprises:

constructing a communication group, the communication group comprising at least two tensor-parallelism communication groups and at least two data-parallelism communication groups, each of the at least two tensor-parallelism communication groups comprising at least two computing devices, tensor-parallelism segmentation being adopted for sparse parameters of each of the at least two computing devices in a same tensor-parallelism communication group, adopting the tensor-parallelism segmentation for dense parameters of each of the at least two computing devices in the same tensor-parallelism communication group; wherein the sparse parameters comprise: expert network parameters; and the dense parameters comprise: backbone network parameters, each of the at least two data-parallelism communication groups comprising two or more computing devices, data parallelism being adopted for each of the two or more computing devices in a same data-parallelism communication group, and for any one of the at least two tensor-parallelism communication groups, each of the at least two computing devices being comprised in a data-parallelism communication group, and a first computing device set formed by the computing devices comprised in all the at least two data-parallelism communication groups being equal to a second computing device set formed by the computing devices comprised in all the at least two tensor-parallelism communication groups; and

training an MoE model based on the communication group using speech data,

wherein any one of the at least two computing devices comprises:

a backbone network configured to perform a first predetermined processing on the speech data to obtain a first processing result of the speech data;

a gating network configured to select an expert network as a routing network from expert networks of the computing devices in the data-parallelism communication group, and send the first processing result to the selected expert network; and

the expert network configured to perform a second predetermined processing on the acquired first processing result to obtain a second processing result of the speech data, and return a return result determined according to the second processing result to the computing device corresponding to the acquired first processing result, the return result comprises: part of a content in the second processing result; and the method further comprises: for any one of the at least two tensor-parallelism communication groups, combining acquired parts of the content belonging to a same second processing result to obtain a complete second processing result, the parts of the content being respectively returned by a same expert network corresponding to the computing devices in the same tensor-parallelism communication group, different computing devices in the same tensor-parallelism communication group correspond to a same backbone network and the same expert network.