IP Library Granted Patent US 11,797,277
Granted Patent B2
US 11,797,277 · App. 17/771,040 · Granted Oct 24, 2023

Neural network model conversion method server, and storage medium

Inventors: Chao Xiong (Shenzhen, CN); Kuenhung Tsoi (Shenzhen, CN); Xinyu Niu (Shenzhen, CN)
Assignee: Shenzhen Corerain Technologies Co., Ltd.
G06F8/427G06F8/74G06N3/063G06N3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,797,277
App. No.
17/771,040
Granted
Oct 24, 2023
Kind
B2
Abstract

A neural network model conversion method, a server, and a storage medium are provided according to embodiments of the present disclosure. The neural network model conversion method includes: parsing a neural network model to obtain initial model information; reconstructing the initial model information to obtain streaming model information; generating a target model information file according to the streaming model information; and running, under a streaming architecture, the neural network model according to the target model information file.

Claims (27)

1. A neural network model conversion method, comprising:

parsing a neural network model to obtain initial model information based on an instruction set architecture chip, the initial model information comprises an initial computation graph and initial model data, the initial computation graph comprises types of first operators and connection relationships among the first operators, and the initial model data comprises corresponding computation parameters of the first operators;

reconstructing the initial model information to obtain streaming model information based on a target streaming architecture chip, the streaming model information comprises a streaming computation graph and streaming model data, the streaming computation graph comprises types of second operators and connection relationships among the second operators, and the streaming model data comprises corresponding computation parameters of the second operators;

generating a target model information file according to the streaming model information, the target model information file is a file that stores neural network model information under the target streaming architecture chip;

generating a target model structure file according to the streaming computation graph, and generating a target model data file according to the streaming model data; and

in the target streaming architecture chip, constructing a target streaming architecture computation graph according to the target model structure file, importing the target model data file into the target streaming architecture computation graph; and running the target streaming architecture computation graph.

2. The method according to claim 1 , wherein the reconstructing the initial model information to obtain streaming model information comprises:

reconstructing an initial computation graph to obtain the streaming computation graph; and

reconstructing initial model data to obtain the streaming model data.

3. A server, comprising:

at least one processor; and

a storage equipment configured to store at least one program,

wherein the at least one program, when being executed by the at least one processor, causes the at least one processor to implement:

parsing a neural network model to obtain initial model information based on an instruction set architecture chip, the initial model information comprises an initial computation graph and initial model data, the initial computation graph comprises types of first operators and connection relationships among the first operators, and the initial model data comprises corresponding computation parameters of the first operators;

reconstructing the initial model information to obtain streaming model information based on a target streaming architecture chip, the streaming model information comprises a streaming computation graph and streaming model data, the streaming computation graph comprises types of second operators and connection relationships among the second operators, and the streaming model data comprises corresponding computation parameters of the second operators;

generating a target model information file according to the streaming model information, the target model information file is a file that stores neural network model information under the target streaming architecture chip;

generating a target model structure file according to the streaming computation graph, and generating a target model data file according to the streaming model data; and

in the target streaming architecture chip, constructing a target streaming architecture computation graph according to the target model structure file, importing the target model data file into the target streaming architecture computation graph; and running the target streaming architecture computation graph.

4. The server according to claim 3 , wherein the at least one program, when being executed by the at least one processor, causes the at least one processor to implement:

reconstructing an initial computation graph to extract the streaming computation graph; and

reconstructing initial model data to extract the streaming model data.

5. A non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when being executed by a processor, implements:

parsing a neural network model to obtain initial model information based on an instruction set architecture chip, the initial model information comprises an initial computation graph and initial model data, the initial computation graph comprises types of first operators and connection relationships among the first operators, and the initial model data comprises corresponding computation parameters of the first operators;

reconstructing the initial model information to obtain streaming model information based on a target streaming architecture chip, the streaming model information comprises a streaming computation graph and streaming model data, the streaming computation graph comprises types of second operators and connection relationships among the second operators, and the streaming model data comprises corresponding computation parameters of the second operators;

generating a target model information file according to the streaming model information, the target model information file is a file that stores neural network model information under the target streaming architecture chip;

generating a target model structure file according to the streaming computation graph, and generating a target model data file according to the streaming model data; and

in the target streaming architecture chip, constructing a target streaming architecture computation graph according to the target model structure file, importing the target model data file into the target streaming architecture computation graph; and running the target streaming architecture computation graph.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2022
From: XIONG, CHAO; TSOI, KUENHUNG; NIU, XINYU
To: SHENZHEN CORERAIN TECHNOLOGIES CO., LTD.
Reel/Frame 059673/0582 →
Continuity (1)
Related Publication 20220365762A1 · Nov 17, 2022