IP Library › Granted Patent US 12,530,624
Granted Patent B2
US 12,530,624 · App. 18/168,723 · Granted Jan 20, 2026

Provisioning resource-efficient artificial intelligence models

Inventors: Yao Yang (Sunnyvale, CA); David Nguyen (Newark, CA)
Assignee: ACCENTURE GLOBAL SOLUTIONS LIMITED
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,624
App. No.
18/168,723
Granted
Jan 20, 2026
Kind
B2
Abstract

Implementations for receiving first user input representative of a first accuracy-to-resource value for an AI model, determining a first training recipe for training of the AI model, the first training recipe including a first set of reduction strategies to be performed during training of the AI model, the first training recipe being determined through genetic search of an initial population to provide an updated population, the first set of reduction strategies being selected from the updated population, providing the first training recipe for training of the AI model to provide a first trained version of the AI model at least partially by executing one or more of pruning and quantization during training of the AI model, and outputting the first trained version of the AI model for inference.

Claims (52)

1 . A computer-implemented method for provisioning resource-efficient artificial intelligence (AI) models, the method comprising:

receiving first user input representative of a first accuracy-to-resource value for an AI model;

determining a first training recipe for training of the AI model, the first training recipe comprising a first set of reduction strategies to be performed during training of the AI model, the first training recipe being determined through genetic search of an initial population to provide an updated population, the first set of reduction strategies being selected from the updated population;

providing the first training recipe for training of the AI model to provide a first trained version of the AI model at least partially by executing one or more of pruning and quantization during training of the AI model; and

outputting the first trained version of the AI model for inference.

2 . The method of claim 1 , wherein genetic search comprises:

determining a fitness score for each set of reduction strategies in a plurality of reduction strategies;

selecting two sets of reduction strategies from the plurality of reduction strategies based on fitness scores;

generating offspring using the two sets of reduction strategies; and

forming the updated population as comprising the two sets of reduction strategies and the offspring.

3 . The method of claim 2 , wherein genetic search further comprises applying one or more mutations to each offspring.

4 . The method of claim 1 , wherein the first set of reduction strategies is selected from the updated population as having a highest fitness score among other sets of reduction strategies in the updated population.

5 . The method of claim 1 , wherein pruning comprises one or more of structured pruning and unstructured pruning.

6 . The method of claim 1 , wherein genetic search is executed until a stop condition is reached.

7 . The method of claim 1 , further comprising:

receiving second user input representative of a second accuracy-to-resource value for the AI model;

determining a second training recipe for training of the AI model, the second training recipe comprising a second set of reduction strategies to be performed during training of the AI model, the second set of reduction strategies being different from the first set of reduction strategies; and

providing the second training recipe for training of the AI model to provide a second trained version of the AI model, the second trained version of the AI model having being a different size from the first trained version of the AI model.

8 . A system, comprising:

one or more processors; and

a computer-readable storage device coupled to the one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for provisioning resource-efficient artificial intelligence (AI) models, the operations comprising:

receiving first user input representative of a first accuracy-to-resource value for an AI model;

determining a first training recipe for training of the AI model, the first training recipe comprising a first set of reduction strategies to be performed during training of the AI model, the first training recipe being determined through genetic search of an initial population to provide an updated population, the first set of reduction strategies being selected from the updated population;

providing the first training recipe for training of the AI model to provide a first trained version of the AI model at least partially by executing one or more of pruning and quantization during training of the AI model; and

outputting the first trained version of the AI model for inference.

9 . The system of claim 8 , wherein genetic search comprises:

determining a fitness score for each set of reduction strategies in a plurality of reduction strategies;

selecting two sets of reduction strategies from the plurality of reduction strategies based on fitness scores;

generating offspring using the two sets of reduction strategies; and

forming the updated population as comprising the two sets of reduction strategies and the offspring.

10 . The system of claim 9 , wherein genetic search further comprises applying one or more mutations to each offspring.

11 . The system of claim 8 , wherein the first set of reduction strategies is selected from the updated population as having a highest fitness score among other sets of reduction strategies in the updated population.

12 . The system of claim 8 , wherein pruning comprises one or more of structured pruning and unstructured pruning.

13 . The system of claim 8 , wherein genetic search is executed until a stop condition is reached.

14 . The system of claim 8 , wherein operations further comprise:

receiving second user input representative of a second accuracy-to-resource value for the AI model;

determining a second training recipe for training of the AI model, the second training recipe comprising a second set of reduction strategies to be performed during training of the AI model, the second set of reduction strategies being different from the first set of reduction strategies; and

providing the second training recipe for training of the AI model to provide a second trained version of the AI model, the second trained version of the AI model having being a different size from the first trained version of the AI model.

15 . Computer-readable storage media coupled to the one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for provisioning resource-efficient artificial intelligence (AI) models, the operations comprising:

receiving first user input representative of a first accuracy-to-resource value for an AI model;

determining a first training recipe for training of the AI model, the first training recipe comprising a first set of reduction strategies to be performed during training of the AI model, the first training recipe being determined through genetic search of an initial population to provide an updated population, the first set of reduction strategies being selected from the updated population;

providing the first training recipe for training of the AI model to provide a first trained version of the AI model at least partially by executing one or more of pruning and quantization during training of the AI model; and

outputting the first trained version of the AI model for inference.

16 . The computer-readable storage media of claim 15 , wherein genetic search comprises:

determining a fitness score for each set of reduction strategies in a plurality of reduction strategies;

selecting two sets of reduction strategies from the plurality of reduction strategies based on fitness scores;

generating offspring using the two sets of reduction strategies; and

forming the updated population as comprising the two sets of reduction strategies and the offspring.

17 . The computer-readable storage media of claim 16 , wherein genetic search further comprises applying one or more mutations to each offspring.

18 . The computer-readable storage media of claim 15 , wherein the first set of reduction strategies is selected from the updated population as having a highest fitness score among other sets of reduction strategies in the updated population.

19 . The computer-readable storage media of claim 15 , wherein pruning comprises one or more of structured pruning and unstructured pruning.

20 . The computer-readable storage media of claim 15 , wherein genetic search is executed until a stop condition is reached.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2023
From: YANG, YAO; NGUYEN, DAVID
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 062705/0465 →
Continuity (1)
Related Publication 20240273399A1 · Aug 15, 2024
References Cited (14)
US 11392829B1 · Pool · 2022 [cited by examiner]
US 12346818B2 · Zhang · 2025 [cited by examiner]
US 20250037028A1 · Tuli · 2025 [cited by examiner]
US 20250139501A1 · Tosi · 2025 [cited by examiner]
GitHub.com [online], “neuralmagic/sparsify,” available on or before Oct. 21, 2021 via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20211021000327/https://github.com/neuralmagic/sparsify>, retrieved … [cited by applicant]
GreenCarCongress.com [online], “Study projects global carbon footprint from ICT will be equivalent to half of transportation's current level by 2040,” Mar. 6, 2018, retrieved on Aug. 9, 2023, retrieved from URL<https://… [cited by applicant]
Hoefler, “Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks,” Journal of Machine Learning Research, Sep. 2021, 22:241, 124 pages. [cited by applicant]
Htor.Inf.Ethz.ch [online], “Sparsity in Deep Learning,” available on or before May 17, 2022 via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20220517181413/https://htor.inf.ethz.ch/sparsity-in-dl/>,… [cited by applicant]
Jiang et al., “Model Pruning Enables Efficient Federated Learning on Edge Devices,” CoRR, submitted on Apr. 6, 2022, arXiv:1909.12326v5, 22 pages. [cited by applicant]
NeuralMagic.com [online], “SparseML,” available on or before Dec. 9, 2022 via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20221209155102/https://docs.neuralmagic.com/products/sparseml/>, retrieved … [cited by applicant]
NeuralMagic.com [online], “Sparsify,” upon information and belief, available no later than Oct. 31, 2022, retrieved on Aug. 9, 2023, retrieved from URL<https://docs.neuralmagic.com/archive/sparsify/>, 12 pages. [cited by applicant]
PYPI.org [online], “Project Description: Sparsify,” available on or before Mar. 7, 2021 via Internet Archive: Wayback Machine URL<https://web.archive.org/web/20210307162826/https://pypi.org/project/sparsify/>, retrieved… [cited by applicant]
TowardsDataScience.com [online], “Introduction to Genetic Algorithms—Including Example Code,” Jul. 7, 2017, retrieved on Aug. 9, 2023, retrieved from URL<https://towardsdatascience.com/introduction-to-genetic-algorithms… [cited by applicant]
Wang et al., “Towards ultra-high performance and energy efficiency of deep learning systems: an algorithm-hardware co-optimization framework,” Presented at Proceedings of AAAI'18: AAAI Conference on Artificial Intellige… [cited by applicant]