IP Library Granted Patent US 12,737,689
Granted Patent B2
US 12,737,689 · App. 18/784,068 · Granted Sep 15, 2026

Scale-permuted machine learning architecture

Inventors: Xianzhi Du (San Jose, CA); Yin Cui (Mountain View, CA); Tsung-Yi Lin (Sunnyvale, CA); Quoc V. Le (Sunnyvale, CA); Pengchong Jin (Mountain View, CA); Mingxing Tan (Newark, CA); Golnaz Ghiasi (Mountain View, CA); Xiaodan Song (Cupertino, CA)
Assignee: GOOGLE LLC
G06N20/00G06F11/3495G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,689
App. No.
18/784,068
Granted
Sep 15, 2026
Kind
B2
Abstract

A computer-implemented method of generating scale-permuted models can generate models having improved accuracy and reduced evaluation computational requirements. The method can include defining, by a computing system including one or more computing devices, a search space including a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective scale. The method can include performing, by the computing system, a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space, the scale-permuted model based at least in part on a candidate permutation of the plurality of candidate permutations.

Claims (49)

1 . A computing system, comprising:

a machine-learned scale-permuted model, the machine-learned scale-permuted model comprising:

a scale-permuted network generated through permutation of a plurality of feature blocks, the scale-permuted network comprising the plurality of feature blocks arranged in a scale-permuted sequence such that a resolution of the plurality of feature blocks varies nonmonotonically throughout the scale-permuted sequence;

wherein the scale-permuted network comprises a sequence of blocks comprising:

a first feature block in the sequence having a first resolution,

a second feature block next in the sequence after the first feature block, the second feature block having a second resolution higher than the first resolution, and

a third feature block next in the sequence after the second feature block, the third feature block having a third resolution lower than the second resolution and different from the first resolution;

one or more processors; and

one or more memory devices storing computer-readable instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:

obtaining input data, the input data comprising an input tensor;

providing the input data to the machine-learned scale-permuted model; and

receiving, as output from the machine-learned scale-permuted model, output data.

2 . The computing system of claim 1 , wherein the machine-learned scale-permuted model comprises one or more cross-scale connections configured to connect a parent block of the plurality of feature blocks, the parent block having a fourth resolution, to a target block of the plurality of feature blocks, the target block having a fifth resolution different from the fourth resolution.

3 . The computing system of claim 2 , wherein the cross-scale connection comprises a scaling factor.

4 . The computing system of claim 1 , wherein the machine-learned scale-permuted model comprises a stem network, the stem network comprising a plurality of feature blocks arranged in a scale-decreasing sequence.

5 . The computing system of claim 1 , wherein the machine-learned scale-permuted model comprises a task-specific combination model.

6 . The computing system of claim 1 , wherein the plurality of feature blocks comprises one or more weight layers, at least one activation function layer, and at least one pooling layer.

7 . The computing system of claim 1 , wherein the scale-permuted model comprises a fourth feature block ordered subsequent to the third feature block, a resolution of the fourth feature block being higher than the third resolution of the third feature block.

8 . The computing system of claim 1 , wherein the scale-permuted model is based at least in part on a candidate permutation of a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective resolution.

9 . The computing system of claim 8 , wherein the candidate permutation was selected by performing a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space.

10 . One or more memory devices storing:

a machine-learned scale-permuted model, the machine-learned scale-permuted model comprising:

a scale-permuted network generated through permutation of a plurality of feature blocks, the scale-permuted network comprising the plurality of feature blocks arranged in a scale-permuted sequence such that a resolution of the plurality of feature blocks varies nonmonotonically throughout the scale-permuted sequence;

wherein the scale-permuted network comprises a sequence of blocks comprising:

a first feature block in the sequence having a first resolution,

a second feature block next in the sequence after the first feature block, the second feature block having a second resolution higher than the first resolution, and

a third feature block next in the sequence after the second feature block, the third feature block having a third resolution lower than the second resolution and different from the first resolution;

computer-readable instructions that, when implemented, cause one or more processors to perform operations, the operations comprising:

obtaining input data, the input data comprising an input tensor;

providing the input data to the machine-learned scale-permuted model; and

receiving, as output from the machine-learned scale-permuted model, output data.

11 . The one or more memory devices of claim 10 , wherein the machine-learned scale-permuted model comprises one or more cross-scale connections configured to connect a parent block of the plurality of feature blocks, the parent block having a fourth resolution, to a target block of the plurality of feature blocks, the target block having a fifth resolution different from the fourth resolution.

12 . The one or more memory devices of claim 11 , wherein the cross-scale connection comprises a scaling factor.

13 . The one or more memory devices of claim 10 , wherein the machine-learned scale-permuted model comprises a stem network, the stem network comprising a plurality of feature blocks arranged in a scale-decreasing sequence.

14 . The one or more memory devices of claim 10 , wherein the machine-learned scale-permuted model comprises a task-specific combination model.

15 . The one or more memory devices of claim 10 , wherein the plurality of feature blocks comprises one or more weight layers, at least one activation function layer, and at least one pooling layer.

16 . The one or more memory devices of claim 10 , wherein the scale-permuted model comprises a fourth feature block ordered subsequent to the third feature block, a resolution of the fourth feature block being higher than the third resolution of the third feature block.

17 . A computer-implemented method, comprising:

obtaining input data, the input data comprising an input tensor;

providing the input data to a machine-learned scale-permuted model, the machine-learned scale-permuted model comprising:

a scale-permuted network generated through permutation of a plurality of feature blocks, the scale-permuted network comprising the plurality of feature blocks arranged in a scale-permuted sequence such that a resolution of the plurality of feature blocks varies nonmonotonically throughout the scale-permuted sequence;

wherein the scale-permuted network comprises a sequence of blocks comprising:

a first feature block in the sequence having a first resolution,

a second feature block next in the sequence after the first feature block, the second feature block having a second resolution higher than the first resolution, and

a third feature block next in the sequence after the second feature block, the third feature block having a third resolution lower than the second resolution and different from the first resolution; and

receiving, as output from the machine-learned scale-permuted model, output data.

18 . The computer-implemented method of claim 17 , wherein the machine-learned scale-permuted model comprises one or more cross-scale connections configured to connect a parent block of the plurality of feature blocks, the parent block having a fourth resolution, to a target block of the plurality of feature blocks, the target block having a fifth resolution different from the fourth resolution.

19 . The computer-implemented method of claim 18 , wherein the cross-scale connection comprises a scaling factor.

20 . The computer-implemented method of claim 17 , wherein the machine-learned scale-permuted model comprises a stem network, the stem network comprising a plurality of feature blocks arranged in a scale-decreasing sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2024
From: DU, XIANZHI; CUI, YIN; SONG, XIAODAN; LIN, TSUNG-YI; LE, QUOC V.; JIN, PENGCHONG; TAN, MINGXING; GHIASI, GOLNAZ
To: GOOGLE LLC
Reel/Frame 068519/0743 →
Continuity (2)
Division 17061355 · Oct 1, 2020
Related Publication 20240378509A1 · Nov 14, 2024
References Cited (57)
US 11748615B1 · Wu · 2023 [cited by examiner]
US 20170360401A1 · Rothberg et al. · 2017 [cited by applicant]
US 20200167586A1 · Gao et al. · 2020 [cited by applicant]
US 20200401871A1 · Tseng · 2020 [cited by applicant]
US 20210019151A1 · Pudipeddi et al. · 2021 [cited by applicant]
Liu et al. Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation. 2019 (Year: 2019). [cited by examiner]
Chen et al., “Detnas: Backbone Search for Object Detection”, 2019 Conference on Neural Information Processing Systems, Vancouver, Canada, Dec. 8-14, 2019, 11 pages. [cited by applicant]
Chen et al., “Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation”, arXiv:1802.02611v3, Aug. 22, 2018, 18 pages. [cited by applicant]
Chollet., “Xception: Deep Learning with Depthwise Separable Convolutions”, 2009 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Honolulu, Hawaii, United States, J… [cited by applicant]
Deng et al., “ImageNet: A Large-Scale Hierarchical Image Database”, 2009 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Miami Beach, Florida, United States, Jun.… [cited by applicant]
Elsken et al. “Neural Architecture Search: A Survey”, Journal of Machine Learning Research, vol. 20, 2019, 21 pages. [cited by applicant]
Ghiasi et al., “DropBlock: A Regularization Method for Convolutional Networks”, Thirty-Second International Conference on Neural Information Processing Systems, Montreal, Canada, Dec. 3-8, 2018, 11 pages. [cited by applicant]
Ghiasi et al., “NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 16-20, 2019, Long Beach, California, United States… [cited by applicant]
Ghosh et al., “Supervised Dimensionality Reduction and Visualization using Centroid-Encoder”, Journal of Machine Learning Research, vol. 23, No. 20, 2022, 34 pages. [cited by applicant]
Goyal et al., “Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour”, arXiv:1706.02677v1, Jun. 8, 2017, 12 pages. [cited by applicant]
He et al., “Bag of Tricks for Image Classification with Convolutional Neural Networks”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long beach, California, United States, Jun. 16-20, 2019, pp. 5… [cited by applicant]
He et al., “Deep Residual Leaming for Image Recognition”, 2016 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Las Vegas, Nevada, United States, Jun. 26-Jul. 1, 2… [cited by applicant]
He et al., “Mask R-CNN”, 2017 Institute of Electrical and Electronics Engineers International Conference on Computer Vision, Venice, Italy, Oct. 22-29, 2017, pp. 2961-2969. [cited by applicant]
He et al., “Rethinking ImageNet Pre-training”, 2019 IEEE/CVF International Conference on Computer Vision, Seoul, South Korea, Oct. 27-Nov. 2, 2019, pp. 4918-4927. [cited by applicant]
Howard et al., “Searching for MobileNetV3”, 2019 IEEE/CVF International Conference on Computer Vision, Seoul, South Korea, Oct. 27-Nov. 2, 2019, pp. 1314-1324. [cited by applicant]
Hu et al., “Squeeze-and-Excitation Networks”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, United States, Jun. 18-22, 2018, pp. 7132-7141. [cited by applicant]
Huang et al., “Deep Networks with Stochastic Depth”, arXiv:1603.09382v3, Jul. 16, 2016, 16 pages. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks”, 2017 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Honolulu, Hawaii, United States, Jul. 21-26, 2017, … [cited by applicant]
Huang et al., “GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism”, arXiv:1811.06965v4, Dec. 12, 2018, 11 pages. [cited by applicant]
Huang et al.. “Speed/Accuracy Trade-offs for Modern Convolutional Object Detectors”, 2017 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Honolulu, Hawaii, United… [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks”, Twenty-Sixth Annual Conference on Neural Information Processing Systems, Lake Tahoe, Nevada, United States, Dec. 3-8, 2012, 9 pages. [cited by applicant]
Lecun et al., “Backpropagation Applied to Handwritten Zip Code Recognition”, Neural Computation, vol. 1, No. 4, Dec. 1989, pp. 541-551. [cited by applicant]
Li et al., “DetNet: Design Backbone for Object Detection”, 2018 European Conference on Computer Vision, Munich, Germany, Sep. 8-14, 2018, 17 pages. [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection”, 2017 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Honolulu, Hawaii, United States, Jul. 21-26, 201… [cited by applicant]
Lin et al., “Focal Loss for Dense Object Detection”, 2017 Institute of Electrical and Electronics Engineers International Conference on Computer Vision, Venice, Italy, Oct. 22-29, 2017, pp. 2980-2988. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context”, arXiv:1405.0312v2, Jul. 5, 2014, 14 pages. [cited by applicant]
Liu et al., “Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, California, United States, Jun. 15-20… [cited by applicant]
Liu et al., “Darts: Differentiable Architecture Search”, Seventh International Conference on Learning Representations, New Orleans, Louisiana, United States, May 6-9, 2019, 13 pages. [cited by applicant]
Liu et al., “Progressive Neural Architecture Search”, Fifteenth European Conference on Computer Vision, Munich, Germany, Sep. Aug. 14, 2018, 16 pages. [cited by applicant]
Medium, “Deep Belief Networks—All you Need to Know”, Aug. 1, 2018, https://icecreamlabs.medium.com/deep-belief-networks-all-you-need-to-know-68aa9a71cc53,on Apr. 2, 2025, 8 pages. [cited by applicant]
Medium, “Deep Learning-Deep Belief Network (DBN)”, Dec. 12, 2018, https://medium.datadriveninvestor.com/deep-learning-deep-belief-network-dbn-ab715b5b8afc, retrieved on Apr. 2, 2025, 7 pages. [cited by applicant]
Medium, “Everything you Need to Know About AutoML and Neural Architecture Search”, Aug. 21, 2018, https://towardsdatascience.com/everything-you-need-to-know-about-automl-and-neural-architecture-search-8db1863682bf, retr… [cited by applicant]
Newell et al., “Stacked Hourglass Networks for Human Pose Estimation”, Fourteenth European Conference on Computer Vision, Amsterdam, The Netherlands, Oct. 8-16, 2016, pp. 483-499. [cited by applicant]
Ramachandran et al., “Searching for Activation Functions”, arXiv:1710.05941v2, Oct. 27, 2017, 13 pages. [cited by applicant]
Real et al., “Regularized Evolution for Image Classifier Architecture Search”, Thirty-Third Association for the Advancement of Artificial Intelligence Conference on Artificial Intelligence, Honolulu, Hawaii, United Stat… [cited by applicant]
Redmon et al., “YOLOv3: An Incremental Improvement”, arXiv:1804.02767v1, Apr. 8, 2018, 6 pages. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge”, International Journal of Computer Vision, vol. 115, Dec. 1, 2015, pp. 211-252. [cited by applicant]
Sandler et al., “Mobilenetv2: Inverted Residuals and Linear Bottlenecks”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, United States, Jun. 18-23, 2018, pp. 4510-4520. [cited by applicant]
Sonoda et al., “Transport Analysis of Infinitely Deep Neural Network”, Journal of Machine Learning Research, vol. 20, 2019, 52 pages. [cited by applicant]
Sun et al., “FishNet: A Versatile Backbone for Image, Region, and Pixel Level Prediction”, Thirthy-Second International Conference on Neural Information Processing Systems, Montreal, Canada, Dec. 3-8, 2018, 11 pages. [cited by applicant]
Szegedy et al., “Going Deeper with Convolutions”, 2015 Institute of Electrical and Electronics Engineers Conference on Computer Vision and Pattern Recognition, Boston, Massachusetts, United States, Jun. 7-12, 2015, 9 pa… [cited by applicant]
Szegedy et al., “Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning”, Thirty-First Association for the Advancement of Artificial Intelligence Conference on Artificial Intelligence, San Fra… [cited by applicant]
Szegedy et al., “Rethinking the Inception Architecture for Computer Vision”, arXiv:1512.00567v3, Dec. 11, 2015, 10 pages. [cited by applicant]
Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”, arXiv:1905.11946v3, Nov. 23, 2019, 10 pages. [cited by applicant]
Tan et al., “MnasNet: Platform-Aware Neural Architecture Search for Mobile”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, California, United States, Jun. 15-20, 2019, pp. 2820-2828. [cited by applicant]
Van Horn et al., “The iNaturalist Species Classification and Detection Dataset”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, United States, Jun. 18-23, 2018, pp. 8769-8778. [cited by applicant]
Wang et al., “Deep High-Resolution Representation Learning for Visual Recognition”, arXiv:1908.07919v2, Mar. 13, 2020, 23 pages. [cited by applicant]
Xie et al., “Exploring Randomly Wired Neural Networks for Image Recognition”, arXiv:1904.01569v2, Apr. 8, 2019, 10 pages. [cited by applicant]
Xu et al., “Auto-FPN: Automatic Network Architecture Adaptation for Object Detection Beyond Classification”, 2019 IEEE/CVF International Conference on Computer Vision, Oct. 27- Nov. 2, 2019, Seoul, South Korea, pp. 6649… [cited by applicant]
Zagoruyko et al., “Wide Residual Networks”, 2016 British Machine Vision Conference, York, United Kingdom, Sep. 19-22, 2016, 12 pages. [cited by applicant]
Zoph et al., “Learning Transferable Architectures for Scalable Image Recognition”, arXiv:1707.07012v4, Apr. 11, 2018, 14 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search with Reinforcement Learning”, Fifth International Conference on Leaming Representations, Toulon, France, Apr. 24-26, 2017, 16 pages. [cited by applicant]