IP Library › Granted Patent US 12,524,266
Granted Patent B2
US 12,524,266 · App. 18/453,298 · Granted Jan 13, 2026

Dynamic batching for inference system for transformer-based generation tasks

Inventors: Gyeongin Yu (Seoul, KR); Geon-Woo Kim (Seoul, KR); Joo Seong Jeong (Seoul, KR); Soojeong Kim (Seoul, KR); Byung-Gon Chun (Seoul, KR)
Assignee: FRIENDLIAI CORP.
G06F9/4881G06F9/5016G06N5/04G06N20/00G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,266
App. No.
18/453,298
Granted
Jan 13, 2026
Kind
B2
Abstract

An inference system applies a machine-learning transformer model to a batch of requests with variable input length or variable target length or variable internal sate length by selectively batching a subset of operations in the transformer model but processing requests in the batch individually for a subset of operations in the transformer model. In one embodiment, the operation to be processed individually is an attention operation of an encoder or a decoder of the transformer model. By selective batching, the inference system can allow batching operations to be performed for a batch of requests with variable input or target length or internal state length to utilize the parallel computation capabilities of hardware accelerators while preventing unnecessary computations that occur for workarounds that restrain the data of a batch of requests to a same length.

Claims (63)

1 . A non-transitory computer-readable storage medium storing computer program instructions executable to perform operations comprising:

receiving, by a serving system, one or more requests for execution, the serving system including one or more execution engines each coupled to access a machine-learning transformer model;

scheduling a batch of requests including the one or more requests for execution on an execution engine;

generating, by the execution engine, a first set of output tokens by applying the transformer model to a first set of inputs for the batch of requests, wherein applying the transformer model comprises applying at least one batch operation to one or more input tensors associated with the batch of requests;

obtaining a sequence of input tokens for a new request;

scheduling a second batch of requests for execution on the execution engine, the second batch of requests including the new request and at least one request in the batch of requests, wherein a length of an internal state for the new request is a length of keys or values for the new request of an attention operation of the machine-learning transformer model, wherein a length of another internal state for the at least one request is a length of keys or values for the at least one request of the attention operation, wherein the length of the internal state for the new request is different from the length of the another internal state for the at least one request; and

generating, by the execution engine, a second set of output tokens by applying the transformer model to a second set of inputs for the second batch, wherein a key or value is added to the internal state for the new request and another key or value is added to the another internal state for the at least one request.

2 . The non-transitory computer-readable storage medium of claim 1 , wherein an input for the at least one request in the second set of inputs comprises at least two tokens.

3 . The non-transitory computer-readable storage medium of claim 1 , wherein the at least two tokens for the at least one request comprises one or more output tokens from the first set of output tokens.

4 . The non-transitory computer-readable storage medium of claim 1 , the operations further comprising:

iteratively generating a third set of output tokens for the at least one request; and

responsive to determining that the third set of output tokens reaches a threshold number, providing the third set of output tokens.

5 . The non-transitory computer-readable storage medium of claim 1 , wherein the second batch of requests is scheduled responsive to determining that the execution engine has memory available to execute the second batch of requests.

6 . The non-transitory computer-readable storage medium of claim 1 , wherein the at least one request is associated with a cache in the execution engine for storing elements of the another internal state for the at least one request, and responsive to determining that the at least one request has been completed, freeing the cache for the at least one request in the execution engine.

7 . The non-transitory computer-readable storage medium of claim 1 ,

wherein the internal state includes a key cache storing keys generated for the new request or a value cache storing values generated for the new request,

wherein the another internal state includes another key cache storing keys generated for the at least one request or another value cache storing values generated for the at least one request, and

wherein a length of the key cache is different from a length of the another key cache or a length of the value cache is different from a length of the another value cache.

8 . The non-transitory computer-readable storage medium of claim 7 , wherein the length of the key cache is a number of keys for the new request and the length of the another key cache is a number of keys for the at least one request, and wherein the length of the value cache is a number of values for the new request and the length of the another value cache is a number of values for the at least one request.

9 . The non-transitory computer-readable storage medium of claim 1 , wherein the length of the sequence of input tokens for the new request is different from the length of the another sequence of input tokens for the at least one request.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein the length of the sequence of input tokens is a number of input tokens in the sequence and the length of the another sequence of input tokens is a number of input tokens in the another sequence.

11 . The non-transitory computer-readable storage medium of claim 1 , wherein applying the at least one batch operation to the one or more input tensors associated with the batch of requests further comprises:

concatenating the one or more input tensors associated with the batch of requests into a concatenated input tensor; and

applying an operation for the batch operation to the concatenated input tensor to generate an output tensor.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein concatenating the one or more input tensors comprises concatenating elements of the one or more input tensors along one dimension.

13 . The non-transitory computer-readable storage medium of claim 1 , wherein each token in the sequence of input tokens represents a text unit.

14 . The non-transitory computer-readable storage medium of claim 1 , wherein applying the transformer model to the first set of inputs further comprises applying an attention operation to the first set of inputs, and wherein the attention operation is configured to obtain queries, keys, values for the first set of inputs and generate attention outputs by combining the queries, the keys, and the values.

15 . The non-transitory computer-readable storage medium of claim 1 , wherein applying the transformer model to the second set of inputs further comprises applying an attention operation to the second set of inputs, and wherein the attention operation is configured to obtain queries, keys, values for the second set of inputs and generate attention outputs by combining the queries, the keys, and the values.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein applying the attention operation to the second set of inputs further comprises:

generating a first attention output for the at least one request by combining a first query, a first key, and a first value for the at least one request, and

separately generating a second attention output for the new request by combining a second query, a second key, and a second value for the new request.

17 . The non-transitory computer-readable storage medium of claim 1 , wherein a length of the sequence of input tokens for the new request is different from a length of another sequence of input tokens for the at least one request.

18 . A method comprising:

receiving, by a serving system, one or more requests for execution, the serving system including one or more execution engines each coupled to access a machine-learning transformer model;

scheduling a batch of requests including the one or more requests for execution on an execution engine;

generating, by the execution engine, a first set of output tokens by applying the transformer model to a first set of inputs for the batch of requests, wherein applying the transformer model comprises applying at least one batch operation to one or more input tensors associated with the batch of requests;

obtaining a sequence of input tokens for a new request;

scheduling a second batch of requests for execution on the execution engine, the second batch of requests including the new request and at least one request in the batch of requests, wherein a length of an internal state for the new request is a length of keys or values for the new request of an attention operation of the machine-learning transformer model, wherein a length of another internal state for the at least one request is a length of keys or values for the at least one request of the attention operation, wherein the length of the internal state for the new request is different from the length of the another internal state for the at least one request; and

generating, by the execution engine, a second set of output tokens by applying the transformer model to a second set of inputs for the second batch, wherein a key or value is added to the internal state for the new request and another key or value is added to the another internal state for the at least one request.

19 . The method of claim 18 , wherein an input for the at least one request in the second set of inputs comprises at least two tokens.

20 . The method of claim 19 , wherein the at least two tokens for the at least one request comprises one or more output tokens from the first set of output tokens.

21 . The method of claim 18 , further comprising:

iteratively generating a third set of output tokens for the at least one request; and

responsive to determining that the third set of output tokens reaches a threshold number, providing the third set of output tokens.

22 . The method of claim 18 , wherein the second batch of requests is scheduled responsive to determining that the execution engine has memory available to execute the second batch of requests.

23 . The method of claim 18 , wherein the at least one request is associated with a cache in the execution engine for storing elements of the another internal state for the at least one request, and responsive to determining that the at least one request has been completed, freeing the cache for the at least one request in the execution engine.

24 . The method of claim 18 ,

wherein the internal state includes a key cache storing keys generated for the new request or a value cache storing values generated for the new request,

wherein the another internal state includes another key cache storing keys generated for the at least one request or another value cache storing values generated for the at least one request, and

wherein a length of the key cache is different from a length of the another key cache or a length of the value cache is different from a length of the another value cache.

25 . The method of claim 24 , wherein the length of the key cache is a number of keys for the new request and the length of the another key cache is a number of keys for the at least one request, and wherein the length of the value cache is a number of values for the new request and the length of the another value cache is a number of values for the at least one request.

26 . The method of claim 18 , wherein the length of the sequence of input tokens for the new request is different from the length of the another sequence of input tokens for the at least one request.

27 . The method of claim 26 , wherein the length of the sequence of input tokens is a number of input tokens in the sequence and the length of the another sequence of input tokens is a number of input tokens in the another sequence.

28 . The method of claim 18 , wherein applying the at least one batch operation to the one or more input tensors associated with the batch of requests further comprises:

concatenating the one or more input tensors associated with the batch of requests into a concatenated input tensor; and

applying an operation for the batch operation to the concatenated input tensor to generate an output tensor.

29 . The method of claim 28 , wherein concatenating the one or more input tensors comprises concatenating elements of the one or more input tensors along one dimension.

30 . The method of claim 18 , wherein each token in the sequence of input tokens represents a text unit.

31 . The method of claim 18 , wherein applying the transformer model to the first set of inputs further comprises applying an attention operation to the first set of inputs, and wherein the attention operation is configured to obtain queries, keys, values for the first set of inputs and generate attention outputs by combining the queries, the keys, and the values.

32 . The method of claim 18 , wherein applying the transformer model to the second set of inputs further comprises applying an attention operation to the second set of inputs, and wherein the attention operation is configured to obtain queries, keys, values for the second set of inputs and generate attention outputs by combining the queries, the keys, and the values.

33 . The method of claim 32 , wherein applying the attention operation to the second set of inputs further comprises:

generating a first attention output for the at least one request by combining a first query, a first key, and a first value for the at least one request, and separately generating a second attention output for the new request by combining a second query, a second key, and a second value for the new request.

34 . The method of claim 18 , wherein a length of the sequence of input tokens for the new request is different from a length of another sequence of input tokens for the at least one request.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2025
From: FRIENDLIAI INC.
To: FRIENDLIAI CORP.
Reel/Frame 072515/0423 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: YU, GYEONGIN; KIM, GEON-WOO; JEONG, JOO SEONG; KIM, SOOJEONG; CHUN, BYUNG-GON
To: FRIENDLIAI INC.; SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
Reel/Frame 065563/0105 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
To: FRIENDLIAI INC.
Reel/Frame 065563/0109 →
Continuity (3)
Continuation 17881549 · Aug 4, 2022
Continuation 17542193 · Dec 3, 2021
Related Publication 20240231902A1 · Jul 11, 2024
References Cited (125)
US 10217053B2 · Bordawekar et al. · 2019 [cited by applicant]
US 10846096B1 · Chung et al. · 2020 [cited by applicant]
US 10990650B1 · Vantrease et al. · 2021 [cited by applicant]
US 11074412B1 · Leeman-Munk · 2021 [cited by examiner]
US 11282160B2 · Barton et al. · 2022 [cited by applicant]
US 11442775B1 · Yu et al. · 2022 [cited by applicant]
US 11521042B2 · Ravindranath · 2022 [cited by applicant]
US 11544113B2 · Wang et al. · 2023 [cited by applicant]
US 11568238B2 · Huang et al. · 2023 [cited by applicant]
US 11716257B1 · Khermosh et al. · 2023 [cited by applicant]
US 11797535B1 · Stefani et al. · 2023 [cited by applicant]
US 11836520B2 · Yu · 2023 [cited by examiner]
US 20170068889A1 · Fougner et al. · 2017 [cited by applicant]
US 20180035240A1 · Ye et al. · 2018 [cited by applicant]
US 20180373976A1 · Woo · 2018 [cited by applicant]
US 20190130273A1 · Keskar et al. · 2019 [cited by applicant]
US 20190370241A1 · Miraldo · 2019 [cited by examiner]
US 20200226453A1 · Luk et al. · 2020 [cited by applicant]
US 20200311341A1 · Chaturvedi et al. · 2020 [cited by applicant]
US 20200311613A1 · Ma et al. · 2020 [cited by applicant]
US 20200410337A1 · Huang et al. · 2020 [cited by applicant]
US 20210012199A1 · Zhang et al. · 2021 [cited by applicant]
US 20210034335A1 · Svyatkovskiy et al. · 2021 [cited by applicant]
US 20210109796A1 · Fozard et al. · 2021 [cited by applicant]
US 20210125033A1 · Zhou · 2021 [cited by applicant]
US 20210141798A1 · Steedman · 2021 [cited by applicant]
US 20210192314A1 · Aarts et al. · 2021 [cited by applicant]
US 20210232773A1 · Wang et al. · 2021 [cited by applicant]
US 20210263779A1 · Haghighat et al. · 2021 [cited by applicant]
US 20210279576A1 · Shazeer et al. · 2021 [cited by applicant]
US 20210357210A1 · Clement et al. · 2021 [cited by applicant]
US 20210373944A1 · Lee et al. · 2021 [cited by applicant]
US 20210397610A1 · Singh et al. · 2021 [cited by applicant]
US 20210406673A1 · Pardeshi et al. · 2021 [cited by applicant]
US 20220066747A1 · Drain et al. · 2022 [cited by applicant]
US 20220066914A1 · Drain et al. · 2022 [cited by applicant]
US 20220067513A1 · Stevens · 2022 [cited by examiner]
US 20220101113A1 · Tam · 2022 [cited by examiner]
US 20220108212A1 · Zhai et al. · 2022 [cited by applicant]
US 20220147838A1 · Gu et al. · 2022 [cited by applicant]
US 20220164626A1 · Bird et al. · 2022 [cited by applicant]
US 20220237368A1 · Tran · 2022 [cited by applicant]
US 20220318601A1 · Yan · 2022 [cited by examiner]
US 20230127306A1 · Emelyanenko · 2023 [cited by examiner]
US 20230153381A1 · Wang et al. · 2023 [cited by applicant]
US 20230168922A1 · Xiaoyan · 2023 [cited by examiner]
CN 106503791A · 2017 [cited by applicant]
CN 108885571A · 2018 [cited by applicant]
CN 110326253A · 2019 [cited by applicant]
CN 110574049A · 2019 [cited by applicant]
CN 111898698A · 2020 [cited by applicant]
CN 112508018A · 2021 [cited by applicant]
CN 113222775A · 2021 [cited by applicant]
CN 113569868A · 2021 [cited by applicant]
CN 114127746A · 2022 [cited by applicant]
KR 1020210145490A · 2021 [cited by applicant]
KR 1020210148586A · 2021 [cited by applicant]
WO WO2021159201A1 · 2021 [cited by applicant]
Gao et al., “Low Latency RNN Inference with Cellular Batching,” EuroSys '18, pp. 1-15. (Year: 2018). [cited by examiner]
Katharopoulos et al., “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention,” arXiv, Aug. 31, 2020, pp. 1-17 (Year: 2020). [cited by examiner]
Lin et al., “A Survey of Transformers”, Jun. 15, 2021, arXiv, pp. 1-40. (Year: 2021). [cited by examiner]
China National Intellectual Property Administration, Office Action with English Translation, Chinese Patent Application No. 202210994036.0, Feb. 7, 2024, 14 pages. [cited by applicant]
China National Intellectual Property Administration, Office Action with English Translation, Chinese Patent Application No. 202210994014.4, Jan. 11, 2024, 13 pages. [cited by applicant]
European Patent Office, European Examination and Communication, European Patent Application No. 22185147.0, Mar. 4, 2024, 12 pages. [cited by applicant]
Faulbrück, L. et al. “Generating Synthetic Comments to Balance Data for Text Classification” Seminar Information Systems, Feb. 7, 2020, 81 pages, Retrieved from the internet <URL: https://humboldt-wi.github.io/blog/rese… [cited by applicant]
Bytedance Inc., “Effective Transformer,” last edited Aug. 8, 2020, 7 pages, Retrieved from the Internet <URL:https://github.com/bytedance/effective_transformer>. [cited by applicant]
Bytedance Inc., “Running BERT without Padding,” last edited Aug. 8, 2020, 7 pages, Retrieved from the Internet <URL:https://github.com/bytedance/effective_transformer#running-bert-without-padding>. [cited by applicant]
Choi, Y. et al. “Prema: A Predictive Multi-task Scheduling Algorithm for Preemptible Neural Processing Units,” IEEE International Symposium on High Performance Computer Architecture, Feb. 22-26, 2020, pp. 220-233. [cited by applicant]
Dai, Z. et al. “Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context,” Carnegie Mellon University, Jun. 2, 2019, pp. 1-20. [cited by applicant]
Doshi, K. “Transformers Explained Visually (Part 1): Overview of Functionality”, Dec. 13, 2020, Retrieved from the internet, <URL:towardsdatascience.com>. [cited by applicant]
Doshi, K. “Transformers Explained Visually (Part 2): How it works, step-by-step”, Jan. 2, 2021, Retrieved from the internet, <URL:towardsdatascience.com>. [cited by applicant]
Doshi, K. “Transformers Explained Visually (Part 3): Multi-head Attention deep dive”, Jan. 16, 2021, Retrieved from the internet, <URL:towardsdatascience.com>. [cited by applicant]
Doshi, K. “Transformers Explained Visually (Part 4): Not Just How, but Why They Work So Well”, Jun. 2, 2021, <URL:towardsdatascience.com>. [cited by applicant]
European Patent Office, Extended European Search Report and Written Opinion, European Patent Application No. 22185149.6, Dec. 21, 2022, 10 pages. [cited by applicant]
Fang, J. et al., “Turbo Transformers: An Efficient GPU Serving System for Transformer Models,” arXiv:2010.05680v4, Feb. 20, 2021, pp. 1-14. [cited by applicant]
Gao, P. et al., “Low Latency RNN Inference with Cellular Batching,” EuroSys '18, Apr. 2018, pp. 1-15. [cited by applicant]
Github, “microsoft/DeepSpeed,” Jan. 19, 2021, pp. 1-9, Retrieved from the Internet <URL:https://github.com/microsoft/DeepSpeed>. [cited by applicant]
Github, “NVIDIA/FasterTransformer,” Apr. 2, 2021, pp. 1-28, Retrieved from the Internet <URL:https://github.com/NVIDIA/FasterTransformer>. [cited by applicant]
Github, “NVIDIA/Megatron-LM,” Aug. 11, 2021, pp. 1-18, Retrieved from the Internet <URL:https://github.com/NVIDIA/Megatron-LM>. [cited by applicant]
Gschwind, M. et al., “A Better Transformer for Fast Transformer Inference,” PyTorch, Jul. 12, 2022, pp. 1-4, Retrieved from the Internet <URL:https://pytorch.org/blog/a-better-transformer-for-fast-transformer-encoder-in… [cited by applicant]
Gschwind, M., “Fast Transformer Inference With Better Transformer,” PyTorch, Jul. 19, 2022, pp. 1-4, Retrieved from the Internet <URL:https://pytorch.org/tutorials/beginner/bettertransformer_tutorial.html>. [cited by applicant]
Guo, H. et al. “ATT: A Fault-Tolerant ReRam Accelerator for Attention-Based Neural Networks,” IEEE 38th International Conference on Computer Design, Oct. 18-21, 2020, pp. 213-221. [cited by applicant]
Li, G. et al., “Easy and Efficient Transformer: Scalable Inference Solution for Large NLP Model,” arXiv:2104.12470v4, Nov. 23, 2021, pp. 1-9. [cited by applicant]
NVIDIA Corporation, “Faster Transformer BERT,” Sep. 30, 2021, 53 pages, Retrieved from the Internet <URL:https://github.com/NVIDIA/FasterTransformer/blob/main/docs/bert_guide.md#standard-bert-and-effective-fastertransfo… [cited by applicant]
NVIDIA Corporation, “Faster Transformer,” Apr. 2, 2021, 15 pages, Retrieved from the Internet <URL:https://github.com/nvidia/FasterTransformer>. [cited by applicant]
NVIDIA, “NVIDIA TensorRT,” Jan. 27, 2021, pp. 1-11, Retrieved from the Wayback Machine <URL:http://web.archive.org/web/20210127111124/https://developer.nvidia.com/tensorrt>. [cited by applicant]
NVIDIA, “NVIDIA Triton Inference Server,” Jan. 25, 2021, pp. 1-6, Retrieved from the Wayback Machine <URL: http://web.archive.org/web/20210125141031/https://developer.nvidia.com/nvidia-triton-inference-server>. [cited by applicant]
Olston, C. et al., “TensorFlow-Serving: Flexible, High-Performance ML Serving,” arXiv:1712.06139v2, Dec. 27, 2017, pp. 1-8. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/IB2022/000664, Mar. 30, 2023, 10 pages. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/IB2022/000666, Apr. 10, 2023, 10 pages. [cited by applicant]
Pytorch, “pytorch,” Dec. 2019, 14 pages, Retrieved from the Internet <URL:https://github.com/pytorch/pytorch/>. [cited by applicant]
Shazeer, N. et al., “Mesh-TensorFlow: Deep Learning for Supercomputers,” arXiv:1811.02084v1, Nov. 5, 2018, pp. 1-16. [cited by applicant]
Shoeybi, M. et al., “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism,” arXiv:1909.08053v4, Mar. 13, 2020, pp. 1-15. [cited by applicant]
Soojeong, K. et al. “Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks,” Fourteenth EuroSys Conference, Mar. 25, 2019, pp. 1-15. [cited by applicant]
Vaswani, A. et al. “Attention Is All You Need,” Advances in Neural Information Processing Systems, No. 30, Dec. 6, 2017, pp. 1-15. [cited by applicant]
Wang, X. et al., “LightSeq: A High-Performance Inference Library for Transformers,” arXiv:2010.13887v4, Apr. 22, 2021, pp. 1-8. [cited by applicant]
Yan, Y. et al. “El-Attention: Memory Efficient Lossless Attention for Generation,” International Conference on Machine Learning, Jul. 2021, pp. 1-11. [cited by applicant]
Zhou, H.Y. et al. “nnFormer: Interleaved Transformer for Volumetric Segmentation,” arXiv preprint arXiv:2109.03201, Sep. 7, 2021, pp. 1-10. [cited by applicant]
United States Office Action, U.S. Appl. No. 17/881,549, Feb. 14, 2023, 18 pages. [cited by applicant]
Choi, Y. et al. “Lazy Batching: An SLA-Aware Batching System for Cloud Machine Learning Inference,” IEEE International Symposium on High-Performance Computer Architecture, Feb. 27, 2021, 14 pages. [cited by applicant]
Exhibit A1—Initial Invalidity Contentions for U.S. Pat. No. 11,442,775 over U.S. Pat. No. 11,797,535, Titled “Use of Batch Mode Function Execution in Database Engines to Enable Efficient Calls to Remote Services” by Ste… [cited by applicant]
Exhibit A2—Initial Invalidity Contentions for U.S. Pat. No. 11,442,775 over U.S. Pat. No. 11,797,535, Titled “Use of Batch Mode Function Execution in Database Engines to Enable Efficient Calls to Remote Services” by Ste… [cited by applicant]
Exhibit A3—Initial Invalidity Contentions for U.S. Pat. No. 11,442,775 over U.S. Pat. No. 11,797,535, Titled “Use of Batch Mode Function Execution in Database Engines to Enable Efficient Calls to Remote Services” by Ste… [cited by applicant]
Exhibit A4—Initial Invalidity Contentions for U.S. Pat. No. 11,442,775 over U.S. Publication No. 2021/0192314, Titled “API for Recurrent Neural Networks” by Bastiaan Joannes Matheus Aarts et al., Filed Mar. 15, 2024, 76… [cited by applicant]
Exhibit A5—Initial Invalidity Contentions for U.S. Pat. No. 11,442,775 over U.S. Publication No. 2021/0192314, Titled “API for Recurrent Neural Networks” by Bastiaan Joannes Matheus Aarts et al., U.S. Publication No. 20… [cited by applicant]
Exhibit A-SR—Initial Invalidity Contentions for U.S. Pat. No. 11,442,775—Disclosure of Secondary References, Filed Mar. 15, 2024, 2 pages. [cited by applicant]
Exhibit B1—Initial Invalidity Contentions for U.S. Pat. No. 11,836,520 over U.S. Pat. No. 11,797,535, Titled “Use of Batch Mode Function Execution in Database Engines to Enable Efficient Calls to Remote Services” by Ste… [cited by applicant]
Exhibit B2—Initial Invalidity Contentions for U.S. Pat. No. 11,836,520 over U.S. Pat. No. 11,797,535, Titled “Use of Batch Mode Function Execution in Database Engines to Enable Efficient Calls to Remote Services” by Ste… [cited by applicant]
Exhibit B3—Initial Invalidity Contentions for U.S. Pat. No. 11,836,520 over U.S. Pat. No. 11,797,535, Titled “Use of Batch Mode Function Execution in Database Engines to Enable Efficient Calls to Remote Services” by Ste… [cited by applicant]
Exhibit B4—Initial Invalidity Contentions for U.S. Pat. No. 11,836,520 over U.S. Publication No. 2021/0192314, Titled “API for Recurrent Neural Networks” by Bastiaan Joannes Matheus Aarts et al., Filed Mar. 15, 2024, 23… [cited by applicant]
Exhibit B5—Initial Invalidity Contentions for U.S. Pat. No. 11,836,520 over U.S. Publication No. 2021/0192314, Titled “API for Recurrent Neural Networks” by Bastiaan Joannes Matheus Aarts et al., U.S. Publication No. 20… [cited by applicant]
Exhibit B-SR—Initial Invalidity Contentions for U.S. Pat. No. 11,836,520—Disclosure of Secondary References, Filed Mar. 15, 2024, 2 pages. [cited by applicant]
Exhibit C—Initial Invalidity Contentions for U.S. Pat. No. 11,442,775 and U.S. Pat. No. 11,836,520—Disclosure of Invalidity Pursuant to 35 U.S.C. § 112, Filed Mar. 15, 2024, 2 pages. [cited by applicant]
United States District Court, District of Delaware, Defendant Hugging Face, Inc.'s Initial Invalidity Contentions, [cited by applicant]
Wang, Y. et al. “Energy-Efficient Inference Service of Transformer-Based Deep Learning Models on GPUs,” 2020 International Conferences on Internet of Things, IEEE Green Computing and Communications, IEEE Cyber, Physical… [cited by applicant]
Yan, Y. et al. “Fastseq: Make Sequence Generation Faster,” arXiv preprint arXiv:2106.04718, Jun. 8, 2021, 9 pages. [cited by applicant]
Benesty, M. “Divide Hugging Face Transformers Training Time by 2 or More with Dynamic Padding and Uniform Length Batching,” Medium.com, 26 pages, Retrieved from the internet <URL:https://towardsdatascience.com/divide-hu… [cited by applicant]
He, P. et al. “DeBERTa: Decoding-enhanced BERT with disentangled attention,” arXiv preprint arXiv:2006.03654, 2021 International Conference on Learning Representations, Oct. 6, 2021, 23 pages. [cited by applicant]
Kitaev, N. et al. “Reformer: The efficient transformer,” arXiv preprint arXiv:2001.04451, 2020 International Conference on Learning Representations, Feb. 18, 2020, 12 pages. [cited by applicant]
Kosec, M. et al. “Packing: Towards 2x NLP BERT Acceleration,” 2021, 12 pages. [cited by applicant]
Liu, Y. et al. “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692, Jul. 26, 2019, 13 pages. [cited by applicant]
Muthuvelu, N. et al. “A dynamic job grouping-based scheduling for deploying applications with fine-grained tasks on global grids,” 2005 Australasian Workshop on Grid Computing and E-Research, vol. 44, Jan. 1, 2005, 8 pa… [cited by applicant]
United States District Court, District of Delaware, Defendant Hugging Face, Inc.'s Supplemental Identification of Invalidity References, [cited by applicant]
Wang, S. et al. “An efficient and non-intrusive GPU scheduling framework for deep learning training systems,” SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, Nov. 9, 2020… [cited by applicant]
Chinese Intellectual Property Office, Office Action, Chinese Patent Application No. 202411209059.1, May 22, 2025, 14 pages. [cited by applicant]