IP Library › Granted Patent US 12,609,715
Granted Patent B2
US 12,609,715 · App. 18/670,615 · Granted Apr 21, 2026

Performing cyclic redundancy checks using parallel computing architectures

Inventor: Andrea Miele (San Jose, CA)
Assignee: NVIDIA Corporation
H03M13/09G06F9/5005G06T1/20H04L1/0061
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,609,715
App. No.
18/670,615
Granted
Apr 21, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to compute cyclic redundancy checks use a graphics processing unit (GPU) to compute cyclic redundancy checks. For example, in at least one embodiment, an input data sequence is distributed among GPU threads for parallel calculation of an overall CRC value for the input data sequence according to various novel techniques described herein.

Claims (28)

1 . A system, comprising:

one or more processors; and

memory including instructions executable by the one or more processors that cause the system to at least:

allocate two or more portions of one or more transmitted data sequences and one or more cyclic redundancy sequences to two or more threads of a graphics processor; and

use the two or more threads of the graphics processor to perform error detection of the one or more transmitted data sequences at least partially in parallel with the one or more cyclic redundancy sequences to generate one or more cyclic redundancy check values for the one or more transmitted data sequences.

2 . The system of claim 1 , wherein the one or more transmitted data sequences comprise one or more transport blocks.

3 . The system of claim 1 , wherein one or more of the cyclic redundancy check values correspond to a combination of transport blocks.

4 . The system of claim 1 , wherein the one or more transmitted data sequences comprise one or more data segments that are to be allocated to the two or more threads of the graphics processor.

5 . The system of claim 1 , wherein the one or more cyclic redundancy sequences comprise one or more generator segments.

6 . The system of claim 5 , wherein the one or more generator segments are to be combined with one or more data segments of the one or more transmitted data sequences.

7 . The system of claim 1 , wherein the error detection is to be performed using a parallel reduction tree to combine with one or more data segments of the one or more transmitted data sequences.

8 . One or more processors, comprising circuitry to:

allocate two or more portions of one or more transmitted data sequences and one or more cyclic redundancy sequences to two or more threads of a graphics processor; and

use the two or more threads of the graphics processor to perform error detection of the one or more transmitted data sequences at least partially in parallel with the one or more cyclic redundancy sequences to generate one or more cyclic redundancy check values for the one or more transmitted data sequences.

9 . The one or more processors of claim 8 , wherein the two or more portions comprise data segments that are to be combined with one or more generator segments of the one or more cyclic redundancy sequences.

10 . The one or more processors of claim 8 , wherein the one or more transmitted data sequences comprise one or more transport blocks.

11 . The one or more processors of claim 8 , wherein the one or more cyclic redundancy sequences comprise generator segments that are to be combined with one or more data segments of the one or more transmitted data sequences.

12 . The one or more processors of claim 8 , wherein the two or more portions comprises data segments having a length of two or more bytes.

13 . The one or more processors of claim 8 , wherein the one or more cyclic redundancy sequences comprise one or more generator segments.

14 . The one or more processors of claim 8 , wherein the one or more transmitted data sequences comprise a binary polynomial, and the one or more cyclic redundancy sequences comprise a generator polynomial.

15 . A method, comprising:

allocating two or more portions of one or more transmitted data sequences and one or more cyclic redundancy sequences to two or more threads of a graphics processor; and

using the two or more threads of the graphics processor to perform error detection of the one or more transmitted data sequences at least partially in parallel with the one or more cyclic redundancy sequences to generate one or more cyclic redundancy check values for the one or more transmitted data sequences.

16 . The method of claim 15 , further comprising correcting one or more errors in the one or more transmitted data sequences.

17 . The method of claim 15 , wherein using the two or more threads of the graphics processor to perform the error detection comprises combining one or more data segments of the one or more transmitted data sequences with one or more generator segments of the one or more cyclic redundancy sequences.

18 . The method of claim 15 , wherein using the two or more threads of the graphics processor to perform the error detection comprises combining the two or more portions with the one or more cyclic redundancy sequences.

19 . The method of claim 15 , wherein the one or more transmitted data sequences comprise one or more transport blocks.

20 . The method of claim 15 , wherein the one or more cyclic redundancy sequences comprise generator segments, and the two or more portions comprise data segments.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2024
From: MIELE, ANDREA
To: NVIDIA CORPORATION
Reel/Frame 067484/0772 →
Continuity (3)
Continuation 17402341 · Aug 13, 2021
Continuation 16559424 · Sep 3, 2019
Related Publication 20250132771A1 · Apr 24, 2025
References Cited (51)
US 4865423A · Doi · 1989 [cited by applicant]
US 5923680A · Brueckheimer et al. · 1999 [cited by applicant]
US 6560742B1 · Dubey et al. · 2003 [cited by applicant]
US 8587609B1 · Wang et al. · 2013 [cited by applicant]
US 9141469B2 · Radhakrishnan et al. · 2015 [cited by applicant]
US 11095307B2 · Miele · 2021 [cited by examiner]
US 11791871B2 · Banuli Nanje Gowda · 2023 [cited by applicant]
US 12003253B2 · Miele · 2024 [cited by examiner]
US 20050066256A1 · Hooper et al. · 2005 [cited by applicant]
US 20080002854A1 · Tehranchi et al. · 2008 [cited by applicant]
US 20080074433A1 · Jiao et al. · 2008 [cited by applicant]
US 20100125777A1 · Wang et al. · 2010 [cited by applicant]
US 20110231636A1 · Olson et al. · 2011 [cited by applicant]
US 20130007573A1 · Radhakrishnan et al. · 2013 [cited by applicant]
US 20130104012A1 · Barner · 2013 [cited by applicant]
US 20130135315A1 · Bares et al. · 2013 [cited by applicant]
US 20140207370A1 · Severson · 2014 [cited by applicant]
US 20170187389A1 · Radhakrishnan et al. · 2017 [cited by applicant]
US 20180067810A1 · Li · 2018 [cited by applicant]
US 20180307984A1 · Koker et al. · 2018 [cited by applicant]
US 20180356235A1 · Jang · 2018 [cited by applicant]
US 20190007495A1 · Zhao et al. · 2019 [cited by applicant]
CN 201153259Y · 2008 [cited by applicant]
CN 101847999A · 2010 [cited by applicant]
CN 107451008A · 2017 [cited by applicant]
CN 109196801A · 2019 [cited by applicant]
JP 2010068429A · 2010 [cited by applicant]
WO 2018175711A1 · 2018 [cited by applicant]
WO 2018230701A1 · 2018 [cited by applicant]
Blem et al., “Instruction Set Extensions for Cyclic Redundancy Check on a Multithreaded Processor,” 7th Workshop on Media and Stream Processors, Dec. 12, 2005, 8 pages. [cited by applicant]
Chi et al., “Exploring Various Levels of Parallelism in High-Performance CRC Algorithms,” IEEE Access 7 (6):32315-32326, Mar. 6, 2019. [cited by applicant]
Chi et al., “VACA: A High-Performance Variable-Length Adaptive CRC Algorithm,” 2017 IEEE 28th Annual International Symposium on Personal, Indoor, and Mobile Radio Communications (PIMRC), Oct. 8, 2017, 6 pages. [cited by applicant]
Cho et al., “Block-Interleaving Based Parallel CRC Computation for Multi-Processor Systems,” 2010 IEEE Workshop on Signal Processing System (SIPS 2010), Oct. 6, 2010, 6 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Electrotechnical Commission, “Functional safety of electrical/electronic/programmable electronic safety-related systems,” IEC Standard 61508-1, Apr. 2014, 23 pages. [cited by applicant]
International Organization for Standardization, “Road vehicles—Functional safety,” ISO Standard 26262, https://www.iso.org/obp/ui/#iso:std:iso:26262:-1:ed-1:v1:en, Nov. 11, 2011, 35 pages. [cited by applicant]
International Search Report and Written Opinion mailed Nov. 19, 2020, Patent Application No. PCT/US2020/048638, 21 pages. [cited by applicant]
Ji et al., “Fast Parallel CRC Algorithm and Implementation on a Configurable Processor,” 2002 IEEE International Conference on Communications, Apr. 28, 2002, 5 pages. [cited by applicant]
Kim et al., “A fast and energy-efficient Hamming decoder for software-defined radio using graphics processing units,” Journal of Supercomputing 71(7):2454-2472, Mar. 5, 2015. [cited by applicant]
Meng, “Analyzing General-Purpose Computing Performance on GPU,” California Polytechnic State University Master's Thesis, Dec. 2015, 79 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Sun et al., “A Table-Based Algorithm for Pipelined CRC Calculation,” 2010 IEEE International Conference on Communications, May 23, 2010, 5 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2202264.4, mailed Sep. 26, 2023, 2 pages. [cited by applicant]
Office Action for Chinese Application No. 202080072279.8, mailed Dec. 11, 2023, 18 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2202264.4, mailed Mar. 28, 2024, 7 pages. [cited by applicant]
Combined Search and Examination Report for United Kingdom Application No. GB2412448.9, mailed Sep. 19, 2024, 6 pages. [cited by applicant]
Zheng et al., “Architecting An LTE Base Station with Graphics Processing Units,” IEEE, Oct. 16, 2013, pp. 219-224. [cited by applicant]
Notice of Decision to Grant for Chinese Application No. 202080072279.8, mailed Sep. 27, 2024, 6 pages. [cited by applicant]
Notice of Intention to Grant for United Kingdom Application No. GB2202264.4, mailed Nov. 29, 2024, 2 pages. [cited by applicant]
Notice of Intention to Grant for United Kingdom Application No. GB2412448.9, mailed Dec. 9, 2024, 2 pages. [cited by applicant]