IP Library › Granted Patent US 12,189,469
Granted Patent B2
US 12,189,469 · App. 18/298,470 · Granted Jan 7, 2025

PCIe device error handling system

Inventors: Wei Liu (Austin, TX); Tuyet-Huong Th Nguyen (Cedar Park, TX)
Assignee: Dell Products L.P.
G06F11/0793G06F11/0745G06F11/0787
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,469
App. No.
18/298,470
Granted
Jan 7, 2025
Kind
B2
Abstract

A PCIe device error handling system includes a BIOS subsystem coupled to a PCIe device and a BMC device. The BIOS subsystem identifies an error in the PCIe device and, in response, begins an SMM that suspends the performance of at least one workload in an operating system, and generates and transmits a PCIe device error information collection instruction associated with the PCIe device to the BMC device. Subsequent to transmitting the PCIe device error information collection instruction, the BIOS subsystem ends the SMM such that the performance of at least one workload is resumed in the operating system. In response to receiving the PCIe device error information collection instruction from the BIOS subsystem, the BMC device retrieves PCIe device error information from the PCIe device while the operating system performs the at least one workload, and stores the PCIe device error information.

Claims (67)

1. A Peripheral Component Interconnect express (PCIe) device error handling system, comprising:

an operating system;

a Peripheral Component Interconnect express (PCIe) device;

a Basic Input/Output System (BIOS) subsystem that is coupled to the PCIe device and that is configured to:

identify an error in the PCIe device and, in response, begin a System Management Mode (SMM) that suspends the performance of at least one workload in the operating system;

generate and transmit a PCIe device error information collection instruction associated with the PCIe device; and

end, subsequent to transmitting the PCIe device error information collection instruction, the SMM such that the performance of at least one workload is resumed in the operating system; and

a Baseboard Management Controller (BMC) device that is coupled to the BIOS and that is configured to:

receive, from the BIOS subsystem, the PCIe device error information collection instruction and, in response, retrieve PCIe device error information from the PCIe device while the operating system performs the at least one workload; and

store the PCIe device error information.

2. The system of claim 1 , wherein the PCIe device error information retrieved by the BMC device includes a first subset of PCIe device error information that is included in at least one register in the PCIe device, and wherein the BIOS subsystem is configured to:

retrieve, in response to identifying the error in the PCIe device, a second subset of the PCIe device error information that is included in at least one register in the PCIe device and that is different than the first subset of the PCIe device error information; and

store the second subset of the PCIe device error information.

3. The system of claim 1 , wherein the BMC device is configured to:

identify, in response to receiving the PCIe device error information collection instruction, at least one port that is connected to the PCIe device; and

store a port identifier for each of the at least one port.

4. The system of claim 1 , wherein the BMC device is configured to:

identify, in response to receiving the PCIe device error information collection instruction, at least one bridge device that is connected to the PCIe device; and

store a bridge device identifier for each of the at least one bridge device.

5. The system of claim 1 , wherein the PCIe device is a Non-Volatile Memory express (NVMe) storage device and the PCIe device error information is retrieved from a queue included in the NVMe storage device.

6. The system of claim 1 , wherein the BMC device is configured to:

perform error causation analysis operations using the PCIe device error information to determine at least one cause of the error identified in the PCIe device; and

provide the at least one cause of the error identified in the PCIe device for display on a display device.

7. An Information Handling System (IHS), comprising:

a Basic Input/Output System (BIOS) processing system;

a BIOS memory system that is coupled to the BIOS processing system and that includes instructions that, when executed by the BIOS processing system, cause the BIOS processing system to provide a BIOS engine that is configured to:

identify an error in a Peripheral Component Interconnect express (PCIe) device that is coupled to the BIOS processing system and, in response, begin a System Management Mode (SMM) that suspends the performance of at least one workload in an operating system;

generate and transmit a PCIe device error information collection instruction associated with the PCIe device; and

end, subsequent to transmitting the PCIe device error information collection instruction, the SMM such that the performance of at least one workload is resumed in the operating system;

a Baseboard Management Controller (BMC) processing system that is coupled to the BIOS processing system and the PCIe device; and

a BMC memory system that is coupled to the BMC processing system and that includes instructions that, when executed by the BMC processing system, cause the BMC processing system to provide a BMC engine that is configured to:

receive, from the BIOS engine, the PCIe device error information collection instruction and, in response, retrieve PCIe device error information from the PCIe device while the operating system performs the at least one workload; and

store the PCIe device error information.

8. The IHS of claim 7 , wherein the PCIe device error information retrieved by the BMC engine includes a first subset of PCIe device error information that is included in at least one register in the PCIe device, and wherein the BIOS engine is configured to:

retrieve, in response to identifying the error in the PCIe device, a second subset of the PCIe device error information that is included in at least one register in the PCIe device and that is different than the first subset of the PCIe device error information; and

store the second subset of the PCIe device error information.

9. The IHS of claim 8 , wherein the first subset of the PCIe device error information includes Base Address Register (BAR) information, command (CMD) register information, and status register information.

10. The IHS of claim 7 , wherein the BMC engine is configured to:

identify, in response to receiving the PCIe device error information collection instruction, at least one port that is connected to the PCIe device; and

store a port identifier for each of the at least one port.

11. The IHS of claim 7 , wherein the BMC engine is configured to:

identify, in response to receiving the PCIe device error information collection instruction, at least one bridge device that is connected to the PCIe device; and

store a bridge device identifier for each of the at least one bridge device.

12. The IHS of claim 7 , wherein the PCIe device is a Non-Volatile Memory express (NVMe) storage device and the PCIe device error information is retrieved from a queue included in the NVMe storage device.

13. The IHS of claim 7 , wherein the BMC engine is configured to:

perform error causation analysis operations using the PCIe device error information to determine at least one cause of the error identified in the PCIe device; and

provide the at least one cause of the error identified in the PCIe device for display on a display device.

14. A method for handling errors in a Peripheral Component Interconnect express (PCIe) device, comprising:

identifying, by a Basic Input/Output System (BIOS) subsystem, an error in a Peripheral Component Interconnect express (PCIe) device that is coupled to the BIOS subsystem and, in response, beginning a System Management Mode (SMM) that suspends the performance of at least one workload in an operating system;

generating and transmitting, by the BIOS subsystem, a PCIe device error information collection instruction associated with the PCIe device;

ending, by the BIOS subsystem subsequent to transmitting the PCIe device error information collection instruction, the SMM such that the performance of at least one workload is resumed in the operating system;

receiving, by a Baseboard Management Controller (BMC) device from the BIOS subsystem, the PCIe device error information collection instruction and, in response, retrieving PCIe device error information from the PCIe device while the operating system performs the at least one workload; and

storing, by the BMC device, the PCIe device error information.

15. The method of claim 14 , wherein the PCIe device error information retrieved by the BMC subsystem includes a first subset of PCIe device error information that is included in at least one register in the PCIe device, and wherein the method includes:

retrieving, by the BIOS subsystem in response to identifying the error in the PCIe device, a second subset of the PCIe device error information that is included in at least one register in the PCIe device and that is different than the first subset of the PCIe device error information; and

storing, by the BIOS subsystem, the second subset of the PCIe device error information.

16. The method of claim 15 , wherein the first subset of the PCIe device error information includes Base Address Register (BAR) information, command (CMD) register information, and status register information.

17. The method of claim 14 , further comprising:

identifying, by the BMC device in response to receiving the PCIe device error information collection instruction, at least one port that is connected to the PCIe device; and

storing, by the BMC device, a port identifier for each of the at least one port.

18. The method of claim 14 , further comprising:

identifying, by the BMC device in response to receiving the PCIe device error information collection instruction, at least one bridge device that is connected to the PCIe device; and

storing, by the BMC device, a bridge device identifier for each of the at least one bridge device.

19. The method of claim 14 , wherein the PCIe device is a Non-Volatile Memory express (NVMe) storage device and the PCIe device error information is retrieved from a queue included in the NVMe storage device.

20. The method of claim 14 , further comprising:

performing, by the BMC device, error causation analysis operations using the PCIe device error information to determine at least one cause of the error identified in the PCIe device; and

providing, by the BMC device, the at least one cause of the error identified in the PCIe device for display on a display device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2023
From: LIU, WEI; NGUYEN, TUYET-HUONG THI
To: DELL PRODUCTS L.P.
Reel/Frame 063287/0202 →
Continuity (1)
Related Publication 20240345915A1 · Oct 17, 2024
References Cited (15)
US 9954727B2 · Su et al. · 2018 [cited by applicant]
US 11366710B1 · Hung · 2022 [cited by examiner]
US 20040003313A1 · Ramirez · 2004 [cited by examiner]
US 20080126650A1 · Swanson · 2008 [cited by examiner]
US 20190278651A1 · Thornley · 2019 [cited by examiner]
US 20200151048A1 · Shantamurthy · 2020 [cited by examiner]
US 20200285534A1 · Chaiken · 2020 [cited by examiner]
US 20210263868A1 · Maddukuri · 2021 [cited by examiner]
US 20220318087A1 · Hong · 2022 [cited by examiner]
US 20220374320A1 · Song · 2022 [cited by examiner]
US 20230195568A1 · Hong · 2023 [cited by examiner]
US 20240004757A1 · Lee · 2024 [cited by examiner]
US 20240037283A1 · Chen · 2024 [cited by examiner]
US 20240103961A1 · Wang · 2024 [cited by examiner]
US 20240176739A1 · Alden · 2024 [cited by examiner]