IP Library Granted Patent US 11,815,984
Granted Patent B2
US 11,815,984 · App. 16/784,768 · Granted Nov 14, 2023

Error handling in an interconnect

Inventors: Tina C. Toupal (Portland, OR); Shamsul Abedin (Portland, OR)
Assignee: Intel Corporation
G06F11/0757G06F9/4812G06F13/4068G06F15/7807H04B10/801H04L12/5601
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,815,984
App. No.
16/784,768
Granted
Nov 14, 2023
Kind
B2
Abstract

A system level error detection and handling of the network IO in a multi-chip-package (MCP) die is provided. The error detection and handling mechanism conceived may be used between a system-on-chip (SoC) die and a different type of die, such as a die manufactured by a third-party (e.g., a high-bandwidth network IO die). To provide a timely indication in case of any part of the network is at fault, a control unit on the SoC die handles error detection on the network IO links using various indicators. After errors are detected, the control unit groups the errors into two categories: a link failure and a virtual channel failure. Such an error handling mechanism may consolidate the actions and provide consistency in hardware behavior.

Claims (38)

1. A tangible, non-transitory, computer-readable medium having instructions stored thereon, wherein the instructions are configured to cause a processor to:

receive an indication of an error for a link between a plurality of dies of a multi-die system, wherein the link is to enable network communications between the plurality of dies;

determine whether the error is a network error associated with the network communications;

when the error is not a network error, logically switch the link off; and

send a not-acknowledged message in response to incoming packets with an error code indicating a type of link error.

2. The tangible, non-transitory, computer-readable medium of claim 1 , wherein the network error occurs during operation of the multi-die system and is not related to an initial boot or restart of the multi-die system, and wherein the processor is configured to invoke an interrupt handler to interrupt the link in response to the network error.

3. The tangible, non-transitory, computer-readable medium of claim 1 , wherein logically switching the link off comprises clearing in-flight packets designated to be transferred over the link.

4. The tangible, non-transitory, computer-readable medium of claim 1 , wherein the link comprises an advanced interface bus link.

5. The tangible, non-transitory, computer-readable medium of claim 4 , wherein the error comprises an advanced interface bus error.

6. The tangible, non-transitory, computer-readable medium of claim 5 , wherein the processor is configured to start a timer after an initial boot or restart of the multi-die system, and wherein the indication of the error comprises a duration of the timer elapsing without receiving a transmitter enable bit or a receiver enable bit over the advanced interface bus link.

7. The tangible, non-transitory, computer-readable medium of claim 1 , wherein the link comprises a silicon photonics link.

8. The tangible, non-transitory, computer-readable medium of claim 7 , wherein the error comprises an optical link training timeout error of the silicon photonics link.

9. The tangible, non-transitory, computer-readable medium of claim 8 , wherein the optical link training timeout error of the silicon photonics link comprises a lack of receiving a clock and data recovery lock status bit associated with the silicon photonics link within a threshold of time after a boot of the multi-die system.

10. The tangible, non-transitory, computer-readable medium of claim 8 , wherein the optical link training timeout error of the silicon photonics link comprises a lack of receiving a phase-locked loop lock status bit within a threshold of time after a boot of the multi-die system.

11. The tangible, non-transitory, computer-readable medium of claim 1 , wherein the error comprises a channel alignment error that occurs when alignment training is not complete within a specific time of a boot of the multi-die system.

12. The tangible, non-transitory, computer-readable medium of claim 1 , wherein the link comprises an optical link, and the error comprises an optical link failure corresponding to a threshold number of virtual channels of the optical link have timed out.

13. The tangible, non-transitory, computer-readable medium of claim 12 , wherein the threshold number comprises all of the virtual channels of the optical link.

14. A method, comprising:

determining that a timeout for a virtual channel of a link between a plurality of dies of a multi-die system;

determining that packets have been replayed over the virtual channel more than a threshold number of attempts and continues to timeout; and

in response to the packets being replayed more than the threshold number of attempts, disabling the virtual channel.

15. The method of claim 14 , wherein the threshold number is a specified number of replays of the packets, ranging from 4 to 8.

16. The method of claim 14 , comprising:

determining that a threshold number of virtual channels of the link have failed; and

in response to a determination that more than the threshold number of the virtual channels of the link have failed, disable the link.

17. The method of claim 14 , comprising:

determining that fewer than a threshold number of virtual channels of the link have failed; and

continue using the link while disabling the virtual channel.

18. The method of claim 14 , wherein disabling the virtual channel comprises responding to incoming requests to transfer data over the virtual channel with non-acknowledged message indicating that the virtual channel has failed using a corresponding error code.

19. The method of claim 14 , wherein disabling the virtual channel comprises clearing in-flight packets that are designated to be transferred over the virtual channel.

20. A method, comprising:

determining that a timeout for a virtual channel of a link between a plurality of dies of a multi-die system;

determining that packets have been replayed over the virtual channel more than a threshold number of attempts and continues to timeout;

in response to the packets being replayed more than the threshold number of attempts, disable the virtual channel;

responding to requests on the disabled virtual channel with a first not-acknowledged response indicating that the virtual channel has failed;

determining that more virtual channels of the link are disabled than a link threshold;

responsive to a determination that more virtual channels of the link are disabled than the link threshold, disabling the link; and

responding to requests on the disabled link with a first not-acknowledged response indicating that the link has failed.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2026
From: INTEL CORPORATION
To: INTEL PRODUCTS IP LLC
Reel/Frame 076025/0681 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2020
From: TOUPAL, TINA C.; ABEDIN, SHAMSUL
To: INTEL CORPORATION
Reel/Frame 051774/0022 →
Continuity (1)
Related Publication 20200174873A1 · Jun 4, 2020