IP Library Granted Patent US 12,363,316
Granted Patent B2
US 12,363,316 · App. 17/555,121 · Granted Jul 15, 2025

Video codec importance indication and radio access network awareness configuration

Inventors: Razvan-Andrei Stoica (Essen, DE); Hossein Bagheri (Urbana, IL); Vijay Nangia (Woodridge, IL)
Assignee: Lenovo (Singapore) Pte. Ltd.
H04N19/164H04N19/117H04N19/136H04N19/188H04N19/189H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,363,316
App. No.
17/555,121
Granted
Jul 15, 2025
Kind
B2
Abstract

Apparatuses, methods, and systems are disclosed for video codec importance indication and RAN awareness configuration. An apparatus includes a processor that detects a plurality of video coded network abstraction layer (“NAL”) units of a video coded stream, extracts semantic information associated with the plurality of the NAL units, combines the extracted semantic information associated with the plurality of NAL units to form a plurality of feature sets that is correspondingly synchronized with the plurality of NAL units that enclose the extracted semantic information, determines an information-to-importance value for each of the plurality of NAL units based on the plurality of feature sets without performing video decoding of the video coded stream, and indicates the determined information-to-importance value for each of the plurality of NAL units of the determined video codec specification to a video coded traffic-aware transceiver for scheduling video traffic based on the indicated information-to-importance value.

Claims (74)

1. A network device apparatus, the apparatus comprising:

at least one memory; and

at least one processor coupled with the at least one memory and configured to cause the apparatus to:

detect a plurality of video coded network abstraction layer (“NAL”) units of a video coded stream according to defined syntax elements of a determined video codec specification;

extract semantic information associated with the plurality of NAL units according to a syntax elements and semantic knowledge base defined within the determined video codec specification;

combine the extracted semantic information associated with the plurality of NAL units to form a plurality of feature sets comprising a plurality of NAL unit type parameters, the plurality of feature sets correspondingly synchronized with the plurality of NAL units that enclose the extracted semantic information;

determine an information-to-importance value for each of the plurality of NAL units based on a NAL unit type, a NAL unit size, and a hierarchical encoding of each of the plurality of feature sets without performing video decoding of the video coded stream; and

indicate the determined information-to-importance value for each of the plurality of NAL units of the determined video codec specification to a video coded traffic-aware transceiver for scheduling video traffic based on the indicated information-to-importance value.

2. The apparatus of claim 1 , wherein the at least one processor is configured to cause the apparatus to parse the plurality of detected video coded NAL units as syntax elements of the determined video codec specification for extracting the semantic information.

3. The apparatus of claim 1 , wherein the at least one processor is configured to cause the apparatus to select a video codec syntax and semantic filter from a set of video codec syntax and semantic filter candidates based on the determined video codec specification and exposing the syntax elements and semantic knowledge base of the video codec specification.

4. The apparatus of claim 3 , wherein the at least one processor is configured to cause the apparatus to perform intrinsic semantic information extraction by applying the selected video codec syntax and semantic filter to at least one of a NAL unit's header and a NAL unit's non-entropy coded payload to extract the semantic information contained within the video coded stream.

5. The apparatus of claim 4 , wherein the intrinsic semantic information extraction is based on processing at least one NAL unit enclosed element selected from a group of:

video layer metadata information corresponding to one or more of a video layer identifier and a video layer priority information;

video layer frames-per-second playback metadata information;

video layer additional metadata for parsing layer-common video coded sequence information;

video sequence resolution corresponding to at least a combination of width and height of a video coded frame of the sequence of video coded frames;

video sequence chroma subsampling information;

video sequence coding metadata information corresponding to one or more of a coding block size, a number of coded rows, and a number of coded columns within the sequence of video coded frames;

video sequence number of maximally allowed reference frames for inter-frame motion prediction;

video sequence additional metadata necessary for parsing sequence-common video coded picture information;

video coded picture encoding mode configuration;

video coded picture segmentation information into picture partitions;

video coded picture additional metadata for parsing slice-common video coded slice information;

video coded slice type;

video coded slice payload size;

video coded slice position within one or more of a video coded picture, a tile, and a sub-picture;

video coded slice number of enclosed video coded rows;

video coded slice payload size of each enclosed video coded row; and

video coded slice reference list wherein the video coded reference list defines at least one reference slice for an inter-slice motion prediction and video encoding for a current slice enclosed within the NAL unit.

6. The apparatus of claim 5 , wherein the at least one processor is configured to cause the apparatus to extend the extracted intrinsic semantic information using an extrinsic source of semantic information synchronized with the plurality of NAL units of the video coded stream as provided by a video encoding function implementing the video codec specification that generated the plurality of NAL units.

7. The apparatus of claim 6 , wherein at least one of the extracted intrinsic semantic information and extrinsic semantic information of the determined video codec is at least one of combined, reduced, and post-processed into a plurality of normalized feature sets that are synchronized with the plurality of NAL units of the video coded stream.

8. The apparatus of claim 1 , wherein the determination of the information-to-importance value for each of the plurality of NAL units of the determined video codec specification comprises processing a hierarchical structure and encapsulation encoding of the video codec specification spanning over at least one hierarchy selected from a group of:

a video layer level;

a video sequence level;

a video coded frame level;

a video coded slice level; and

a video coded slice segment level.

9. The apparatus of claim 8 , wherein the at least one processor is configured to cause the apparatus to model a graphical directed information flow of the processing of a NAL unit's hierarchical structure and encapsulation encoding towards the determination of the information-to-importance value for each of the plurality of NAL units based at least on an information-to-importance functional kernel and a recursive accumulation of information flow across directed edges of the graphical directed information flow model formed of the plurality of NAL units as nodes.

10. The apparatus of claim 9 , wherein the information-to-importance functional kernel is defined for any level depth of the graphical directed information flow model processing as a functional parameter of a subset of an extracted feature set associated with a graph node i on the level as .

11. The apparatus of claim 10 , wherein the recursive accumulation of information flow across the directed edges of the graphical directed information flow model formed of NAL unit nodes is limited to a weighted aggregation of intra-level

directed edges and a direct lower level inter-level directed edges such that for a node i at level , wherein an accumulated importance value is obtained as

= +Σ j≠i +Σ k ,

wherein a weight of the aggregation between the node and a node j≠i at level is a scalar , and a weight of the aggregation between node and a node k at level

+1 is a scalar .

12. The apparatus of claim 11 , wherein the at least one processor is configured to cause the apparatus to limit the information-to-importance functional kernel

to a model-common information-to-importance functional kernel

.

13. The apparatus of claim 9 , wherein the information-to-importance functional kernel is a scalar constant for graphical directed information flow model nodes and levels.

14. The apparatus of claim 9 , wherein the at least one processor is configured to cause the apparatus to determine an importance value of the information-to-importance functional kernel for the video coded slice or slice segment level inversely proportional to a video compression gain relative to an uncompressed video picture resolution content under constant quality encodings such that frames that are compressed above a predefined threshold have low importance values and frames that are compressed below the predefined threshold have high importance values.

15. The apparatus of claim 9 , wherein the at least one processor is configured to cause the apparatus to determine an importance value of the information-to-importance functional kernel for the video coded slice or slice segment level based on at least one selected from a group comprising:

a video coded sequence of pictures;

a video frame type;

a video slice type; and

a video slice reference list.

16. The apparatus of claim 9 , wherein the at least one processor is configured to cause the apparatus to simplify a full connectivity of a plurality of video coding layers of at least one of spatial and temporal layer type to ignore graphical directed information flow model connections between layers at a lowest level and decoupling and offsetting the plurality of video coding layers by a constant scalar to determine their importance levels and their enclosed NAL units' importance levels.

17. The apparatus of claim 9 , wherein the at least one processor is configured to cause the apparatus to optimize a plurality of information-to-importance functional kernels and associated graphical directed information flow weights using statistical learning based on a generic set of training data of video coded sequences to at least one of minimize expected metric and maximize expected metric of visual rendering quality based on the plurality of feature sets associated with each of the plurality of NAL units of the video coded stream.

18. The apparatus of claim 1 , wherein the at least one processor is configured to cause the apparatus to signal the information-to-importance values of the plurality of NAL units to at least one of a radio access network (“RAN”) and a user equipment (“UE”) for optimization of transmission, retransmission, and reception procedures associated with the video coded stream over at least one of a limited capacity and varying capacity wireless link, the signaling performed using at least one indication selected from a group of a radio resource control (“RRC”) semi-static indication, a downlink control information (“DCI”) indication, and a dynamic header information indication, the optimization of transmission, retransmission, and reception procedures comprising at least one of packet scheduling, radio power control, radio frequency and resource allocation, physical layer modulation and coding scheme selection, hybrid automatic repeat request retransmissions preemptions, and multi-antenna and multiple transmission reception point transmissions, beamforming, and spatial multiplexing procedures.

19. A method of a network device, the method comprising:

detecting a plurality of video coded network abstraction layer (“NAL”) units of a video coded stream according to defined syntax elements of a determined video codec specification;

extracting semantic information associated with the plurality of NAL units according to a syntax elements and semantic knowledge base defined within the determined video codec specification;

combining the extracted semantic information associated with the plurality of NAL units to form a plurality of feature sets comprising a plurality of NAL unit type parameters, the plurality of feature sets correspondingly synchronized with the plurality of NAL units that enclose the extracted semantic information;

determining an information-to-importance value for each of the plurality of NAL units based on a NAL unit type, a NAL unit size, and a hierarchical encoding of each of the plurality of feature sets without performing video decoding of the video coded stream; and

indicating the determined information-to-importance value for each of the plurality of NAL units of the determined video codec specification to a video coded traffic-aware transceiver for scheduling video traffic based on the indicated information-to-importance value.

20. A remote network device apparatus, the apparatus comprising:

at least one memory; and

at least one processor coupled with the at least one memory and configured to cause the apparatus to:

encode and compress an uncompressed video sequence to a video coded stream formed of a plurality of network abstraction layer (“NAL”) units using a selected video codec specification;

extract semantic information associated with the plurality of NAL units according to a syntax elements and semantic knowledge base defined within the selected video codec specification;

combine the extracted semantic information associated with the plurality of NAL units to form a plurality of feature sets comprising a plurality of NAL unit type parameters, the plurality of feature sets correspondingly synchronized with the plurality of NAL units that enclose the extracted semantic information;

determine an information-to-importance value for each of the plurality of NAL units based on a NAL unit type, a NAL unit size, and a hierarchical encoding of each of the plurality of feature sets without performing video decoding of the video coded stream;

annotate the plurality of NAL units with the information-to-importance values for forming a plurality of application data units (“ADUs”) of the video coded stream for packet-switched communication networks;

signal the information-to-importance values for the plurality of ADUs and the plurality of NAL units;

process video coded traffic awareness information to determine optimized radio scheduling and control procedures for transmit and receive operations for the video coded stream ADUs; and

transmit the video coded stream ADUs to a radio access network (“RAN”) based on video coded traffic-aware optimizations.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2025
From: LENOVO (SINGAPORE) PTE. LTD.
To: EDGEWOOD IP, LLC
Reel/Frame 074118/0091 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: LENOVO (UNITED STATES) INC.
To: LENOVO (SINGAPORE) PTE. LTD.
Reel/Frame 061880/0110 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 12, 2022
From: STOICA, RAZVAN-ANDREI; BAGHERI, HOSSEIN; NANGIA, VIJAY
To: LENOVO (UNITED STATES) INC.
Reel/Frame 058635/0349 →
Continuity (1)
Related Publication 20230199198A1 · Jun 22, 2023
References Cited (74)
US 6404817B1 · Saha et al. · 2002 [cited by applicant]
US 6754278B1 · Suh et al. · 2004 [cited by applicant]
US 6909450B2 · Paulin · 2005 [cited by applicant]
US 6964021B2 · Jun et al. · 2005 [cited by applicant]
US 7886201B2 · Shi et al. · 2011 [cited by applicant]
US 7974341B2 · Chen et al. · 2011 [cited by applicant]
US 7987415B2 · Niu et al. · 2011 [cited by applicant]
US 8599316B2 · Deever · 2013 [cited by applicant]
US 8605221B2 · Deever · 2013 [cited by applicant]
US 8621320B2 · Yang et al. · 2013 [cited by applicant]
US 8654834B2 · Cho et al. · 2014 [cited by applicant]
US 8989559B2 · Chen et al. · 2015 [cited by applicant]
US 9386326B2 · Noru et al. · 2016 [cited by applicant]
US 9432298B1 · Smith · 2016 [cited by applicant]
US 10015486B2 · Kumar et al. · 2018 [cited by applicant]
US 10664687B2 · Suri et al. · 2020 [cited by applicant]
US 10764574B2 · Teo et al. · 2020 [cited by applicant]
US 11562565B2 · Cristache · 2023 [cited by applicant]
US 11917206B2 · Stoica et al. · 2024 [cited by applicant]
US 20020146074A1 · Ariel et al. · 2002 [cited by applicant]
US 20090213938A1 · Lee et al. · 2009 [cited by applicant]
US 20110090921A1 · Anthru · 2011 [cited by examiner]
US 20190075308A1 · Wei · 2019 [cited by examiner]
US 20200106554A1 · Kannan et al. · 2020 [cited by applicant]
US 20200157058A1 · Lee et al. · 2020 [cited by applicant]
US 20200195946A1 · Choi · 2020 [cited by examiner]
US 20200287654A1 · Xi et al. · 2020 [cited by applicant]
US 20200351936A1 · Kunt et al. · 2020 [cited by applicant]
US 20210028893A1 · Hwang et al. · 2021 [cited by applicant]
US 20210099715A1 · Topiwala et al. · 2021 [cited by applicant]
US 20210174155A1 · Smith et al. · 2021 [cited by applicant]
US 20210320956A1 · Berliner et al. · 2021 [cited by applicant]
US 20220413741A1 · Livis et al. · 2022 [cited by applicant]
EP 3041195A1 · 2008 [cited by applicant]
EP 2574010A2 · 2012 [cited by applicant]
WO WO2008066257A1 · 2008 [cited by examiner]
WO 2020139921A1 · 2020 [cited by applicant]
3GPP, “3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; System Architecture for the 5G System (5GS); Stage 2 (Release 17)”, 3GPP TS 23.501 V17.2.0, Sep. 2021, pp. 1-542. [cited by applicant]
ETSI, “Digital cellular telecommunications system (Phase 2+) (GSM); Universal Mobile Telecommunications System (UMTS) LTE; 5G; Transparent end-to-end packet switched streaming service (Release 16)”, 3GPP TS 26.244, ESTI… [cited by applicant]
ETSI, “Multiplexing and channel coding”, 3GPP TS 38.212, ESTI TS 138 212 V16.5.0, Apr. 2021 pp. 1-155. [cited by applicant]
ETSI, “Physical layer procedures for data”, 3GPP TS 38.214, ESTI TS 138 214 V16.5.0, Apr. 2021 pp. 1-173. [cited by applicant]
Rivaz, “AV1 Bitstream & Decoding Process Specification”, 2018 The Alliance for Open Media, Jan. 2019, pp. 1-681. [cited by applicant]
Rahmani, “Compressed domain visual information retrieval based on I-frames in HEVC”, SpringerLink, Mar. 4, 2016, pp. 1-14. [cited by applicant]
ETSI, “Extended Reality (XR) in 5G”, 3GPP TR 26.928 ESTI TR 126 928 V16.1.0, Jan. 2021 pp. 1-133. [cited by applicant]
ITU-T, “Series H: Audiovisual and Multimedia Systems: Infrastructure of audiovisual services—Coding of moving video! Advanced video coding for generic audiovisual services”, International Telecommunication Union, H.264,… [cited by applicant]
ITU-T, “Series H: Audiovisual and Multimedia Systems: Infrastructure of audiovisual services—Coding of moving video; High efficiency video coding”, International Telecommunication Union, H.265, Aug. 2021, pp. 1-716. [cited by applicant]
ITU-T, “Series H: Audiovisual and Multimedia Systems: Infrastructure of audiovisual services—Coding of moving video; Versatile video coding”, International Telecommunication Union, H.266, Aug. 2020, pp. 1-516. [cited by applicant]
3GPP, “FS_XRTraffic: Permanent document, v0.8.0”, Qualcomm Incorporated (Rapporteur), 3GPP TSG-SA4 Meeting #115e, S4211210, Aug. 27, 2021, pp. 1-130. [cited by applicant]
PCT/IB2022/062503, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or the Declaration”, International Searching Authority, Feb. 27, 2023,… [cited by applicant]
Ashfaq et al., “Hybrid Automatic repeat Request (HARQ) in Wireless Communications Systems and Standards: A Contemporary Survey”, IEEE Communications Surveys & Tutorials, vol. 23, No. 4, July 2, 1021, pp. 1-42. [cited by applicant]
Ducla-Soares et al., “Error resilience and concealment performance for MPEG-4 frame-based video coding”, Signal Processing. Image Communication, Elsevier Science Publishers, May 1, 1999, pp. 1-26. [cited by applicant]
Gao, “Flow-edge Guided Video Completion”, European Conference on Computer Vision, Aug. 2020, pp. 1-17. [cited by applicant]
Ericsson, “CSI Feedback Enhancements for IIoT/URLLC”, 3GPP TSG-RAN WG1 Meeting #104-e Tdoc R1-2100269, Jan. 25-Feb. 5, 2021, pp. 1-12. [cited by applicant]
Qualcomm Inc., “CSI enhancement for IOT and URLLC”, 3GPP TSG RAN WG1 #104-e R1-2101460, Jan. 25-Feb. 5, 2021, pp. 1-18. [cited by applicant]
Mediatek Inc., “Further Potential XR Enhancements”, 3GPP TSG RAN WG1 Meeting #106bis-e R1-2109556, Oct. 11-19, 2021. [cited by applicant]
Moderator of XR Enhancements (Nokia), Moderator's Summary on “[RAN94e-R18Prep-11] Enhancements for XR”, 3GPP TSG RAN#94e RP-212671, Dec. 6-17, 2021, pp. 1-54. [cited by applicant]
Nokia, “New SID on XR Enhancements for NR”, 3GPP TSG RAN Meeting #94e RP-212711, Dec. 6-17, 2021, pp. 1-5. [cited by applicant]
Sankisa, et. al., “Video error concealment using deep neural networks”, 2018 25th IEEE International Conference on Image Processing (ICIP) , Oct. 2018, pp. 1-5. [cited by applicant]
Shahriari et al., “Adaptive Error Concealment With Radial Basis Neuro-Fuzzy Networks For Video Communication over Lossy Channels,” First International Conference on Industrial and Information Systems, Aug. 8-11, 2006, p… [cited by applicant]
Shannon, “A Mathematical Theory of Communication” The Bell System Technical Journal, vol. 27, Oct. 1948, pp. 1-55. [cited by applicant]
Strinati, et. al., “6G Networks: Beyond Shannon Towards Semantic and Goal-Oriented Communications”, Computer Networks, Feb. 17, 2021, pp. 1-52. [cited by applicant]
3GPP, “3rd Generation Partnership Project; Technical Specification Group Radio Access Network; NR; Medium Access Control (MAC) protocol specification (Release 16)”, 3GPP TS 38.321 V16.3.0, Dec. 2020, pp. 1-156. [cited by applicant]
3GPP, “3rd Generation Partnership Project; Technical Specification Group Radio Access Network; NR; Radio Resource Control (RRC) protocol specification (Release 16)”, 3GPP TS 38.331 V16.1.0, Jul. 2020, pp. 1-906. [cited by applicant]
Ulyanov, et. al., “Deep Image Prior”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 1-9. [cited by applicant]
Zargari et al., “Visual Information Retrieval in HEVC Compressed Domain”, 2015 23rd Iranina Conference on Electrical Engineering (ICEE), May 10-14, 2015, pp. 1-6. [cited by applicant]
Zhang, et. al., “Intelligent Image and Video Compression: Communicating Pictures”, Academic Press; 2nd edition, Apr. 7, 2021, pp. 1-2. [cited by applicant]
PCT/IB2022/062496, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or the Declaration”, International Searching Authority, Mar. 17, 2023,… [cited by applicant]
PCT/IB2022/062500, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or the Declaration”, International Searching Authority, Mar. 27, 2023,… [cited by applicant]
Yip et al., “Joint source and channel coding for H.264 compliant stereoscopic video transmission”, Electrical and Computer Engineering, IEEE, May 1, 2005, pp. 1-5. [cited by applicant]
Dung et al., “Unequal Error Protection for H.26L Video Transmission”, Wireless Personal Multimedia Communications, IEEE, Oct. 27, 2002, pp. 1-6. [cited by applicant]
Naghdinezhad et al., “Frame distortion estimation for unequal error protection methods in scalable video coding (SVC)”, Signal Processing, Elsevier Science Publishers, Jul. 30, 2014, pp. 1-16. [cited by applicant]
Perera et al., “QoE aware resource allocation for video communications over LTE based mobile networks”, 10th International Conference on Heterogeneous Networking for Quality, Reliability, Security and Robustness, ICST, … [cited by applicant]
U.S. Appl. No. 18/317,781 Office Action Summary, USPTO, Jul. 25, 2024, pp. 1-32. [cited by applicant]
U.S. Appl. No. 18/317,781 Office Action Summary, USPTO, Nov. 12, 2024, pp. 1-8. [cited by applicant]