IP Library Granted Patent US 12,284,397
Granted Patent B2
US 12,284,397 · App. 17/937,220 · Granted Apr 22, 2025

Lossy compression of video content into a graph representation

Inventors: Dotan Di Castro (Haifa, IL); Eitan Kosman (Haifa, IL)
Assignee: ROBERT BOSCH GMBH
H04N19/90H04N19/172H04N19/186
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,284,397
App. No.
17/937,220
Granted
Apr 22, 2025
Kind
B2
Abstract

A method for lossily compressing a sequence of video frames into a representation, wherein each video frame comprises pixels that carry color values. The method includes: segmenting each video frame into superpixels, wherein these superpixels are groups of pixels that share at least one predetermined common property; assigning, to each superpixel in each video frame, at least one attribute derived from the pixels belonging to the respective superpixel; and combining superpixels as nodes in a graph representation, wherein superpixels in a same video frame are connected by spatial edges associated with at least one quantity that is a measure for a distance between these superpixels; and in response to superpixels in adjacent video frames in the sequence meeting at least one predetermined relatedness criterion, these superpixels are connected by temporal edges.

Claims (39)

1. A method for lossily compressing a sequence of video frames into a representation, wherein each video frame for the video frames includes pixels that carry color values, the method comprising the following steps:

segmenting each video frame into superpixels, wherein the superpixels are groups of pixels that share at least one predetermined common property;

assigning, to each superpixel in each video frame, at least one attribute derived from the pixels belonging to the respective superpixel;

combining superpixels as nodes in a graph representation, wherein:

superpixels in a same video frame are connected by spatial edges associated with at least one quantity that is a measure for a distance between the superpixels in the same video frame, and

in response to superpixels in adjacent video frames in the sequence meeting at least one predetermined relatedness criterion, the superpixels in the adjacent video frames are connected by temporal edges, wherein the relatedness criterion is a threshold value of a distance between the superpixels in adjacent video frames in the sequence, wherein the temporal edges connect the superpixels in the adjacent video frames when the relatedness criterion is met between the superpixels;

providing the graph representation to a graph neural network (GNN); and

obtaining, from the GNN, a processing result for the sequence of video frames.

2. The method of claim 1 , wherein the attribute assigned to each superpixel includes a minimum color value or a maximum color value or a mean color value or a median color value or another aggregate value derived from the color values of pixels belonging to the superpixel.

3. The method of claim 1 , wherein the measure for the distance between superpixels includes an Euclidean distance between spatial coordinates of superpixels.

4. The method of claim 3 , wherein the spatial coordinates of each superpixel include spatial coordinates of a centroid of the pixels belonging to the superpixel.

5. The method of claim 1 , wherein the measure for the distance between respective superpixels includes a difference computed between histograms of properties of individual pixels belonging to the respective superpixels.

6. The method of claim 1 , wherein the relatedness criterion includes a proximity with respect to spatial coordinates of the superpixels, and/or a similarity of attributes assigned to these superpixels.

7. The method of claim 1 , further comprising:

pre-selecting, given a first superpixel in a video frame, superpixels from an adjacent video frame in the sequence that meet a first relatedness criterion with respect to proximity; and

choosing, from the pre-selected superpixels, a superpixel whose assigned attributes are most similar to those of the first superpixel as a superpixel to connect to the first superpixel by one of the temporal edge.

8. The method of claim 1 , further comprising: in response to determining that a superpixel of the superpixels belongs to a background or other area of the video frame that is not relevant to an application at hand, excluding and/or removing the superpixel from the graph representation.

9. The method of claim 1 , wherein the GNN is configured to map the graph representation to one or more classification scores with respect to a given set of available classes.

10. The method of claim 1 , further comprising:

computing, from the processing result obtained from the GNN, an actuation signal; and

actuating, with the actuation signal, a vehicle and/or a quality inspection system and/or a classification system and/or a surveillance system.

11. The method of claim 1 , further comprising:

retrieving, from at least one database, media content or other information, stored in association with the graph representation.

12. A non-transitory machine-readable storage medium on which is stored a computer program for lossily compressing a sequence of video frames into a representation, wherein each video frame for the video frames includes pixels that carry color values, the computer program, when executed by one or more computers, causing the one or more computers to perform the following steps:

segmenting each video frame into superpixels, wherein the superpixels are groups of pixels that share at least one predetermined common property;

assigning, to each superpixel in each video frame, at least one attribute derived from the pixels belonging to the respective superpixel; and

combining superpixels as nodes in a graph representation, wherein:

superpixels in a same video frame are connected by spatial edges associated with at least one quantity that is a measure for a distance between the superpixels in the same video frame,

in response to superpixels in adjacent video frames in the sequence meeting at least one predetermined relatedness criterion, the superpixels in the adjacent video frames are connected by temporal edges, wherein the relatedness criterion is a threshold value of a distance between the superpixels in adjacent video frames in the sequence, wherein the temporal edges connect the superpixels in the adjacent video frames when the relatedness criterion is met between the superpixels;

providing the graph representation to a graph neural network (GNN); and

obtaining, from the GNN, a processing result for the sequence of video frames.

13. One or more computers configured to lossily compress a sequence of video frames into a representation, wherein each video frame for the video frames includes pixels that carry color values, the one or more computers configured to:

segment each video frame into superpixels, wherein the superpixels are groups of pixels that share at least one predetermined common property;

assign, to each superpixel in each video frame, at least one attribute derived from the pixels belonging to the respective superpixel; and

combine superpixels as nodes in a graph representation, wherein:

superpixels in a same video frame are connected by spatial edges associated with at least one quantity that is a measure for a distance between the superpixels in the same video frame,

in response to superpixels in adjacent video frames in the sequence meeting at least one predetermined relatedness criterion, the superpixels in the adjacent video frames are connected by temporal edges, wherein the relatedness criterion is a threshold value of a distance between the superpixels in adjacent video frames in the sequence, wherein the temporal edges connect the superpixels in the adjacent video frames when the relatedness criterion is met between the superpixels;

providing the graph representation to a graph neural network (GNN); and

obtaining, from the GNN, a processing result for the sequence of video frames.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2023
From: DI CASTRO, DOTAN; KOSMAN, EITAN
To: ROBERT BOSCH GMBH
Reel/Frame 062407/0926 →
Priority Claims (1)
EP 21201632 · Oct 8, 2021 · regional
Continuity (1)
Related Publication 20230115248A1 · Apr 13, 2023
References Cited (23)
US 11908180B1 · Ho · 2024 [cited by examiner]
US 20140362240A1 · Klivington · 2014 [cited by examiner]
US 20160155024A1 · Partis · 2016 [cited by examiner]
US 20160171707A1 · Schwartz · 2016 [cited by examiner]
US 20170094288A1 · Hannuksela · 2017 [cited by examiner]
US 20180278957A1 · Fracastoro · 2018 [cited by examiner]
US 20190058887A1 · Wang · 2019 [cited by examiner]
US 20200175325A1 · Nie · 2020 [cited by examiner]
US 20200184655A1 · Xiang · 2020 [cited by examiner]
US 20210350620A1 · Bronstein · 2021 [cited by examiner]
US 20220245802A1 · Wang · 2022 [cited by examiner]
US 20220261590A1 · Brahma · 2022 [cited by examiner]
US 20220284552A1 · Yang · 2022 [cited by examiner]
US 20220400200A1 · Lagnado · 2022 [cited by examiner]
US 20230065773A1 · Dimitriou · 2023 [cited by examiner]
US 20240404254A1 · Ouyang · 2024 [cited by examiner]
Grundmann et al., “Efficient Hierarchical Graph-Based Video Segmentation,” 2010 IEEE Computer Society Conference On Computer Vision and Pattern Recognition, 2010, pp. 2141-2148. [cited by applicant]
Zhang et al., “Saliency-Guided Unsupervised Feature Learning for Scene Classification,” IEEE Transactions On Geoscience and Remote Sensing, vol. 53, No. 4, 2015, pp. 2175-2184. [cited by applicant]
Yu et al., “Efficient Video Segmentation Using Parametric Graph Partitioning,” 2015 IEEE International Conference On Computer Vision, 2015, pp. 3155-3163. [cited by applicant]
Rogers, “Normalizing the X.Y Coordinates for Different Image Sizes,” 2019, pp. 1. https://forum.image.sc/t/normalizing-the-x-y-coordinates-for-different-image-sizes/24160. [cited by applicant]
Lv, “An Improved Slic Superpixels Using Reciprocal Nearest Neighbor Clustering,” International Journal of Signal Processing, Image Processing and Pattern Recognition, vol. 8, No. 5, 2015, pp. 239-248. [cited by applicant]
Jain et al., “Supervoxel-Consistent Foreground Propagation in Video,” ECCV 2014, Part IV, NCS 8692, 2014, pp. 656-671. [cited by applicant]
Guo et al., “Depth Estimation From a Single Image in Pedestrian Candidate Generation,” 2016 IEEE 11th Conference On Industrial Electronics and Applications (ICIEA), 2016, pp. 1005-1008. [cited by applicant]