IP Library › Granted Patent US 11,157,744
Granted Patent B2
US 11,157,744 · App. 16/742,944 · Granted Oct 26, 2021

Automated detection and approximation of objects in video

Inventors: Udi Barzelay (Haifa, IL); Tal Hakim (Haifa, IL); Daniel Nechemia Rotman (Haifa, IL); Dror Porat (Haifa, IL)
Assignee: International Business Machines Corporation
G06K9/00744G06F16/7837G06K9/6256G06K9/6262G06N3/04G06N3/08G06T7/70G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,157,744
App. No.
16/742,944
Granted
Oct 26, 2021
Kind
B2
Abstract

Automated detection and approximation of objects in a video, including: (a) sampling a provided digital video, to obtain a set of sampled frames; (b) applying an object detection algorithm to the sampled frames, to detect objects appearing in the sampled frames; (c) based on the detections in the sampled frames, applying an object approximation algorithm to each sequence of frames that lie between the sampled frames, to approximately detect objects appearing in each of the sequences; (d) applying a trained regression model to each of the sequences, to estimate a quality of the approximate detection of objects in the respective sequence; (e) applying the object detection algorithm to one or more frames in those of the sequences whose quality of the approximate detection is below a threshold, to detect objects appearing in those frames.

Claims (63)

1. A method comprising:

(a) sampling a provided digital video, to obtain a set of sampled frames;

(b) applying an object detection algorithm to the sampled frames, to detect objects appearing in the sampled frames;

(c) based on the detections in the sampled frames, applying an object approximation algorithm to each sequence of frames that lie between the sampled frames, to approximately detect objects appearing in each of the sequences;

(d) applying a trained regression model to each of the sequences, to estimate a quality of the approximate detection of objects in the respective sequence by the object approximation algorithm;

(e) applying the object detection algorithm to one or more frames in those of the sequences whose quality of the approximate detection is below a threshold, to detect objects appearing in those frames;

(f) defining multiple sub-sequences that are different from the sequences, wherein each of the multiple sub-sequences comprises frames that lie between every adjacent pair of frames to which the object detection algorithm has been applied in steps (b) and (e); and

(g) re-applying the object approximation algorithm to each of the multiple sub-sequences.

2. The method according to claim 1 , further comprising:

obtaining a training set of digital videos;

for each of the digital videos of the training set:

applying the object detection algorithm to all frames of the respective digital video, to detect objects appearing in the frames of the respective digital video,

sampling the respective digital video, to obtain a set of sampled frames of the respective digital video,

applying the object approximation algorithm to frames of the respective digital video that lie between the sampled frames of the respective digital video, to approximately detect objects appearing in those frames that lie between the sampled frames of the respective digital video, and

for each of the frames that lie between the sampled frames of the respective digital video, comparing the approximate detection by the object approximation algorithm and the detection by the object detection algorithm, to estimate an accuracy of the approximate detection by the object approximation algorithm,

extracting features from frames of the respective digital video; and

training the regression model based on the estimated accuracy and the extracted features.

3. The method according to claim 2 , wherein the provided digital video and the digital videos of the training set are of a same genre.

4. The method according to claim 1 , wherein the threshold is determined according to a budget of computing resources that is available to operate the object detection algorithm.

5. The method according to claim 1 , wherein the object approximation algorithm is selected from the group consisting of: a tracking-based algorithm, an interpolation-based algorithm, an extrapolation-based algorithm, a duplication-based algorithm, and an Artificial Neural Network (ANN)-based algorithm.

6. The method according to claim 1 , executed on at least one hardware processor.

7. A system comprising:

(i) at least one hardware processor; and

(ii) a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by said at least one hardware processor to:

(a) sample a provided digital video, to obtain a set of sampled frames;

(b) apply an object detection algorithm to the sampled frames, to detect objects appearing in the sampled frames;

(c) based on the detections in the sampled frames, apply an object approximation algorithm to each sequence of frames that lie between the sampled frames, to approximately detect objects appearing in each of the sequences;

(d) apply a trained regression model to each of the sequences, to estimate a quality of the approximate detection of objects in the respective sequence by the object approximation algorithm;

(e) apply the object detection algorithm to one or more frames in those of the sequences whose quality of the approximate detection is below a threshold, to detect objects appearing in those frames;

(f) define multiple sub-sequences that are different from the sequences, wherein each of the multiple sub-sequences comprises frames that lie between every adjacent pair of frames to which the object detection algorithm has been applied in steps (b) and (e); and

(g) re-apply the object approximation algorithm to each of the multiple sub-sequences.

8. The system according to claim 7 , wherein the program code is further executable by said at least one hardware processor to:

obtain a training set of digital videos;

for each of the digital videos of the training set:

apply the object detection algorithm to all frames of the respective digital video, to detect objects appearing in the frames of the respective digital video,

sample the respective digital video, to obtain a set of sampled frames of the respective digital video,

apply the object approximation algorithm to frames of the respective digital video that lie between the sampled frames of the respective digital video, to approximately detect objects appearing in those frames that lie between the sampled frames of the respective digital video, and

for each of the frames that lie between the sampled frames of the respective digital video, compare the approximate detection by the object approximation algorithm and the detection by the object detection algorithm, to estimate an accuracy of the approximate detection by the object approximation algorithm,

extract features from frames of the respective digital video; and

train the regression model based on the estimated accuracy and the extracted features.

9. The system according to claim 8 , wherein the provided digital video and the digital videos of the training set are of a same genre.

10. The system according to claim 7 , wherein the threshold is determined according to a budget of computing resources that is available to operate the object detection algorithm.

11. The system according to claim 7 , wherein the object approximation algorithm is selected from the group consisting of: a tracking-based algorithm, an interpolation-based algorithm, an extrapolation-based algorithm, a duplication-based algorithm, and an Artificial Neural Network (ANN)-based algorithm.

12. A computer program product comprising a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by at least one hardware processor to:

(a) sample a provided digital video, to obtain a set of sampled frames;

(b) apply an object detection algorithm to the sampled frames, to detect objects appearing in the sampled frames;

(c) based on the detections in the sampled frames, apply an object approximation algorithm to each sequence of frames that lie between the sampled frames, to approximately detect objects appearing in each of the sequences;

(d) apply a trained regression model to each of the sequences, to estimate a quality of the approximate detection of objects in the respective sequence by the object approximation algorithm;

(e) apply the object detection algorithm to one or more frames in those of the sequences whose quality of the approximate detection is below a threshold, to detect objects appearing in those frames;

(f) define multiple sub-sequences that are different from the sequences, wherein each of the multiple sub-sequences comprises frames that lie between every adjacent pair of frames to which the object detection algorithm has been applied in steps (b) and (e); and

(g) re-apply the object approximation algorithm to each of the multiple sub-sequences.

13. The computer program product according to claim 12 , wherein the program code is further executable by said at least one hardware processor to:

obtain a training set of digital videos;

for each of the digital videos of the training set:

apply the object detection algorithm to all frames of the respective digital video, to detect objects appearing in the frames of the respective digital video,

sample the respective digital video, to obtain a set of sampled frames of the respective digital video,

apply the object approximation algorithm to frames of the respective digital video that lie between the sampled frames of the respective digital video, to approximately detect objects appearing in those frames that lie between the sampled frames of the respective digital video, and

for each of the frames that lie between the sampled frames of the respective digital video, compare the approximate detection by the object approximation algorithm and the detection by the object detection algorithm, to estimate an accuracy of the approximate detection by the object approximation algorithm,

extract features from frames of the respective digital video; and

train the regression model based on the estimated accuracy and the extracted features.

14. The computer program product according to claim 13 , wherein the provided digital video and the digital videos of the training set are of a same genre.

15. The computer program product according to claim 12 , wherein the threshold is determined according to a budget of computing resources that is available to operate the object detection algorithm.

16. The computer program product according to claim 12 , wherein the object approximation algorithm is selected from the group consisting of: a tracking-based algorithm, an interpolation-based algorithm, an extrapolation-based algorithm, a duplication-based algorithm, and an Artificial Neural Network (ANN)-based algorithm.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2020
From: BARZELAY, UDI; HAKIM, TAL; NECHEMIA ROTMAN, DANIEL; PORAT, DROR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051517/0262 →
Continuity (1)
Related Publication 20210216780A1 · Jul 15, 2021
Cited By (1)
US 12,499,674