IP Library Granted Patent US 11,087,137
Granted Patent B2
US 11,087,137 · App. 16/278,377 · Granted Aug 10, 2021

Methods and systems for identification and augmentation of video content

Inventors: Vinay Polavarapu (Irving, TX); Christian Egeler (Roseland, NJ); Sergey Virodov (Fair Lawn, NJ); William Patrick Dildine (Hoboken, NJ); Adwait Ashish Murudkar (Highland Park, NJ); Paul Duree (Livermore, CA); Gang Lin (Colorado Springs, CO); Paul Cannon (Frederick, MD); Alexander Zilbershtein (Buffalo Grove, IL); Pai Moodlagiri (Clinton, NJ)
Assignee: Verizon Patent and Licensing Inc.
G06K9/00744G06K9/00671G06K9/00718G06K9/3233G06K9/6256G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,087,137
App. No.
16/278,377
Granted
Aug 10, 2021
Kind
B2
Abstract

An exemplary object identification system detects, based on a machine learning model, an object depicted within a video frame. The system identifies, based on the detecting of the object, a class label of the object and a region of interest, within the video frame, of the object. The system identifies, within the region of interest of the object, set of features of the object. The system compares the set of features of the object with a plurality of predefined features within a data store associated with the class label of the object. The system determines, based on the comparing of the set of features of the object with the plurality of predefined features within the data store, whether the object is configured to trigger an augmentation of video content associated with the video frame. Corresponding methods and systems are also disclosed.

Claims (69)

1. A method comprising:

detecting, by an object identification system and based on a machine learning model, an object depicted within a video frame;

identifying, by the object identification system and based on the detecting of the object, a class label of the object and a region of interest, within the video frame, of the object;

identifying, by the object identification system and within the region of interest of the object, a set of features of the object;

selecting, by the object identification system based on the class label of the object and from a plurality of data stores each associated with a different class label, a data store associated with the class label of the object;

comparing, by the object identification system, the set of features of the object with a plurality of predefined features within the data store associated with the class label of the object; and

determining, by the object identification system and based on the comparing of the set of features of the object with the plurality of predefined features within the data store, whether the object is configured to trigger an augmentation of video content associated with the video frame.

2. The method of claim 1 , further comprising rendering, by an augmentation system in response to a determination that the object is configured to trigger the augmentation of the video content associated with the video frame, an augmented video frame with augmentation content included within the augmented video frame, wherein the augmentation content is based on the set of features of the object.

3. The method of claim 1 , wherein each of the plurality of data stores includes a distinct plurality of predefined features.

4. The method of claim 1 , further comprising:

detecting, by the object identification system and based on the machine learning model, a second object depicted within the video frame;

identifying, by the object identification system and based on the detecting of the second object, a class label of the second object and a region of interest, within the video frame, of the second object;

identifying, by the object identification system and within the region of interest of the second object, a set of features of the second object;

comparing, by the object identification system, the set of features of the second object with a second plurality of predefined features within a second data store associated with the class label of the second object, wherein the comparing of the set of features of the second object with the second plurality of predefined features within the second data store is performed in parallel with the comparing of the set of features of the object with the plurality of predefined features within the data store; and

determining, by the object identification system based on the comparing of the set of features of the second object with the second plurality of predefined features within the second data store, whether the second object is configured to trigger a second augmentation of second video content associated with the video frame, wherein the determining whether the second object is configured to trigger the second augmentation of the second video content associated with the video frame is performed in parallel with the determining whether the object is configured to trigger the augmentation of the video content associated with the video frame.

5. The method of claim 4 , further comprising receiving, by the object identification system over a low-latency network connection, data representative of the video frame from a device that captured the video frame;

wherein the performing of the determining whether the object is configured to trigger the augmentation of the video content associated with the video frame and the determining whether the second object is configured to trigger a second augmentation of second video content associated with the video frame in parallel facilitates concurrent low-latency network-based augmentations of the video content based on the object and the second object depicted in the video frame.

6. The method of claim 1 , further comprising:

receiving, by the object identification system, data representative of a new object to be detected by the machine learning model;

identifying, by the object identification system in response to the receiving of the data representative of the new object, one or more features of the new object; and

storing, by the object identification system, the one or more identified features of the new object within one of the plurality of data stores.

7. The method of claim 1 , wherein the object identification system is implemented within a cloud-based network edge server having a low-latency network connection with a device that captured the video frame.

8. A non-transitory computer-readable medium having stored thereon computer-readable instructions, which instructions, when executed by a processor, perform the method of claim 1 .

9. A system comprising:

at least one memory storing instructions; and

at least one processor communicatively coupled to the at least one memory and configured to execute the instructions to:

detect, based on a machine learning model, an object depicted within a video frame;

identify, based on the detection of the object, a class label of the object and a region of interest, within the video frame, of the object;

identify, within the region of interest of the object, a set of features of the object;

select, based on the class label of the object and from a plurality of data stores each associated with a different class label, a data store associated with the class label of the object;

compare the set of features of the object with a plurality of predefined features within the data store associated with the class label of the object; and

determine, based on the comparison of the set of features of the object with the plurality of predefined features within the data store, whether the object is configured to trigger an augmentation of video content associated with the video frame.

10. The system of claim 9 , further comprising at least one other processor configured to render, in response to a determination that the object is configured to trigger the augmentation of the video content associated with the video frame, an augmented video frame with augmentation content included within the augmented video frame, wherein the augmentation content is based on the set of features of the object.

11. The system of claim 9 , each of the plurality of data stores includes a distinct plurality of predefined features.

12. The system of claim 9 , wherein the at least one processor is further configured to:

detect, based on the machine learning model, a second object depicted within the video frame;

identify, based on the detection of the second object, a class label of the second object and a region of interest, within the video frame, of the second object;

identify, within the region of interest of the second object, set of features of the second object;

compare the set of features of the second object with a second plurality of predefined features within a second data store associated with the class label of the second object, wherein the comparison of the set of features of the second object with the second plurality of predefined features within the second data store is performed in parallel with the comparison of the set of features of the object with the plurality of predefined features within the data store; and

determine, based on the comparison of the set of features of the second object with the second plurality of predefined features within the second data store, whether the second object is configured to trigger a second augmentation of second video content associated with the video frame, wherein the determination of whether the second object is configured to trigger the second augmentation of the second video content associated with the video frame is performed in parallel with the determination of whether the object is configured to trigger the augmentation of the video content associated with the video frame.

13. The system of claim 12 , wherein:

the at least one processor is implemented within a network edge server having a low-latency network connection with a device that captured the video frame; and

the performance of the determination of whether the object is configured to trigger the augmentation of the video content associated with the video frame and the determination of whether the second object is configured to trigger a second augmentation of second video content associated with the video frame in parallel facilitates concurrent low-latency network-based augmentations of the video content based on the object and the second object depicted in the video frame.

14. The system of claim 9 , wherein the at least one processor is further configured to:

receive data representative of a new object to be detected by the machine learning model;

identify, in response to the reception of the data representative of the new object, one or more features of the new object; and

store the one or more identified features of the new object within one of the plurality of data stores.

15. The system of claim 9 , wherein the system is implemented within a cloud-based network edge server having a low-latency network connection with a device that captured the video frame.

16. A system comprising:

at least one physical computing device that includes at least one processor and that is configured to:

detect, based on a machine learning model, an object depicted within a video frame,

identify, based on the detection of the object, a class label of the object and a region of interest, within the video frame, of the object,

identify, within the region of interest of the object, a set of features of the object,

select, based on the class label of the object, a data store from a plurality of data stores each storing a respective plurality of predefined features,

compare the set of features of the object with the plurality of predefined features stored within the selected data store, and

determine, based on the comparison of the set of features of the object with the plurality of predefined features stored within the data store, whether the object is configured to trigger an augmentation of video content associated with the video frame.

17. The system of claim 16 , wherein the at least one physical computing device further implements an augmentation system configured to render, in response to a determination that the object is configured to trigger the augmentation of the video content associated with the video frame, an augmented video frame with augmentation content included within the augmented video frame, wherein the augmentation content is based on the set of features of the object.

18. The system of claim 16 , wherein:

the at least one physical computing device is further configured to:

detect, based on the machine learning model, a second object depicted within the video frame;

identify, based on the detection of the second object, a class label of the second object and a region of interest, within the video frame, of the second object;

identify, within the region of interest of the second object, set of features of the second object;

compare the set of features of the second object with a second plurality of predefined features within a second data store associated with the class label of the second object, wherein the comparison of the set of features of the second object with the second plurality of predefined features within the second data store is performed in parallel with the comparison of the set of features of the object with the plurality of predefined features within the data store; and

determine, based on the comparison of the set of features of the second object with the second plurality of predefined features within the second data store, whether the second object is configured to trigger a second augmentation of second video content associated with the video frame, wherein the determination of whether the second object is configured to trigger the second augmentation of the second video content associated with the video frame is performed in parallel with the determination of whether the object is configured to trigger the augmentation of the video content associated with the video frame.

19. The system of claim 16 , wherein the at least one physical computing device is further configured to:

receive data representative of a new object to be detected by the machine learning model;

identify, in response to the reception of the data representative of the new object, a set of features of the new object; and

store the identified set of features of the new object within one of the plurality of data stores.

20. The system of claim 16 , wherein the at least one physical computing device is implemented by at least one cloud-based network edge server having a low-latency network connection with a device that captured the video frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2019
From: POLAVARAPU, VINAY; EGELER, CHRISTIAN; VIRODOV, SERGEY; DILDINE, WILLIAM PATRICK; MURUDKAR, ADWAIT ASHISH; DUREE, PAUL; LIN, GANG; CANNON, PAUL; ZILBERSHTEIN, ALEXANDER; MOODLAGIRI, PAI
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 048360/0658 →
Continuity (1)
Related Publication 20200265238A1 · Aug 20, 2020