IP Library Granted Patent US 11,875,264
Granted Patent B2
US 11,875,264 · App. 16/823,227 · Granted Jan 16, 2024

Almost unsupervised cycle and action detection

Inventors: Krishnendu Chaudhury (Saratoga, CA); Ananya Honnedevasthana Ashok (Bangalore, IN); Sujay Narumanchi (Bangalore, IN); Devashish Shankar (Gwalior, IN); Ashish Mehra (Sunnyvale, CA)
Assignee: R4N63R Capital LLC
G06N3/084G06F18/214G06F18/2163G06F18/23G06F18/2431G06F18/24137G06V10/761G06V10/763G06V10/7715G06V10/82G06V20/41G06V20/49G06V40/20G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,875,264
App. No.
16/823,227
Granted
Jan 16, 2024
Kind
B2
Abstract

An event detection method can include encoding a plurality of training video snippets into low dimensional descriptors of the training video snippets in a code space. The low dimensional descriptors of the training video snippets can be decoded into corresponding reconstructed video snippets. One or more parameters of the encoding and decoding can be adjusted based on one or more a loss functions to reduce a reconstruction error between the one or more training video snippets and the corresponding one or more reconstructed video snippets, to reduce a class entropy of the plurality of event classes of the code space, to increase fit of the training video snippet, and/or to increase compactness of the code space. The method can further include encoding one or more labeled video snippets of a plurality of event classes into low dimensional descriptors of the labeled video snippets in the code space. The plurality of event classes can be mapped to class clusters corresponding to the low dimensional descriptors of the labeled video snippets. After training, query video snippets can be encoded into corresponding low dimensional descriptors in the code space. The low dimensional descriptors of the query video snippets can be classified based on their respective proximity to a nearest one of a plurality of class cluster of the code space. An event class of the query video snippet can be determined based on the class cluster classification.

Claims (61)

1. An event detection method comprising:

receiving a query video snippet;

encoding the query video snippet into a low dimensional descriptor of a code space, wherein the code space includes a plurality of class clusters characterized by one or more of a minimum trending reconstruction error, a minimum trending class entropy, a maximum trending fit of video snippets and a maximum trending compactness of the code space;

classifying the low dimensional descriptor of the query video snippet based on its proximity to a nearest one of a plurality of class clusters of the code space; and

outputting an indication of an event class of the query video snippet based on the classified class cluster of the low dimensional descriptor of the query video snippet.

2. The event detection method of claim 1 , wherein the low dimensional descriptors are encouraged to belong to any one of a plurality of class clusters, but discouraged from being between the plurality of class clusters in the code space.

3. The event detection method of claim 2 , wherein the event classes include a cycle class and a not cycle class.

4. The event detection method of claim 2 , wherein the event classes include a plurality of action cycle classes.

5. The event detection method of claim 1 , wherein the plurality of class clusters are mapped to corresponding event classes.

6. The event detection method of claim 1 , further comprising:

encoding a plurality of training video snippets into low dimensional descriptors of the training video snippets in the code space;

decoding the low dimensional descriptors of the training video snippets into corresponding reconstructed video snippets; and

adjusting one or more parameters of the encoding and decoding based on a loss function including one or more objective functions selected from a group including to reduce a reconstruction error between the one or more training video snippets and the corresponding one or more reconstructed video snippets, to reduce entropy class entropy of the plurality of event classes of the code space, to increase fit (of the training video snippet), and to increase compactness of the code space.

7. The event detection method of claim 6 , further comprising:

encoding one or more labeled video snippets of a plurality of event classes into low dimensional descriptors of the labeled video snippets in the code space; and

mapping the plurality of event classes to class clusters corresponding to the low dimensional descriptors of the labeled video snippets.

8. The event detection method of claim 6 , further comprising:

decoding the low dimensional descriptors of the query video snippets into corresponding reconstructed video snippets; and

further adjusting the one or more parameters of the encoding and decoding based on the loss function.

9. The event detection method of claim 1 , further comprising:

receiving indications of event classes of a plurality of query video snippets;

determining segments of contiguous video snippets having a same event class; and

outputting an indication of the segments of contiguous video snippets of the same event classes.

10. A event detection method comprising:

encoding a plurality of training video snippets into low dimensional descriptors of the training video snippets in a code space, including one or more labeled video snippets of a plurality of event classes;

mapping the plurality of event classes to class clusters corresponding to the low dimensional descriptors of the labeled video snippets;

decoding the low dimensional descriptors of the training video snippets into corresponding reconstructed video snippets; and

adjusting one or more parameters of the encoding and decoding based on a loss function including one or more objective functions selected from a group including to reduce a reconstruction error between the one or more training video snippets and the corresponding one or more reconstructed video snippets, to reduce a class entropy of the plurality of event classes of the code space, to increase fit of the training video snippet, and to increase compactness of the code space.

11. The event detection method of claim 10 , wherein the low dimensional descriptors are encouraged to belong to any one of a plurality of class clusters, but discouraged from being between the plurality of class clusters in the code space.

12. The event detection method of claim 10 , wherein the plurality of event classes include a cycle class and a not cycle class.

13. The event detection method of claim 10 , wherein the plurality of event classes include a plurality of action cycle classes.

14. An event detection device comprising:

a neural network encoder configured to encode a query video snippet into a low dimensional descriptor of a code space, wherein the code space includes a plurality of class clusters characterized by one or more of a minimum trending reconstruction error, a minimum trending class entropy, a maximum trending fit of video snippets and a maximum compactness of the code space; and

a class cluster classifier configured to classify the low dimensional descriptor of the query video snippet based on its proximity to a nearest one of a plurality of class clusters of the code space and output an indication of an event class of the query video snippet based on the classified class cluster of the low dimensional descriptor of the query video snippet.

15. The event detection device according to claim 14 , wherein the plurality of class clusters are mapped to corresponding event classes.

16. The event detection device according to claim 15 , wherein the event classes include a cycle class and a not cycle class.

17. The event detection device according to claim 15 , wherein the event classes include a plurality of action cycle classes.

18. The event detection device according to claim 14 , further comprising:

the neural network encoder further configured to encode one or more labeled video snippets of a plurality of event classes into low dimensional descriptors of the labeled video snippets in the code space; and

the class cluster classifier further configured to map the plurality of event classes to class clusters corresponding to the low dimensional descriptors of the labeled video snippets.

19. The event detection device according claim 14 , further comprising:

the neural network encoder further configured to encode a plurality of training video snippets into low dimensional descriptors of the training video snippets in the code space;

a neural network decoder configured to decode the low dimensional descriptors of the training video snippets into corresponding reconstructed video snippets; and

a loss function configured to adjust one or more parameters of the neural network encoder and the neural network decoder based on one or more objective functions selected from a group including to reduce a reconstruction error between the one or more training video snippets and the corresponding one or more reconstructed video snippets, to reduce class entropy of the plurality of event classes of the code space, to increase fit of the training video snippets, and to increase compactness of the code space.

20. The event detection device according claim 19 , further comprising:

the neural network decoder further configured to decode the low dimensional descriptors of the query video snippets into corresponding reconstructed video snippets; and

the loss function further configured to adjust the one or more parameters of the neural network encoder and the neural network decoder based on the one or more objective functions.

21. An event detection device comprising:

a neural network encoder configured to encode a plurality of training video snippets into low dimensional descriptors of the training video snippets in the code space, including one or more labeled video snippets of a plurality of event classes;

a class cluster classifier configured to map the plurality of event classes to class clusters corresponding to the low dimensional descriptors of the labeled video snippets;

a neural network decoder configured to decode the low dimensional descriptors of the training video snippets into corresponding reconstructed video snippets; and

a loss function configured to adjust one or more parameters of the neural network encoder and the neural network decoder based on one or more objective functions selected from a group including to reduce a reconstruction error between the one or more training video snippets and the corresponding one or more reconstructed video snippets, to reduce entropy class entropy of the plurality of event classes of the code space, to increase fit of the training video snippets, and to increase compactness of the code space.

22. The event detection device according claim 21 , wherein the plurality of class clusters are mapped to corresponding event classes.

23. The event detection device according claim 22 , wherein the event classes include a cycle class and a not cycle class.

24. The event detection device according claim 22 , wherein the event classes include a plurality of action cycle classes.

25. The event detection device according claim 21 , further comprising:

the neural network encoder further configured to encode a query video snippet into a low dimensional descriptor of the code space, wherein the code space includes a plurality of class clusters characterized by one or more of a minimum trending reconstruction error, a minimum trending class entropy, a maximum trending fit of video snippets and a maximum compactness of the code space; and

the class cluster classifier further configured to classify the low dimensional descriptor of the query video snippet based on its proximity to a nearest one of a plurality of class clusters of the code space and output an indication of an event class of the query video snippet based on the classified class cluster of the low dimensional descriptor of the query video snippet.

26. The event detection device according to claim 25 , further comprising:

a patch creator configured to subdivide frame of the query video snippet into a small number of patches and labeling patches with overlapping foreground objects; and

the neural network encoder further configured to encode the patches with overlapping foreground objects of the query video snippet into low dimensional descriptor of the code space.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2023
From: DRISHTI TECHNOLOGIES, INC.
To: R4N63R CAPITAL LLC
Reel/Frame 065626/0244 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2020
From: CHAUDHURY, KRISHNENDU; ASHOK, ANANYA HONNEDEVASTEHANA; NARUMANCHI, SUJAY; SHANKAR, DEVASHISH; MEHRA, ASHISH
To: DRISHTI TECHNOLOGIES, INC.
Reel/Frame 052811/0908 →
Continuity (2)
Provisional Application 62961407 · Jan 15, 2020
Related Publication 20210216777A1 · Jul 15, 2021