IP Library Granted Patent US 11,887,384
Granted Patent B2
US 11,887,384 · App. 17/164,954 · Granted Jan 30, 2024

In-cabin occupant behavoir description

Inventors: Zilong Hu (San Jose, CA); Lei Zhang (Campbell, CA); Qun Gu (San Jose, CA)
Assignee: Black Sesame Technologies Inc.
G06V20/59G06T7/70G06V10/40G06V20/41G06V20/47G06V40/20B60Q9/00G06T2207/10016G06T2207/30196G06T2207/30268
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,887,384
App. No.
17/164,954
Granted
Jan 30, 2024
Kind
B2
Abstract

A method of describing a temporal event, including receiving a video sequence of the temporal event, extracting at least one physical characteristic of an at least one occupant within the video sequence, extracting at least one action of the at least one occupant within the video sequence, extracting at least one interaction of the at least one occupant with a secondary occupant within the video sequence, determining a safety level of the temporal event within a vehicle based on at least one of the at least one action and the at least one interaction and describing the at least one physical characteristic of the at least one occupant and at least one of the at least one action and the at least one interaction of the at least one occupant.

Claims (34)

1. A method of describing a temporal event, comprising:

receiving a live stream video sequence of the temporal event recorded by an in-cabin camera of a vehicle;

dividing the video sequence into multiple video clips;

extracting at least one physical characteristic of an at least one occupant within video clip;

extracting at least one action of the at least one occupant within the video clip based on a previous action with a previous video clip;

extracting at least one interaction of the at least one occupant with a secondary occupant within the video clip based on a previous action with a previous video clip;

determining a safety level of the temporal event within a vehicle based on the at least one action and the at least one interaction;

describing the at least one physical characteristic of the at least one occupant and the at least one action and the at least one interaction of the at least one occupant; and

wherein the extracting at least one action and extracting at least one interaction is performed by a convolutional gated recurrent unit (GRU), wherein the output of the GRU of the previous video clip is sent to the GRU of the at least one video clip; wherein the GRU allows action and interaction description for occupants detected within the video sequence by retaining spatial information during a forward pass to process temporal features.

2. The method of claim 1 wherein the at least one physical characteristic of the at least one occupant includes a location.

3. The method of claim 1 wherein the at least one physical characteristic of the at least one occupant includes at least one of a gender and an age.

4. The method of claim 1 wherein the at least one physical characteristic of the at least one occupant includes at least one emotional state.

5. The method of claim 1 wherein the at least one action of the at least one occupant includes at least one of calling, talking and arguing.

6. The method of claim 1 further comprising alarming if the safety level of the temporal event exceeds a predetermined safety threshold.

7. The method of claim 1 further comprising generating at least one action label for the at least one action of the at least one occupant.

8. The method of claim 1 further comprising generating at least one interaction label for the at least one action of the at least one occupant.

9. The method of claim 1 further comprising generating a scene summary of the video sequence of the temporal event of the at least one occupant.

10. A method of describing a temporal event, comprising:

receiving a live stream video sequence of the temporal event recorded by an in-cabin camera of a vehicle;

dividing the video sequence into multiple video clips;

extracting at least one spatial characteristic of an at least one occupant within the video clip based on a previous action with a previous video clip;

extracting at least one temporal action of the at least one occupant within the video clip based on a previous action with a previous video clip;

extracting at least one temporal interaction of the at least one occupant with a secondary occupant within video clip;

determining a safety level of the temporal event within a vehicle based on the at least one temporal action and the at least one temporal interaction of the at least one occupant;

describing the at least one spatial characteristic of the at least one occupant and the at least one temporal action and the at least one temporal interaction of the at least one occupant; and

wherein the extracting at least one temporal action and extracting at least one temporal interaction is performed by a convolutional gated recurrent unit (GRU), wherein the output of the GRU of the previous video clip is sent to the GRU of the at least one video clip; wherein the GRU allows action and interaction description for occupants detected within the video sequence by retaining spatial information during a forward pass to process temporal features.

11. The method of claim 10 wherein the at least one spatial characteristic of the at least one occupant includes a location.

12. The method of claim 10 wherein the at least one spatial characteristic of the at least one occupant includes at least one of a gender and an age.

13. The method of claim 10 wherein the at least one spatial characteristic of the at least one occupant includes at least one emotional state.

14. The method of claim 10 wherein the at least one temporal action of the at least one occupant includes at least one of calling, talking and arguing.

15. The method of claim 10 further comprising alarming if the safety level of the temporal event exceeds a predetermined safety threshold.

16. The method of claim 10 further comprising generating at least one action label for the at least one temporal action of the at least one occupant.

17. The method of claim 10 further comprising generating at least one interaction label for the at least one temporal action of the at least one occupant.

18. The method of claim 10 further comprising generating a scene summary of the video sequence of the temporal event of the at least one occupant.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2021
From: BLACK SESAME INTERNATIONAL HOLDING LIMITED
To: BLACK SESAME TECHNOLOGIES INC.
Reel/Frame 058301/0364 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2021
From: HU, ZILONG; ZHANG, LEI; GU, QUN
To: BLACK SESAME INTERNATIONAL HOLDING LIMITED
Reel/Frame 057903/0389 →
Continuity (1)
Related Publication 20220245388A1 · Aug 4, 2022