IP Library Granted Patent US 11,532,169
Granted Patent B1
US 11,532,169 · App. 17/347,682 · Granted Dec 20, 2022

Distracted driving detection using a multi-task training process

Inventors: Ali Hassan (Lahore, PK); Ijaz Akhter (Lahore, PK); Muhammad Faisal (Jhawarian, PK); Afsheen Rafaqat Ali (Gujranwala, PK); Ahmed Ali (Lahore, PK)
Assignee: MOTIVE TECHNOLOGIES, INC.
G06V20/597G06K9/6257G06K9/6264G06V10/22G06V10/95
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,532,169
App. No.
17/347,682
Granted
Dec 20, 2022
Kind
B1
Abstract

Disclosed are a multi-task training technique and resulting model for detecting distracted driving. In one embodiment, a method is disclosed comprising inputting a plurality of labeled examples into a multi-task network, the multi-task network comprising: a backbone network, the backbone network generating one or more feature vectors corresponding to each of the labeled examples, and a plurality of prediction heads coupled to the backbone network; minimizing a joint loss based on outputs of the plurality of prediction heads, the minimizing the joint loss causing a change in parameters of the backbone network; and storing a distraction classification model after minimizing the joint loss, the distraction classification model comprising the parameters of the backbone network and parameters of at least one of the prediction heads.

Claims (34)

1. A method comprising:

inputting a plurality of labeled examples into a multi-task network, the multi-task network comprising:

a backbone network comprising a convolutional neural network (CNN) and a feature pyramid network (FPN) coupled to the CNN, the backbone network generating one or more feature vectors corresponding to each of the labeled examples, and

a plurality of prediction heads coupled to the backbone network, wherein a subset of the plurality of prediction heads receives input from the CNN and a second subset of the plurality of prediction heads receives input from the FPN;

minimizing a joint loss based on outputs of the plurality of prediction heads, the minimizing the joint loss causing a change in parameters of the backbone network; and

storing a distraction classification model after minimizing the joint loss, the distraction classification model comprising the parameters of the backbone network and parameters of at least one of the prediction heads.

2. The method of claim 1 , wherein the CNN comprises an EfficientNet.

3. The method of claim 1 , wherein the FPN comprises a bi-directional FPN.

4. The method of claim 1 , wherein a subset of the plurality of prediction heads comprises a distraction classification prediction head, the distraction classification head comprising a convolutional layer, pooling layer, and fully connected layer.

5. The method of claim 1 , wherein a second subset of the plurality of prediction heads includes one or more of an object detection prediction head and a pose estimation prediction head.

6. The method of claim 5 , wherein the object detection prediction head comprises a bounding box regression network and an object class prediction network, each of the bounding box regression network and the object class prediction network comprising deep neural networks having a plurality of hidden layers, each hidden layer in the hidden layers comprising a convolutional layer, a batch normalization layer, and a batch activation layer, wherein the bounding box regression network outputs coordinates of a bounding box enclosing a detected object and the object class prediction network outputs a class corresponding to the detected object.

7. The method of claim 5 , wherein the pose estimation prediction head comprises a deep neural network, the deep neural network comprising a plurality of hidden layers and an output layer, each of the hidden layers comprising a convolutional layer, a batch normalization layer, and an activation layer, and the output layer comprising a convolutional layer.

8. The method of claim 1 , wherein storing a distraction classification model after minimizing the joint loss comprises storing parameters of the CNN and at least one of the prediction heads.

9. A non-transitory computer-readable storage medium for tangibly storing computer program instructions capable of being executed by a computer processor, the computer program instructions defining operations of:

inputting a plurality of labeled examples into a multi-task network, the multi-task network comprising:

a backbone network comprising a convolutional neural network (CNN) and a feature pyramid network (FPN) coupled to the CNN, the backbone network generating one or more feature vectors corresponding to each of the labeled examples, and

a plurality of prediction heads coupled to the backbone network, wherein a subset of the plurality of prediction heads receives input from the CNN and a second subset of the plurality of prediction heads receives input from the FPN;

minimizing a joint loss based on outputs of the plurality of prediction heads, the minimizing the joint loss causing a change in parameters of the backbone network; and

storing a distraction classification model after minimizing the joint loss, the distraction classification model comprising the parameters of the backbone network and parameters of at least one of the prediction heads.

10. The non-transitory computer-readable storage medium of claim 9 , wherein the CNN comprises an EfficientNet.

11. The non-transitory computer-readable storage medium of claim 9 , wherein the FPN comprises a bi-directional FPN.

12. The non-transitory computer-readable storage medium of claim 9 , wherein a subset of the plurality of prediction heads comprises a distraction classification prediction head, the distraction classification head comprising a convolutional layer, pooling layer, and fully connected layer.

13. The non-transitory computer-readable storage medium of claim 9 , wherein a second subset of the plurality of prediction heads includes one or more of an object detection prediction head and a pose estimation prediction head.

14. The non-transitory computer-readable storage medium of claim 9 , wherein storing a distraction classification model after minimizing the joint loss comprises storing parameters of the CNN and at least one of the prediction heads.

15. A device comprising:

a processor configured to:

input a plurality of labeled examples into a multi-task network, the multi-task network comprising: a backbone network comprising a convolutional neural network (CNN) and a feature pyramid network (FPN) coupled to the CNN, the backbone network generating one or more feature vectors corresponding to each of the labeled examples, and a plurality of prediction heads coupled to the backbone network, wherein a subset of the plurality of prediction heads receives input from the CNN and a second subset of the plurality of prediction heads receives input from the FPN,

minimize a joint loss based on outputs of the plurality of prediction heads, the minimizing the joint loss causing a change in parameters of the backbone network, and

store a distraction classification model after minimizing the joint loss, the distraction classification model comprising the parameters of the backbone network and parameters of at least one of the prediction heads.

16. The device of claim 15 , wherein storing a distraction classification model after minimizing the joint loss comprises storing parameters of the backbone network and at least one of the prediction heads.

17. The device of claim 15 , wherein the CNN comprises an EfficientNet.

18. The device of claim 15 , wherein the FPN comprises a bi-directional FPN.

19. The device of claim 15 , wherein a subset of the plurality of prediction heads comprises a distraction classification prediction head, the distraction classification head comprising a convolutional layer, pooling layer, and fully connected layer.

20. The device of claim 15 , wherein a second subset of the plurality of prediction heads includes one or more of an object detection prediction head and a pose estimation prediction head.

Assignments (3)
CHANGE OF NAME Recorded Apr 12, 2022
From: KEEP TRUCKIN, INC.
To: MOTIVE TECHNOLOGIES, INC.
Reel/Frame 059965/0872 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2022
From: ALI, AHMED
To: KEEP TRUCKIN, INC.
Reel/Frame 058549/0222 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2021
From: HASSAN, ALI; AKHTER, IJAZ; FAISAL, MUHAMMAD; ALI, AFSHEEN RAFAQAT
To: KEEP TRUCKIN, INC.
Reel/Frame 056542/0192 →
Cited By (33)
US 12,197,610 US 12,213,090 US 12,228,944 US 12,253,617 US 12,256,021 US 12,260,616 US 12,269,498 US 12,289,181 US 12,306,010 US 12,327,445 US 12,328,639 US 12,344,168 US 12,346,712 US 12,367,718 US 12,426,007 US 12,445,285 US 12,450,329 US 12,479,446 US 12,501,178 US 12,511,947 US 12,534,097 US 12,561,624 US 12,565,143 US 12,626,200 US 12,630,050 US 12,646,402 US 12,651,529 US 12,662,152 US 12,665,989 US 12,671,464 US 12,675,419 US 12,701,004 US 12,730,425