IP Library Granted Patent US 10,522,147
Granted Patent B2
US 10,522,147 · App. 15/850,814 · Granted Dec 31, 2019

Device and method for generating text representative of lip movement

Inventors: Stephen Varner (Xenia, OH); Wei Lin (Lake Zurich, IL); Randy L. Ekl (Downers Grove, IL); Daniel A. Law (Glencoe, IL)
Assignee: MOTOROLA SOLUTIONS, INC.
G10L15/25G06K9/00335G06K9/00711G10L15/26G10L25/57G10L25/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,522,147
App. No.
15/850,814
Granted
Dec 31, 2019
Kind
B2
Abstract

A device and method for generating text representative of lip movement is provided. One or more portions of video data are determined that include: audio with an intelligibility rating below a threshold intelligibility rating; and lips of a human face. A lip-reading algorithm is applied to the one or more portions of the video data to determine text representative of detected lip movement in the one or more portions of the video data. The text representative of the detected lip movement is stored in a memory. A transcript that includes the text representative of the detected lip movement may be generated. Captioned video data may be generated from the video data and the text representative of detected lip movement.

Claims (41)

1. A device comprising:

a controller and a memory, the controller configured to:

determine one or more portions of video data that include: audio with an intelligibility rating below a threshold intelligibility rating; and lips of a human face;

apply a lip-reading algorithm to the one or more portions of the video data to determine text representative of detected lip movement in the one or more portions of the video data;

convert the audio of the one or more portions of the video data to respective text;

combine the text representative of the detected lip movement with the respective text converted from the audio to generate combined text by: replacing words in the respective text converted from the audio that have respective intelligibility ratings below the threshold intelligibility rating with corresponding words from the text representative of the detected lip movement; and

store, in the memory, the combined text.

2. The device of claim 1 , wherein the controller is further configured to:

generate a transcript of the combined text.

3. The device of claim 1 , wherein the controller is further configured to:

select the lip-reading algorithm based on sensor data indicative of one or more of: a level of excitement of a person with whom the lips are associated; and a heart rate of the person.

4. The device of claim 1 , wherein the controller is further configured to:

select the one or more portions of the video data based on one or more of video metadata and context data indicating one or more of a time and a location of an incident that corresponds to the one or more portions of the video data.

5. The device of claim 1 , wherein the controller is further configured to:

select the one or more portions of the video data by performing video analytics on the video data.

6. The device of claim 1 , wherein the controller is further configured to:

select the one or more portions of the video data based on sensor data indicative of an incident that corresponds to the one or more portions of the video data.

7. The device of claim 1 , wherein the controller is further configured to:

select the one or more portions of the video data based on context data indicating one or more of: a severity of an incident that corresponds to the one or more portions of the video data; and a role of a person that captured the video data.

8. The device of claim 1 , wherein the controller is further configured to:

generate text captions for the video data from the combined text.

9. A method comprising:

determining, at a computing device, one or more portions of video data that include: audio with an intelligibility rating below a threshold intelligibility rating; and lips of a human face;

applying, at the computing device, a lip-reading algorithm to the one or more portions of the video data to determine text representative of detected lip movement in the one or more portions of the video data;

converting the audio of the one or more portions of the video data to respective text;

combining the text representative of the detected lip movement with the respective text converted from the audio to generate combined text by: replacing words in the respective text converted from the audio that have respective intelligibility ratings below the threshold intelligibility rating with corresponding words from the text representative of the detected lip movement; and

storing, in a memory, the combined text.

10. The method of claim 9 , further comprising:

generating a transcript of the combined text.

11. The method of claim 9 , further comprising:

selecting the lip-reading algorithm based on sensor data indicative of one or more of: a level of excitement of a person with whom the lips are associated; and a heart rate of the person.

12. The method of claim 9 , further comprising:

selecting the one or more portions of the video data based on one or more of video metadata and context data indicating one or more of a time and a location of an incident that corresponds to the one or more portions of the video data.

13. The method of claim 9 , further comprising:

selecting the one or more portions of the video data by performing video analytics on the video data.

14. The method of claim 9 , further comprising:

selecting the one or more portions of the video data based on sensor data indicative of an incident that corresponds to the one or more portions of the video data.

15. The method of claim 9 , further comprising:

selecting the one or more portions of the video data based on context data indicating one or more of: a severity of an incident that corresponds to the one or more portions of the video data; and a role of a person that captured the video data.

16. The method of claim 9 , further comprising:

generating text captions for the video data from the combined text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2018
From: VARNER, STEPHEN; LIN, WEI; EKL, RANDY L.; LAW, DANIEL A.
To: MOTOROLA SOLUTIONS, INC.
Reel/Frame 044691/0692 →
Continuity (1)
Related Publication 20190198022A1 · Jun 27, 2019