IP Library Granted Patent US 10,109,300
Granted Patent B2
US 10,109,300 · App. 15/063,928 · Granted Oct 23, 2018

System and method for enhancing speech activity detection using facial feature detection

Inventors: Brant Jameson Vasilieff (Glendale, CA); Patrick John Ehlen (San Francisco, CA); Jay Henry Lieske (Los Angeles, CA)
Assignee: NUANCE COMMUNICATIONS, INC.
G10L25/78G10L15/20G10L25/57H04N7/183H04N1/00403
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,109,300
App. No.
15/063,928
Filed
Mar 8, 2016
Granted
Oct 23, 2018
Kind
B2
Art Unit
2487
USPC
348/77
Abstract

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing audio. A system configured to practice the method monitors, via a processor of a computing device, an image feed of a user interacting with the computing device and identifies an audio start event in the image feed based on face detection of the user looking at the computing device or a specific region of the computing device. The image feed can be a video stream. The audio start event can be based on a head size, orientation or distance from the computing device, eye position or direction, device orientation, mouth movement, and/or other user features. Then the system initiates processing of a received audio signal based on the audio start event. The system can also identify an audio end event in the image feed and end processing of the received audio signal based on the end event.

Claims (34)

1. A method, comprising:

capturing video content by a camera;

analyzing, by a system including a processor, the video content, wherein the analyzing comprises detecting a user is depicted in the video content and detecting that a computing device is also depicted in the video content;

responsive to a first determination by the system that the user is looking at the computing device and a second determination by the system that a mouth of the user is moving, determining, by the system, an audio start event of a voice signal of the user, wherein the first and second determinations are based on the analyzing of the video content and the detecting of the user and the computing device; and

responsive to the audio start event, initiating processing of the voice signal.

2. The method of claim 1 , wherein the computing device is a mobile computing device.

3. The method of claim 1 , wherein the video content shows the user looking at a specific region of the computing device.

4. The method of claim 1 , wherein the determining of the audio start event is based on head orientation, eye position, eye direction, device orientation, or any combination thereof.

5. The method of claim 1 , further comprising:

determining an audio end event based on the analyzing of the video content; and

causing the computing device to cease the processing of the voice signal based on the audio end event.

6. The method of claim 5 , wherein the determining of the audio end event is based on a third determination by the system that the video content shows the user looking away from the computing device.

7. The method of claim 1 , wherein the processing of the voice signal is by the computing device.

8. The method of claim 1 , wherein the processing of the voice signal comprises performing speech recognition of the voice signal.

9. The method of claim 1 , wherein the processing of the voice signal comprises transmitting the voice signal to a second device separate from the computing device.

10. The method of claim 1 , further comprising ignoring a portion of an audio signal that is received prior to the audio start event, the audio signal including the voice signal.

11. An apparatus comprising:

a processor; and

a non-transitory computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

analyzing video content captured by a camera to determine that the video content includes a depiction of a user and that the video content includes a depiction of a computing device;

responsive to a first determination of a distance of the user to a screen of the computing device and a second determination that a mouth of the user is moving, determining an audio start event of a voice signal of the user, wherein the first and second determinations are based on the analyzing of the video content that depicts the user and the computing device; and

responsive to the audio start event, initiating processing of the voice signal.

12. The apparatus of claim 11 , wherein the computing device is a mobile computing device.

13. The apparatus of claim 11 , wherein the video content shows the user looking at a specific region of the computing device.

14. The apparatus of claim 11 , wherein the determining of the audio start event is based on head orientation, eye position, eye direction, device orientation, or any combination thereof.

15. The apparatus of claim 11 , wherein the operations further comprise ignoring a portion of an audio signal that is received prior to the audio start event, the audio signal including the voice signal.

16. A non-transitory computer-readable storage device having instructions which, when executed by a processor, cause the processor to perform operations comprising:

analyzing video content captured by a camera, the video content depicting a user in the video content and a computing device depicted in the video content;

responsive to a first determination that the user is looking at the computing device and a second determination that a mouth of the user is moving, determining an audio start event of a voice signal of the user, wherein the first and second determinations are based on the analyzing of the video content; and

responsive to the audio start event, initiating processing of the voice signal.

17. The non-transitory computer-readable storage device of claim 16 , wherein the operations further comprise:

ignoring a portion of an audio signal that is received prior to the audio start event, the audio signal including the voice signal.

18. The non-transitory computer-readable storage device of claim 16 , wherein the video content shows the user looking at a specific region of the computing device.

19. The non-transitory computer-readable storage device of claim 16 , wherein the determining of the audio start event is based on head orientation, eye position, eye direction, device orientation, or any combination thereof.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2016
From: VASILIEFF, BRANT JAMESON; EHLEN, PATRICK JOHN; LIESKE, JAY HENRY, JR.
To: AT&T INTELLECTUAL PROPERTY I, LP
Reel/Frame 038200/0573 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2016
From: VASILIEFF, BRANT JAMESON; EHLEN, PATRICK JOHN; LIESKE, JAY HENRY, JR.
To: AT&T INTELLECTUAL PROPERTY I, LP
Reel/Frame 038074/0599 →
Continuity (2)
Continuation 13184986 · Jul 18, 2011
Related Publication 20160189733A1 · Jun 30, 2016