IP Library Granted Patent US 9,318,129
Granted Patent B2
US 9,318,129 · App. 13/184,986 · Granted Apr 19, 2016

System and method for enhancing speech activity detection using facial feature detection

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,318,129
App. No.
13/184,986
Granted
Apr 19, 2016
Kind
B2
Abstract

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing audio. A system configured to practice the method monitors, via a processor of a computing device, an image feed of a user interacting with the computing device and identifies an audio start event in the image feed based on face detection of the user looking at the computing device or a specific region of the computing device. The image feed can be a video stream. The audio start event can be based on a head size, orientation or distance from the computing device, eye position or direction, device orientation, mouth movement, and/or other user features. Then the system initiates processing of a received audio signal based on the audio start event. The system can also identify an audio end event in the image feed and end processing of the received audio signal based on the end event.

Claims (48)

1. A method for detecting use of a mobile computing device, comprising:

identifying, a processor of a remote computing device, in an image feed from a surveillance camera:

a user; and

the mobile computing device;

identifying, the remote computing device and within the image feed, an interaction between the user and the mobile computing device;

receiving, at the remote computing device and while monitoring the image feed, an audio signal having a first voice signal associated with the user and a second voice signal associated with a non-user;

identifying, at the remote computing device, an audio start event of the first voice signal based on a distance of the user to a screen of the mobile computing device and on mouth movement of the user in the image feed; and

based on the audio start event, initiating processing of the first voice signal by the processor of the remote computing device.

2. The method of claim 1 , wherein the image feed is a video stream.

3. The method of claim 1 , wherein the image feed comprises the user looking at a specific region of the mobile computing device.

4. The method of claim 1 , wherein the audio start event is further identified based on one of head orientation, eye position, eye direction, device orientation, and other user features.

5. The method of claim 1 , further comprising:

identifying an audio end event in the image feed; and

ending processing of the first voice signal based on the audio end event.

6. The method of claim 5 , wherein identifying of the audio end event in the image feed is based on one of the user looking away from the mobile computing device, a linear model that combines events, and ending mouth movement of the user.

7. The method of claim 5 , wherein the audio end event is of a different type from the audio start event.

8. The method of claim 1 , wherein processing of the first voice signal comprises performing speech recognition of the first voice audio signal.

9. The method of claim 1 , wherein processing of the first voice signal comprises transmitting the first voice signal to a second device separate from the computing device.

10. The method of claim 1 , further comprising ignoring a portion of the audio signal that is received prior to the audio start event.

11. A system for detecting use of a mobile computing device comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

identifying, by a remote computing device, in an image feed from a surveillance camera:

a user; and

the mobile computing device;

identifying, by the remote computing device and within the image feed, an interaction between the user and the mobile computing device;

receiving, at the remote computing device and while monitoring the image feed, an audio signal having a first voice signal associated with the user and a second voice signal associated with a non-user;

identifying, at the remote computing device, an audio start event of the first voice signal based on a distance of the user to a screen of the mobile computing device and on mouth movement of the user in the image feed; and

based on the audio start event, initiating processing of the first voice signal by the remote computing device.

12. The system of claim 11 , wherein the image feed is a video stream.

13. The system of claim 11 , wherein the image feed further comprises the user looking at a specific region of the mobile computing device.

14. The system of claim 11 , wherein the audio start event is based on one of head orientation, eye position, eye direction, device orientation, and a linear model that combines events.

15. A computer-readable storage device having instructions stored for detecting use of a mobile computing device which, when executed by a computing device, cause the computing device to perform operations comprising:

identifying, a remote computing device, in an image feed from a surveillance camera:

a user; and

a mobile computing device;

identifying, the remote computing device and within the image feed, an interaction between the user and the mobile computing device;

receiving, at the remote computing device and while monitoring the image feed, an audio signal having a first voice signal associated with the user and a second voice signal associated with a non-user;

identifying, at the remote computing device, an audio start event of the first voice signal based on a distance of the user to a screen of the mobile computing device and on mouth movement of the user in the image feed; and

based on the audio start event, initiating processing of the first voice signal by the remote computing device.

16. The computer-readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, result in the operations further comprising:

identifying an audio end event in the image feed; and

ending processing of the first voice signal based on the end event.

17. The computer-readable storage device of claim 16 , wherein identifying the audio end event in the image feed is based on one of the user looking away from the mobile computing device and ending mouth movement of the user.

18. The computer-readable storage device of claim 16 , wherein the audio end event is of a different type from the audio start event.

19. The method of claim 1 , wherein the distance of the user to the screen is calculated according to a relative head size of the user.

20. The system of claim 11 , wherein the distance of the user to the screen is calculated according to a relative head size of the user.

21. The computer-readable storage device of claim 15 , wherein the distance of the user to the screen is calculated according to a relative head size of the user.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041504/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2011
From: VASILIEFF, BRANT JAMESON; EHLEN, PATRICK JOHN; LIESKE, JAY HENRY, JR.
To: AT&T INTELLECTUAL PROPERTY I, LP
Reel/Frame 026609/0961 →