IP Library Granted Patent US 11,158,320
Granted Patent B2
US 11,158,320 · App. 16/852,376 · Granted Oct 26, 2021

Methods and systems for speech detection

Inventor: Patricia Scanlon (Dublin, IE)
Assignee: Soapbox Labs Ltd.
G10L15/24G06F3/012G06F3/167G06F21/32G06K9/00228G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,158,320
App. No.
16/852,376
Granted
Oct 26, 2021
Kind
B2
Abstract

Methods and systems for processing user input to a computing system are disclosed. The computing system has access to an audio input and a visual input such as a camera. Face detection is performed on an image from the visual input, and if a face is detected this triggers the recording of audio and making the audio available to a speech processing function. Further verification steps can be combined with the face detection step for a multi-factor verification of user intent to interact with the system.

Claims (42)

1. A method of processing user input to a computing system having an audio input and a visual input, the method comprising:

receiving, at the computing system, an audio input signal from said audio input;

performing a determination of whether the user has demonstrated an intent to interact with the computing system via the audio input, wherein determining whether the user has demonstrated the intent to interact with the computing system via the audio input comprises:

performing a face detection method on an image received from the visual input after the audio input signal has been received; and

determining whether said face detection method has detected a face after the audio input signal has been received;

responsive to the determination that the user has demonstrated the intent to interact with the computing system via the audio input, confirming whether the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable by (i) performing additional verification operations comprising two or more of matching the face against a user profile of the user, determining whether the face is detected at an expected distance from a camera, or determining whether the face is detected at an expected angle with respect to the camera, and (ii) determining whether a weighted sum of results of the additional verification operations satisfies a threshold; and

responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable:

recording the audio input signal from said audio input; and

making said audio input signal available to a speech processing function.

2. The method of claim 1 , wherein said additional verification operations comprise a gaze direction detection operation to verify that the user is looking in a predefined direction or range of directions.

3. The method of claim 1 , wherein said additional verification operations comprise a mouth movement detection operation to verify that the user's mouth is moving.

4. The method of claim 3 , wherein said mouth movement detection operation further verifies that the mouth movement of the user corresponds to a movement pattern typical of speech.

5. The method of claim 1 , wherein said additional verification operations comprise an audio detection operation to verify that the audio input is receiving sound from the environment of the user.

6. The method of claim 5 , wherein said audio detection operation further verifies that the characteristics of detected sound are consistent with speech.

7. The method of claim 5 , wherein said audio detection operation further verifies that the direction from which sound is detected is consistent with the direction of the detected face.

8. The method of claim 5 , wherein said audio detection operation further verifies that the characteristics of detected sound are consistent with a speech profile stored for a given user.

9. The method of claim 1 , wherein said face detection operation verifies that the detected face is oriented in a predetermined direction or range of directions.

10. The method of claim 1 , wherein the face detection operation verifies that the visual input is at or below the level of the user's eyes or nose.

11. The method of claim 1 , further comprising performing said speech processing function on the recorded signal.

12. The method of claim 1 , further comprising sending said recorded audio signal to a remote computing device for speech processing.

13. The method of claim 1 , further comprising buffering said audio input signal, and (i) responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is not reliable, overwriting or discarding the buffered signal; and (ii) responsive to responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable, retrieving said audio input signal from the buffer.

14. The method of claim 12 , wherein said buffer is of sufficient capacity to store an audio input signal of a duration at least as long as the time required to determine the face detection and optionally the additional verification operations.

15. A computing system for processing user input having an audio input and a visual input, the system comprising:

a memory; and

a processor, coupled to the memory, to:

receive, at the computing system, an audio input signal from said audio input;

perform a determination of whether the user has demonstrated an intent to interact with the computing system via the audio input, wherein to determine whether the user has demonstrated the intent to interact with the computing system via the audio input, the processor is further to:

perform a face detection method on an image received from the visual input after the audio input signal has been received; and

determine whether said face detection method has detected a face after the audio input signal has been received;

responsive to the determination that the user has demonstrated the intent to interact with the computing system via the audio input, confirm whether the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable by (i) performing additional verification operations comprising two or more of matching the face against a user profile of the user, determining whether the face is detected at an expected distance from a camera, or determining whether the face is detected at an expected angle with respect to the camera, and (ii) determining whether a weighted sum of results of the additional verification operations satisfies a threshold; and

responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable:

record the audio input signal from said audio input; and

make said audio input signal available to a speech processing function.

16. A non-transitory computer readable medium comprising instructions, which when executed by a processor, cause the processor to perform a method of processing user input to a computing system having an audio input and a visual input, the method comprising:

receiving, at the computing system, an audio input signal from said audio input;

performing a determination of whether the user has demonstrated an intent to interact with the computing system via the audio input, wherein determining whether the user has demonstrated the intent to interact with the computing system via the audio input comprises:

performing a face detection method on an image received from the visual input after the audio input signal has been received; and

determining whether said face detection method has detected a face after the audio input signal has been received;

responsive to the determination that the user has demonstrated the intent to interact with the computing system via the audio input, confirming whether the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable by (i) performing additional verification operations comprising two or more of matching the face against a user profile of the user, determining whether the face is detected at an expected distance from a camera, or determining whether the face is detected at an expected angle with respect to the camera, and (ii) determining whether a weighted sum of results of the additional verification operations satisfies a threshold; and

responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable:

recording the audio input signal from said audio input; and

making said audio input signal available to a speech processing function.

Assignments (2)
PATENT SECURITY AGREEMENT Recorded May 30, 2024
From: SOAPBOX LABS LIMITED
To: GOLDMAN SACHS BANK USA
Reel/Frame 067577/0354 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2020
From: SCANLON, PATRICIA
To: SOAPBOX LABS LTD.
Reel/Frame 052738/0806 →
Priority Claims (1)
EP 17197186 · Oct 18, 2017 · regional
Continuity (2)
Continuation PCTEP2018078469 · Oct 18, 2018
Related Publication 20200286484A1 · Sep 10, 2020
Cited By (1)
US 12,387,619