IP Library Granted Patent US 11,699,442
Granted Patent B2
US 11,699,442 · App. 17/510,310 · Granted Jul 11, 2023

Methods and systems for speech detection

Inventor: Patricia Scanlon (Dublin, IE)
Assignee: SoapBox Labs Ltd.
G10L15/24G06F3/012G06F3/013G06F3/167G06F21/32G06V40/161G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,699,442
App. No.
17/510,310
Granted
Jul 11, 2023
Kind
B2
Abstract

Methods and systems for processing user input to a computing system are disclosed. The computing system has access to an audio input and a visual input such as a camera. Face detection is performed on an image from the visual input, and if a face is detected this triggers the recording of audio and making the audio available to a speech processing function. Further verification steps can be combined with the face detection step for a multi-factor verification of user intent to interact with the system.

Claims (47)

1. A method of processing user input to a computing system having an audio input and a visual input, the method comprising:

receiving, at the computing system, an audio input signal from the audio input;

performing a determination of whether a user has demonstrated an intent to interact with the computing system via the audio input, wherein performing the determination of whether the user has demonstrated the intent to interact with the computing system via the audio input comprises:

determining whether a face has been detected using the visual input; and

responsive to the determination that the user has demonstrated the intent to interact with the computing system via the audio input, confirming whether the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable by (i) performing additional verification operations comprising two or more of matching the face against a user profile of the user, determining whether the face is detected at an expected distance from a camera, or determining whether the face is detected at an expected angle with respect to the camera, and (ii) determining whether a weighted combination of results of the additional verification operations satisfies a threshold; and

responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable:

recording the audio input signal from the audio input.

2. The method of claim 1 , wherein determining whether the face has been detected using the visual input comprises:

performing a face detection method on an image received from the visual input after the audio input signal has been received; and

determining whether the face detection method has detected the face after the audio input signal has been received.

3. The method of claim 1 , wherein the additional verification operations further comprise a gaze direction detection operation to verify that the user is looking in a predefined direction or range of directions.

4. The method of claim 1 , wherein the additional verification operations further comprise a mouth movement detection operation to verify that the user's mouth is moving.

5. The method of claim 4 , wherein the mouth movement detection operation further verifies that the mouth movement of the user corresponds to a movement pattern typical of speech.

6. The method of claim 1 , wherein the additional verification operations comprise an audio detection operation to verify that the audio input is receiving sound from an environment of the user.

7. The method of claim 6 , wherein the audio detection operation further verifies that characteristics of detected sound are consistent with speech.

8. The method of claim 6 , wherein the audio detection operation further verifies that the direction from which sound is detected is consistent with the direction of the detected face.

9. The method of claim 6 , wherein the audio detection operation further verifies that characteristics of detected sound are consistent with a speech profile stored for a given user.

10. The method of claim 1 , further comprising:

making the recorded audio input signal available to a speech processing function.

11. The method of claim 1 , wherein the weighted combination is a weighted sum of the results of the additional verification operations.

12. The method of claim 1 , wherein determining whether the face has been detected comprises verifying whether the face is oriented in a predetermined direction or range of directions.

13. The method of claim 1 , wherein determining whether the face has been detected comprises verifying whether the visual input is at or below a level of the user's eyes or nose.

14. The method of claim 1 , further comprising sending the recorded audio input signal to a remote computing device for speech processing.

15. The method of claim 1 , further comprising:

buffering the audio input signal; and

performing one of the following:

(i) responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is not reliable, overwriting or discarding the buffered signal; or

(ii) responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable, retrieving the audio input signal from the buffer.

16. The method of claim 15 , wherein the buffer is of sufficient capacity to store an audio input signal of a duration at least as long as a time required to determine whether the face has been detected and optionally the additional verification operations.

17. A computing system for processing user input having an audio input and a visual input, the system comprising:

a memory; and

a processor, coupled to the memory, to perform a method comprising:

receiving, at the computing system, an audio input signal from the audio input;

performing a determination of whether a user has demonstrated an intent to interact with the computing system via the audio input, wherein performing the determination of whether the user has demonstrated the intent to interact with the computing system via the audio input comprises:

determining whether a face has been detected using the visual input; and

responsive to the determination that the user has demonstrated the intent to interact with the computing system via the audio input, confirming whether the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable by (i) performing additional verification operations comprising two or more of matching the face against a user profile of the user, determining whether the face is detected at an expected distance from a camera, or determining whether the face is detected at an expected angle with respect to the camera, and (ii) determining whether a weighted combination of results of the additional verification operations satisfies a threshold; and

responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable:

recording the audio input signal from the audio input.

18. A non-transitory computer readable medium comprising instructions, which when executed by a processor, cause the processor to perform a method of processing user input to a computing system having an audio input and a visual input, the method comprising:

receiving, at the computing system, an audio input signal from the audio input;

performing a determination of whether a user has demonstrated an intent to interact with the computing system via the audio input, wherein performing the determination of whether the user has demonstrated the intent to interact with the computing system via the audio input comprises:

determining whether a face has been detected using the visual input; and

responsive to the determination that the user has demonstrated the intent to interact with the computing system via the audio input, confirming whether the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable by (i) performing additional verification operations comprising two or more of matching the face against a user profile of the user, determining whether the face is detected at an expected distance from a camera, or determining whether the face is detected at an expected angle with respect to the camera, and (ii) determining whether a weighted combination of results of the additional verification operations satisfies a threshold; and

responsive to confirming that the determination that the user has demonstrated the intent to interact with the computing system via the audio input is reliable:

recording the audio input signal from the audio input.

19. The non-transitory computer readable medium of claim 18 , wherein the additional verification operations further comprise a gaze direction detection operation to verify that the user is looking in a predefined direction or range of directions.

20. The non-transitory computer readable medium of claim 18 , wherein the additional verification operations further comprise a mouth movement detection operation to verify that the user's mouth is moving, wherein the mouth movement detection operation further verifies that the mouth movement of the user corresponds to a movement pattern typical of speech.

Assignments (2)
PATENT SECURITY AGREEMENT Recorded May 30, 2024
From: SOAPBOX LABS LIMITED
To: GOLDMAN SACHS BANK USA
Reel/Frame 067577/0354 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: SCANLON, PATRICIA
To: SOAPBOX LABS LTD.
Reel/Frame 059201/0796 →
Priority Claims (1)
EP 17197186 · Oct 18, 2017 · regional
Continuity (3)
Continuation 16852376 · Apr 17, 2020
Continuation PCTEP2018078469 · Oct 18, 2018
Related Publication 20220189483A1 · Jun 16, 2022