IP Library Granted Patent US 11,842,736
Granted Patent B2
US 11,842,736 · App. 18/167,653 · Granted Dec 12, 2023

Subvocalized speech recognition and command execution by machine learning

Inventors: Yaroslav Volovich (Cambridge, GB); Ant Oztaskent (London, GB); Blaise Aguera-Arcas (Seattle, WA)
Assignee: Google LLC
G10L15/22G10L15/1815H04R1/08H04R1/1016H04R1/1041G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,842,736
App. No.
18/167,653
Granted
Dec 12, 2023
Kind
B2
Abstract

Provided is an in-ear device and associated computational support system that leverages machine learning to interpret sensor data descriptive of one or more in-ear phenomena during subvocalization by the user. An electronic device can receive sensor data generated by at least one sensor at least partially positioned within an ear of a user, wherein the sensor data was generated by the at least one sensor concurrently with the user subvocalizing a subvocalized utterance. The electronic device can then process the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

Claims (72)

1. An electronic device, comprising:

one or more sensors;

one or more processors; and

one or more non-transitory computer-readable media that store instructions that when executed by the one or more processors cause the electronic device to perform operations, the operations comprising:

receiving sensor data generated by the one or more sensors concurrently with a user subvocalizing a subvocalized utterance, the subvocalized utterance being pronounced but inaudible speech caused by passing little to no air through a varying formation of a mouth of the user without excitation of the vocal cords of the user; and

processing the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

2. The electronic device of claim 1 , wherein the one or more sensors comprise one or more of:

an accelerometer;

a gyroscope;

a RADAR device:

a SONAR device;

a LASER microphone;

an infrared sensor; or

a barometer.

3. The electronic device of claim 1 , wherein at least one of the one or more sensors is at least partially positioned within an ear of the user.

4. The electronic device of claim 1 , further comprising one or more microphones located external to an ear canal of the user, wherein the operations further comprise:

receiving, with the one or more microphones located external to the ear canal of the user, audio signals external to the ear canal of the user; and

filtering, by the one or more processors, the received audio signals out of the received sensor data from the one or more sensors.

5. The electronic device of claim 4 , wherein the received audio signals external to the ear canal of the user contain an audible-speech from the user.

6. The electronic device of claim 1 , wherein the processing of the sensor data with the machine-learned subvocalization interpretation model to generate the interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model includes using an audible-speech recognition process.

7. The electronic device of claim 6 , wherein the audible speech recognition process is incorporated into the machine-learned subvocalization interpretation model.

8. The electronic device of claim 1 , wherein the output of the machine-learned subvocalization interpretation model is translated from one language into a different language.

9. The electronic device of claim 1 , wherein the operations further comprise using an input from the one or more sensors to determine if the electronic device is being worn by the user.

10. The electronic device of claim 9 , wherein the processing the sensor data with the machine-learned subvocalization interpretation model to generate the interpretation of the subvocalized utterance is respondent to the determination that the electronic device is being worn by the user.

11. The electronic device of claim 1 , the operations further comprising modifying the machine-learned subvocalization interpretation model using the received sensor data.

12. The electronic device of claim 11 , wherein the modifying comprises adjusting the weights and/or biases in the machine-learned subvocalization interpretation model.

13. A method, comprising:

generating, by one or more sensors, sensor data representative of a subvocalized utterance made by a user, the subvocalized utterance being pronounced but inaudible speech caused by passing little to no air through a varying formation of a mouth of the user without excitation of the vocal cords of the user; and

processing the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

14. The method of claim 13 , wherein the one or more sensors comprise one or more of:

an accelerometer;

a gyroscope;

a RADAR device:

a SONAR device;

a LASER microphone;

an infrared sensor; or

a barometer.

15. The method of claim 13 , wherein at least one of the one or more sensors is at least partially positioned within an ear of the user.

16. The method of claim 13 , further comprising:

receiving, by at least one or more microphones located external to an ear canal of the user, audio signals external to the ear canal of the user; and

filtering the received audio signals out of the sensor data from the one or more sensors.

17. The method of claim 16 , wherein the received audio signals external to the ear canal of the user contain an audible-speech from the user.

18. The method of claim 13 , wherein the processing the sensor data with the machine-learned subvocalization interpretation model to generate the interpretation of the subvocalized utterance as the output of the machine-learned subvocalization interpretation model includes using an audible-speech recognition process.

19. The method of claim 18 , wherein the audible-speech recognition process is incorporated into the machine-learned subvocalization interpretation model.

20. The method of claim 13 , wherein the output of the machine-learned subvocalization interpretation model is translated from one language into a different language.

21. The method of claim 13 , further comprising using an input from the one or more sensors to determine if the one or more sensors are being worn by the user.

22. The method of claim 21 , wherein the processing the sensor data with the machine-learned subvocalization interpretation model to generate the interpretation of the subvocalized utterance is respondent to the determination that the one or more sensors are being worn by the user.

23. The method of claim 13 , further comprising modifying the machine-learned subvocalization interpretation model using the received sensor data.

24. The method of claim 23 , wherein the modifying comprises adjusting the weights and/or biases in the machine-learned subvocalization interpretation model.

25. A non-transitory computer-readable medium storing instructions, which when executed by one or more processors cause the one or more processors to:

receive sensor data generated by one or more sensors concurrently with a user subvocalizing a subvocalized utterance, the subvocalized utterance being pronounced but inaudible speech caused by passing little to no air through a varying formation of a mouth of the user without excitation of the vocal cords of the user; and

process the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

26. The non-transitory computer-readable medium of claim 25 , wherein the one or more sensors comprise one or more of:

an accelerometer;

a gyroscope;

a RADAR device:

a SONAR device;

a LASER microphone;

an infrared sensor; or

a barometer.

27. The non-transitory computer-readable medium of claim 25 , wherein the one or more processors are further caused to:

receive, from one or more microphones located external to an ear canal of the user, audio signals external to the ear canal of the user; and

filter the received audio signals out of the received sensor data from the one or more sensors.

28. The non-transitory computer-readable medium of claim 27 , wherein the received audio signals external to the ear canal of the user contain an audible-speech from the user.

29. The non-transitory computer-readable medium of claim 25 , wherein the processing the sensor data with the machine-learned subvocalization interpretation model to generate the interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model includes using an audible-speech recognition process.

30. The non-transitory computer-readable medium of claim 29 , wherein the audible-speech recognition process is incorporated into the machine-learned subvocalization interpretation model.

31. The non-transitory computer-readable medium of claim 25 , wherein the one or more processors are further caused to translate the output of the machine-learned subvocalization interpretation model from one language into a different language.

32. The non-transitory computer-readable medium of claim 25 , wherein the one or more processors are further caused to use an input from the one or more sensors to determine if the one or more sensors are being worn by the user.

33. The non-transitory computer-readable medium of claim 32 , wherein the one or more processors are caused to process the sensor data with the machine-learned subvocalization interpretation model to generate the interpretation of the subvocalized utterance respondent to the determination that the one or more sensors are being worn by the user.

34. The non-transitory computer-readable medium of claim 25 , wherein the one or more processors are further caused to modify the machine-learned subvocalization interpretation model using the received sensor data.

35. The non-transitory computer-readable medium of claim 34 , wherein the modifying comprises adjusting the weights and/or biases in the machine-learned subvocalization interpretation model.

36. The non-transitory computer-readable medium of claim 25 , wherein at least one of the one or more sensors is at least partially positioned within an ear of the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2023
From: VOLOVICH, YAROSLAV; OZTASKENT, ANT; AGUERA-ARCAS, BLAISE
To: GOOGLE LLC
Reel/Frame 062751/0288 →
Continuity (3)
Continuation 17103345 · Nov 24, 2020
Provisional Application 62948989 · Dec 17, 2019
Related Publication 20230186917A1 · Jun 15, 2023
Cited By (2)
US 12,198,698 US 12,469,488