IP Library Granted Patent US 11,580,978
Granted Patent B2
US 11,580,978 · App. 17/103,345 · Granted Feb 14, 2023

Machine learning for interpretation of subvocalizations

Inventors: Yaroslav Volovich (Cambridge, GB); Ant Oztaskent (London, GB); Blaise Aguera-Arcas (Seattle, WA)
Assignee: Google LLC
G10L15/22G10L15/1815H04R1/08H04R1/1016H04R1/1041G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,580,978
App. No.
17/103,345
Granted
Feb 14, 2023
Kind
B2
Abstract

Provided is an in-ear device and associated computational support system that leverages machine learning to interpret sensor data descriptive of one or more in-ear phenomena during subvocalization by the user. An electronic device can receive sensor data generated by at least one sensor at least partially positioned within an ear of a user, wherein the sensor data was generated by the at least one sensor concurrently with the user subvocalizing a subvocalized utterance. The electronic device can then process the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

Claims (42)

1. An electronic device, comprising:

one or more processors; and

one or more non-transitory computer-readable media that store instructions that when executed by the one or more processors cause the electronic device to perform operations, the operations comprising:

receiving sensor data generated by at least one sensor at least partially positioned within an ear of a user, the sensor data generated by the at least one sensor concurrently with the user subvocalizing a subvocalized utterance, the subvocalized utterance being pronounced but inaudible speech caused by passing little to no air through a varying formation of a mouth of the user without excitation of vocal cords of the user; and

processing the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

2. The electronic device of claim 1 , wherein the at least one sensor comprises one or more microphones that convert a sound wave located within an ear canal of the ear of the user to the sensor data.

3. The electronic device of claim 2 , wherein the sound wave located within the ear canal of the ear of the user is generated by an eardrum of the user.

4. The electronic device of claim 2 wherein, when the one or more microphones are placed within the ear of the user, the one or more microphones are directed toward the eardrum of the user.

5. The electronic device of claim 1 , wherein the at least one sensor comprises one or more of:

an accelerometer;

a gyroscope;

a RADAR device:

a SONAR device;

a LASER microphone;

an infrared sensor; or

a barometer.

6. The electronic device of claim 1 , wherein the electronic device is sized and shaped to be at least partially positioned within the ear of the user.

7. The electronic device of claim 1 , wherein the electronic device is coupled to an ancillary support device that is physically separate from the at least one sensor.

8. The electronic device of any of claim 1 , wherein the interpretation of the subvocalized utterance output by the machine-learned subvocalization interpretation model comprises a transcript of the subvocalized utterance in textual form.

9. The electronic device of claim 1 , wherein the interpretation of the subvocalized utterance output by the machine-learned subvocalization interpretation model comprises a classification of the subvocalized utterance into one or more of a plurality of categories.

10. The electronic device of claim 9 , wherein the plurality of categories compose a plurality of defined commands.

11. The electronic device of claim 1 , wherein the operations further comprise:

determining, by an artificial intelligence-based personal assistant system, at least one operation to perform based at least in part on the interpretation of the subvocalized utterance output by the machine-learned subvocalization interpretation model.

12. The electronic device of claim 1 , wherein the passing little to no air through the varying formation of the mouth of the user results in a nearly-inaudible whisper or “silent speech”.

13. The electronic device of claim 1 , wherein the passing little to no air through the varying formation of the mouth of the user passes no air.

14. An ear bud, comprising:

at least one sensor;

one or more processors; and

one or more non-transitory computer-readable media that store instructions that when executed by the one or more processors cause the ear bud to perform operations, the operations composing:

receiving sensor data generated by the at least one sensor, the sensor data generated by the at least one sensor concurrently with the user subvocalizing a subvocalized utterance and with the at least one sensor at least partially positioned within an ear of a user, the subvocalized utterance being pronounced but inaudible speech caused by passing little to no air through a varying formation of a mouth of the user without excitation of vocal cords of the user; and

processing the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

15. The ear bud of claim 14 , wherein the at least one sensor comprises a microphone that converts a sound wave located within an ear canal of the ear of the user to the sensor data.

16. The ear bud of claim 15 , wherein the sound wave located within the ear canal of the ear of the user is generated by an eardrum of the user.

17. The ear bud of claim 15 , wherein, when the microphone is placed within the ear of the user, the microphone is directed to the eardrum of the user.

18. The ear bud of claim 14 , wherein the interpretation of the subvocalized utterance output by the machine-learned subvocalization interpretation model comprises a transcript of the subvocalized utterance in textual form.

19. The ear bud of claim 14 , wherein the interpretation of the subvocalized utterance output by the machine-learned subvocalization interpretation model comprises a classification of the subvocalized utterance into one or more of a plurality of categories.

20. The ear bud of claim 19 , wherein the plurality of categories comprise a plurality of defined commands.

21. The ear bud of claim 14 , wherein the operations further comprise:

determining, by an artificial intelligence-based personal assistant system, at least one operation to perform based at least in part on the interpretation of the subvocalized utterance output by the machine-learned subvocalization interpretation model.

22. A method, comprising:

generating, by at least one sensor at least partially positioned within an ear of a user, sensor data representative of a subvocalized utterance made by the user, the subvocalized utterance being pronounced but inaudible speech caused by passing little to no air through a varying formation of a mouth of the user without excitation of vocal cords of the user; and

processing the sensor data with a machine-learned subvocalization interpretation model to generate an interpretation of the subvocalized utterance as an output of the machine-learned subvocalization interpretation model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2020
From: VOLOVICH, YAROSLAV; OZTASKENT, ANT; AGUERA-ARCAS, BLAISE
To: GOOGLE LLC
Reel/Frame 054467/0935 →
Continuity (2)
Provisional Application 62948989 · Dec 17, 2019
Related Publication 20210183383A1 · Jun 17, 2021
Cited By (2)
US 12,198,698 US 12,469,488