IP Library › Granted Patent US 12,073,844
Granted Patent B2
US 12,073,844 · App. 17/601,042 · Granted Aug 27, 2024

Audio-visual hearing aid

Inventors: Anatoly Efros (Rishon LeZion, IL); Noam Etzion-Rosenberg (Binyamina, IL); Tal Remez (Jerusalem, IL); Oran Lang (Givatayim, IL); Inbar Mosseri (Raanana, IL); Israel Or Weinstein (Tel Aviv, IL); Benjamin Schlesinger (Ramat Hasharon, IL); Michael Rubinstein (Natick, MA); Ariel Ephrat (Efrat, IL); Yukun Zhu (Shoreline, WA); Stella Laurenzo (Seattle, WA); Amit Pitaru (Brooklyn, NY); Yossi Matias (Tel Aviv, IL)
Assignee: Google LLC
G10L21/0208G10L17/00G10L21/0272G10L25/57G10L2021/02087
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,073,844
App. No.
17/601,042
Granted
Aug 27, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for audio-visual speech separation. A method includes: receiving, by a user device, a first indication of one or more first speakers visible in a current view recorded by a camera of the user device, in response, generating a respective isolated speech signal for each of the one or more first speakers that isolates speech of the first speaker in the current view and sending the isolated speech signals for each of the one or more first speakers to a listening device operatively coupled to the user device, receiving, by the user device, a second indication of one or more second speakers visible in the current view recorded by the camera of the user device, and in response generating and sending a respective isolated speech signal for each of the one or more second speakers to the listening device.

Claims (53)

1. A method comprising:

receiving, by a user device, a first indication of one or more first speakers visible in a current view recorded by a camera of the user device;

in response to receiving the first indication, generating a respective isolated speech signal for each of the one or more first speakers that isolates speech of each of the one or more first speakers in the current view and sending the isolated speech signals for each of the one or more first speakers to a listening device operatively coupled to the user device, wherein sending the isolated speech signals for each of the one or more first speakers to the listening device comprises, for each first speaker of the one or more first speakers:

identifying a respective location of the first speaker relative to a location of the listening device that is configured to receive audio input from a plurality of audio channels; and

sending an isolated speech signal to a respective audio channel of the plurality of audio channels in accordance with the respective location of the first speaker corresponding to the isolated speech signal;

while generating the respective isolated speech signal for each of the one or more first speakers, receiving, by the user device, a second indication of one or more second speakers visible in the current view recorded by the camera of the user device; and

in response to the second indication, generating and sending a respective isolated speech signal for each of the one or more second speakers to the listening device.

2. The method of claim 1 , further comprising:

for each of one or more of the first speakers, processing a respective isolated speech signal for the speaker to generate a transcription of the speech of the speaker; and

displaying the transcription while sending the isolated speech signal of the first speaker.

3. The method of claim 1 , wherein the one or more first speakers indicated are speakers at or near the center of the current view recorded by the camera.

4. The method of claim 1 , wherein the generating and the sending of the isolated speech signals of the one or more first speakers comprises generating and sending an isolated speech signal of a first speaker of the one or more first speakers only while the first speaker is visible in the current view recorded by the camera.

5. The method of claim 1 , wherein the method further comprises receiving an indication of a preferred speaker of the one or more first speakers, and whenever generating and sending isolated speech signals for more than one first speaker, generating and sending an isolated speech signal for the preferred speaker at the exclusion of the other first speakers.

6. The method of claim 5 , wherein receiving the indication of the preferred speaker comprises receiving, at the user device, a user input selecting the preferred speaker.

7. The method of claim 1 ,

wherein receiving the first indication comprises receiving, at the user device, a first user input indicating the one or more first speakers; and

wherein receiving the second indication comprises receiving, at the user device, a second user input indicating the one or more second speakers.

8. The method of claim 7 , wherein the first user input and/or the second user input is a user selection received via a display operatively coupled to the user device.

9. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

receiving, by a user device, a first indication of one or more first speakers visible in a current view recorded by a camera of the user device;

in response to receiving the first indication, generating a respective isolated speech signal for each of the one or more first speakers that isolates speech of each of the one or more first speakers in the current view and sending the isolated speech signals for each of the one or more first speakers to a listening device operatively coupled to the user device, wherein sending the isolated speech signals for each of the one or more first speakers to the listening device comprises, for each first speaker of the one or more first speakers:

identifying a respective location of the first speaker relative to a location of the listening device that is configured to receive audio input from a plurality of audio channels; and

sending an isolated speech signal to a respective audio channel of the plurality of audio channels in accordance with the respective location of the first speaker corresponding to the isolated speech signal;

while generating the respective isolated speech signal for each of the one or more first speakers, receiving, by the user device, a second indication of one or more second speakers visible in the current view recorded by the camera of the user device; and

in response to the second indication, generating and sending a respective isolated speech signal for each of the one or more second speakers to the listening device.

10. The one or more non-transitory computer-readable storage media of claim 9 , further comprising:

for each of one or more of the first speakers, processing a respective isolated speech signal for the speaker to generate a transcription of the speech of the speaker; and

displaying the transcription while sending the isolated speech signal of the first speaker.

11. The one or more non-transitory computer-readable storage media of claim 9 , wherein the operations further comprise receiving an indication of a preferred speaker of the one or more first speakers, and whenever generating and sending isolated speech signals for more than one first speaker, generating and sending an isolated speech signal for the preferred speaker at the exclusion of the other first speakers.

12. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations, the operations comprising:

receiving, by a user device, a first indication of one or more first speakers visible in a current view recorded by a camera of the user device;

in response to receiving the first indication, generating a respective isolated speech signal for each of the one or more first speakers that isolates speech of each of the one or more first speakers in the current view and sending the isolated speech signals for each of the one or more first speakers to a listening device operatively coupled to the user device, wherein sending the isolated speech signals for each of the one or more first speakers to the listening device comprises, for each first speaker of the one or more first speakers:

identifying a respective location of the first speaker relative to a location of the listening device that is configured to receive audio input from a plurality of audio channels; and

sending an isolated speech signal to a respective audio channel of the plurality of audio channels in accordance with the respective location of the first speaker corresponding to the isolated speech signal;

while generating the respective isolated speech signal for each of the one or more first speakers, receiving, by the user device, a second indication of one or more second speakers visible in the current view recorded by the camera of the user device; and

in response to the second indication, generating and sending a respective isolated speech signal for each of the one or more second speakers to the listening device.

13. The system of claim 12 , further comprising:

for each of one or more of the first speakers, processing a respective isolated speech signal for the speaker to generate a transcription of the speech of the speaker; and

displaying the transcription while sending the isolated speech signal of the first speaker.

14. The system of claim 12 , wherein the operations further comprise receiving an indication of a preferred speaker of the one or more first speakers, and whenever generating and sending isolated speech signals for more than one first speaker, generating and sending an isolated speech signal for the preferred speaker at the exclusion of the other first speakers.

15. The method of claim 7 , wherein the first user input comprises a tactile input for selecting the one or more first speakers.

16. The method of claim 1 , further comprising:

in response to receiving the first indication, displaying a visual indicator that indicates that the one or more first speakers are selected.

17. The non-transitory computer-readable storage media of claim 9 ,

wherein receiving the first indication comprises receiving, at the user device, a first user input indicating the one or more first speakers; and

wherein receiving the second indication comprises receiving, at the user device, a second user input indicating the one or more second speakers.

18. The non-transitory computer-readable storage media of claim 9 , wherein the operations further comprise:

in response to receiving the first indication, displaying a visual indicator that indicates that the one or more first speakers are selected.

19. The system of claim 12 ,

wherein receiving the first indication comprises receiving, at the user device, a first user input indicating the one or more first speakers; and

wherein receiving the second indication comprises receiving, at the user device, a second user input indicating the one or more second speakers.

20. The system of claim 12 , wherein the operations further comprise:

in response to receiving the first indication, displaying a visual indicator that indicates that the one or more first speakers are selected.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 23, 2021
From: EFROS, ANATOLY; ETZION-ROSENBERG, NOAM; REMEZ, TAL; LANG, ORAN; MOSSERI, INBAR; WEINSTEIN, ISRAEL OR; SCHLESINGER, BENJAMIN; RUBINSTEIN, MICHAEL; EPHRAT, ARIEL; ZHU, YUKUN; LAURENZO, STELLA; PITARU, AMIT; MATIAS, YOSSI
To: GOOGLE LLC
Reel/Frame 058472/0288 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2021
From: EFROS, ANATOLY; ETZION-ROSENBERG, NOAM; REMEZ, TAL; LANG, ORAN; MOSSERI, INBAR; WEINSTEIN, ISRAEL OR; SCHLESINGER, BENJAMIN; RUBINSTEIN, MICHAEL; EPHRAT, ARIEL; ZHU, YUKUN; LAURENZO, STELLA; PITARU, AMIT; MATIAS, YOSSI
To: GOOGLE LLC
Reel/Frame 058319/0949 →
Continuity (1)
Related Publication 20230267942A1 · Aug 24, 2023