IP Library Granted Patent US 10,721,571
Granted Patent B2
US 10,721,571 · App. 16/264,297 · Granted Jul 21, 2020

Separating and recombining audio for intelligibility and comfort

Inventors: Dwight Crow (San Francisco, CA); Shlomo Zippel (San Francisco, CA); Andrew Song (San Francisco, CA); Emmett McQuinn (San Francisco, CA); Zachary Rich (San Francisco, CA)
Assignee: WHISPER.AI, Inc.
H04R25/505H04R25/43H04R2225/43H04R2225/55
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,721,571
App. No.
16/264,297
Granted
Jul 21, 2020
Kind
B2
Abstract

Audio enhancement systems, devices, methods, and computer program products are disclosed. In particular embodiments, audio is separated by source at an auxiliary processing device, primary voice presence and/or relevancy per source is determined and used to determine enhancement data that is sent to one or more earpieces and used for enhancing audio at the earpiece. These and other embodiments are disclosed herein.

Claims (56)

1. A method of enhancing audio for at least one earpiece of a user comprising:

receiving, at an auxiliary processing device, audio data corresponding to audio sensed at the at least one earpiece;

using the received audio data to identify a plurality of respective portions of the audio data as corresponding to a plurality of respective separate audio sources;

using the received audio data to compute audio source parameters for each respective separate audio source;

using the computed audio source parameters to estimate whether one of the respective separate audio sources comprises a primary voice source to which the user is attending; and

if one of the respective separate audio sources is estimated to comprise a primary voice source, then performing processing comprising:

determining all respective separate audio sources that are not the primary voice source to be secondary noise sources;

using a preferred signal-to-noise ratio corresponding to the user and a volume of the primary voice source to determine a maximum combined noise value for the secondary noise sources;

determining and applying volume weights to each of the secondary noise sources to maintain a combined noise value for the secondary noise sources that is equal to or less than the maximum secondary noise value; and

using the volume weights to determine enhancement data to send to the at least one earpiece for enhancing audio played for the user at the at least one earpiece.

2. The method of claim 1 wherein the enhancement data comprises enhanced audio.

3. The method of claim 1 wherein the enhancement data comprises filter updates for filters to be applied by the at least one earpiece to enhance the audio sensed at the at least one earpiece.

4. The method of claim 1 wherein the volume of the primary voice source is in a preferred range corresponding to the user.

5. The method of claim 4 wherein the volume of the primary voice source is a lowest volume in the preferred range corresponding to the user.

6. The method of claim 1 further comprising:

receiving additional data including non-audio data at the auxiliary processing device; and

using the additional data in computing the audio source parameters.

7. The method of claim 6 wherein the additional data comprises data regarding head movement of the user.

8. The method of claim 6 wherein the additional data comprises data regarding a change in speech of the user.

9. The method of claim 6 wherein the additional data comprises data regarding an amount of immediately preceding time that an audio source of the respective separate audio sources has been making sound in a vicinity of the user.

10. The method of claim 6 wherein the additional data comprises data regarding a fraction of an historical time period that an audio source of the respective separate audio sources has been making sound in a vicinity of the user.

11. The method of claim 6 wherein the additional data comprises data regarding an identity of a voice source that is an audio source of the respective separate audio sources.

12. The method of claim 6 wherein the additional data comprises data regarding a recognized word or words spoken by an audio source of the respective separate audio sources.

13. A hearing aid system comprising:

at least one earpiece configured to sense audio and to transmit audio data corresponding to the sensed audio; and

an auxiliary processing device configured to receive the audio data from the at least one earpiece and to perform processing comprising:

using the received audio data to identify a plurality of respective portions of the audio data as corresponding to a plurality of respective separate audio sources;

using the received audio data to compute audio source parameters for each respective separate audio source;

using the computed audio source parameters to estimate whether one of the respective separate audio sources comprises a primary voice source to which the user is attending; and

if one of the respective separate audio sources is estimated to comprise a primary voice source, then performing processing comprising:

determining all respective separate audio sources that are not the primary voice source to be secondary noise sources;

using a preferred signal-to-noise ratio corresponding to the user and a volume of the primary voice source to determine a maximum combined noise value for the secondary noise sources;

determining and applying volume weights to each of the secondary noise sources to maintain a combined noise value for the secondary noise sources that is equal to or less than the maximum secondary noise value; and

using the volume weights to determine enhancement data to send to the at least one earpiece for enhancing audio played for the user at the at least one earpiece.

14. The hearing aid system of claim 13 wherein the auxiliary processing device is further configured to receive additional data including non-audio data and to perform processing comprising using the additional data in computing the audio source parameters.

15. The hearing aid system of claim 14 wherein the additional data comprises data regarding head movement of the user.

16. The hearing aid system of claim 14 wherein the additional data comprises data regarding an amount of immediately preceding time that an audio source of the respective separate audio sources has been making sound in a vicinity of the user.

17. The hearing aid system of claim 14 wherein the additional data comprises data regarding a fraction of an historical time period that an audio source of the respective separate audio sources has been making sound in a vicinity of the user.

18. The hearing aid system of claim 13 wherein the enhancement data comprises enhanced audio.

19. The hearing aid system of claim 13 wherein the enhancement data comprises filter updates for filters to be applied by the at least one earpiece to enhance the audio sensed at the at least one earpiece.

20. A computer program product comprising a non-transitory computer readable medium storing executable instruction code that, when executed by a processor, performs processing comprising:

using audio data corresponding to audio sensed by at least one earpiece to identify a plurality of respective portions of the audio data as corresponding to a plurality of respective separate audio sources;

using the audio data to compute audio source parameters for each respective separate audio source;

using the computed audio source parameters to estimate whether one of the respective separate audio sources comprises a primary voice source to which the user is attending; and

if one of the respective separate audio sources is estimated to comprise a primary voice source, then performing processing comprising:

determining all respective separate audio sources that are not the primary voice source to be secondary noise sources;

using a preferred signal-to-noise ratio corresponding to the user and a volume of the primary voice source to determine a maximum combined noise value for the secondary noise sources;

determining and applying volume weights to each of the secondary noise sources to maintain a combined noise value for the secondary noise sources that is equal to or less than the maximum secondary noise value; and

using the volume weights to determine enhancement data for enhancing audio played for the user at the at least one earpiece.

21. The computer program product of claim 20 wherein the volume of the primary voice source is in a preferred range corresponding to the user.

22. The computer program product of claim 20 wherein the volume of the primary voice source is a lowest volume in the preferred range corresponding to the user.

23. The computer program product of claim 20 wherein the executable instruction code, when executed by a processor, performs processing comprising using additional data including non-audio data in computing the audio source parameters.

24. The computer program product of claim 23 wherein the additional data comprises data regarding head movement of the user.

25. The computer program product of claim 23 wherein the additional data comprises data regarding change in speech of the user.

26. The computer program product of claim 23 wherein the additional data comprises data regarding an amount of immediately preceding time that an audio source of the respective separate audio sources has been making sound in a vicinity of the user.

27. The computer program product of claim 23 wherein the additional data comprises data regarding a fraction of an historical time period that an audio source of the respective separate audio sources has been making sound in a vicinity of the user.

Assignments (2)
CHANGE OF NAME Recorded Jan 17, 2024
From: WHISPER.AI, INC.
To: WHISPER.AI, LLC
Reel/Frame 066342/0669 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2019
From: CROW, DWIGHT; ZIPPEL, SHLOMO; SONG, ANDREW; MCQUINN, EMMETT; RICH, ZACHARY
To: WHISPER.AI, INC.
Reel/Frame 049407/0346 →
Continuity (3)
Continuation PCTUS2018057418 · Oct 24, 2018
Provisional Application 62576373 · Oct 24, 2017
Related Publication 20190166435A1 · May 30, 2019
Cited By (16)
US 12,196,835 US 12,231,851 US 12,323,769 US 12,356,153 US 12,356,154 US 12,356,156 US 12,363,489 US 12,395,800 US 12,418,756 US 12,457,450 US 12,574,691 US 12,598,434 US 12,610,200 US 12,616,417 US 12,634,642 US 12,713,188