IP Library Granted Patent US 12,375,853
Granted Patent B2
US 12,375,853 · App. 18/423,933 · Granted Jul 29, 2025

Audio encoding with compressed ambience

Inventors: Tomlinson Holman (Palm Springs, CA); Christopher T. Eubank (Santa Barbara, CA); Joshua D. Atkins (Los Angeles, CA); Soenke Pelzer (San Jose, CA); Dirk Schroeder (Sunnyvale, CA)
Assignee: Apple Inc.
H04R5/027G10L19/167G10L21/0216H04R3/005H04R3/04H04R5/033H04R5/04G10L2021/02082G10L2021/02166H04R2420/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,375,853
App. No.
18/423,933
Granted
Jul 29, 2025
Kind
B2
Abstract

An audio device can sense sound in a physical environment using a plurality of microphones to generate a plurality of microphone signals. Clean speech can be extracted from microphone signals. Ambience can be extracted from the microphone signals. The clean speech can be encoded at a first compression level. The ambience can be encoded at a second compression level that is higher than the first compression level. Other aspects are also described and claimed.

Claims (29)

1. A method performed by an audio device, comprising:

receiving, in a bit stream, a) an encoded speech signal containing speech sensed by a plurality of microphones in a physical environment, the encoded speech signal compressed to have a first bit rate, b) an encoded ambient signal containing ambient sound sensed by the plurality of microphones in the physical environment, the encoded ambient signal is compressed to have a second bit rate that is lower than the first bit rate; and c) one or more acoustic parameters of the physical environment;

decoding the encoded speech signal and the encoded ambient signal; and

applying the one or more acoustic parameters to a decoded speech signal for playback through a plurality of speakers.

2. The method of claim 1 , wherein the one or more acoustic parameters includes one or more binaural room impulse responses (BRIRs).

3. The method of claim 2 , wherein the BRIRs are applied to the decoded speech signal to spatialize the speech for playback through a left headphone speaker and a right headphone speaker of the plurality of speakers.

4. The method of claim 1 , wherein the one or more acoustic parameters includes a reverberation time or a pattern of early reflections of the physical environment.

5. The method of claim 1 , wherein applying the one or more acoustic parameters to the decoded speech signal generates a speech signal with a reverberant component for playback through the plurality of speakers.

6. The method of claim 1 , wherein the audio device includes a plurality of microphones that are integral to the audio device; the audio device being one or more of: a head-worn device, a mobile device with a display, a smart speaker, or a virtual reality headset; and the bit stream is received from another audio device through a communication protocol.

7. The method of claim 6 , wherein the audio device has a wireless transmitter, and the communication protocol is a wireless communication protocol.

8. The method of claim 1 , wherein the one or more acoustic parameters includes a reverberation decay time or a pattern of early reflections of the physical environment.

9. The method of claim 1 , wherein the one or more acoustic parameters includes one or more impulse responses of the physical environment.

10. A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations, comprising:

receiving, in a bit stream, a) an encoded speech signal containing speech sensed by a plurality of microphones in a physical environment, the encoded speech signal is compressed to have a first bit rate, b) an encoded ambient signal containing ambient sound sensed by the plurality of microphones in the physical environment, the encoded ambient signal is compressed to have a second bit rate that is lower than the first bit rate; and c) one or more acoustic parameters of the physical environment;

decoding the encoded speech signal and the encoded ambient signal; and

applying the one or more acoustic parameters to a decoded speech signal for playback through a plurality of speakers.

11. The non-transitory computer readable medium storing instructions of claim 10 , the operations further comprising:

decoding, from the bit stream, one or more spatial parameters associated with a) the ambient sound, or b) the speech, the one or more spatial parameters defining spatial locations of the ambient sound or the speech in the physical environment; and

applying the one or more spatial parameters to a decoded ambient signal or the decoded speech signal.

12. The non-transitory computer readable medium storing instructions of claim 10 , wherein a bit rate of the encoded speech signal is 96 KB/sec or greater.

13. The non-transitory computer readable medium storing instructions of claim 10 , wherein a bit rate of the encoded ambient signal is less than one tenth of a bit rate of the encoded speech signal.

14. The non-transitory computer readable medium storing instructions of claim 10 , wherein the encoded speech signal does not contain reverberant or ambient sound components.

15. The non-transitory computer readable medium storing instructions of claim 10 , the operations further comprising:

rendering a video stream onto a display of an audio device, the video stream including an avatar or real-life depiction of a speaker and the physical environment.

16. An audio device, comprising: a plurality of microphones that form a microphone array that generate a plurality of microphone signals representing sound sensed in a physical environment; and one or more processors configured to: extract clean speech from the plurality of microphone signals; extract ambience from the plurality of microphone signals; determine, based on the plurality of microphone signals, one or more acoustic parameters of the physical environment, wherein the one or more acoustic parameters include one or more of: a reverberation time, a pattern of early reflections, or one or more impulse responses of the physical environment; and encode, in a bit stream a) the clean speech by compressing the clean speech into an encoded speech signal at a first bit rate,) the ambience by compressing the ambience into an encoded ambience signal at a second bit rate that is lower than the first bit rate, and c) the one or more acoustic parameters of the physical environment, the one or more acoustic parameters encoded to be applied to the clean speech by a receiving device.

17. The audio device of claim 16 , wherein the plurality of microphones is integral to the audio device being a head-worn device, a mobile device with display, a smart speaker, or a virtual reality headset, and wherein the audio device is to transmit the bit stream to a second device through a wireless communication protocol.

18. The audio device of claim 16 , wherein the clean speech does not contain reverberant or ambient sound components.

19. The audio device of claim 16 , wherein the one or more impulse responses includes a binaural room impulse response (BRIR).

20. The audio device of claim 16 , wherein the one or more acoustic parameters are determined based on a) one or more images of the physical environment, and b) measured reverberation of the physical environment based on the plurality of microphone signals.

Continuity (4)
Continuation 17360825 · Jun 28, 2021
Continuation PCTUS2020055774 · Oct 15, 2020
Provisional Application 62927244 · Oct 29, 2019
Related Publication 20240163609A1 · May 16, 2024
References Cited (44)
US 6351733B1 · Saunders et al. · 2002 [cited by applicant]
US 9807498B1 · Harmke et al. · 2017 [cited by applicant]
US 11523244B1 · Meade et al. · 2022 [cited by applicant]
US 11930337B2 · Holman · 2024 [cited by examiner]
US 20050163323A1 · Oshikiri · 2005 [cited by applicant]
US 20050267746A1 · Jelinek et al. · 2005 [cited by applicant]
US 20080008342A1 · Sauk · 2008 [cited by examiner]
US 20080281602A1 · Van Schijndel et al. · 2008 [cited by applicant]
US 20090040289A1 · Hetherington · 2009 [cited by examiner]
US 20090094037A1 · Gao · 2009 [cited by examiner]
US 20090111507A1 · Chen · 2009 [cited by examiner]
US 20110060599A1 · Kim et al. · 2011 [cited by applicant]
US 20140086414A1 · Vilermo et al. · 2014 [cited by applicant]
US 20150356978A1 · Dickins · 2015 [cited by examiner]
US 20160337779A1 · Davidson et al. · 2016 [cited by applicant]
US 20160345116A1 · Yen et al. · 2016 [cited by applicant]
US 20170078819A1 · Habets et al. · 2017 [cited by applicant]
US 20180206038A1 · Tengelsen et al. · 2018 [cited by applicant]
US 20180232471A1 · Schissler et al. · 2018 [cited by applicant]
US 20190116448A1 · Schmidt et al. · 2019 [cited by applicant]
US 20190189144A1 · Dusan · 2019 [cited by applicant]
US 20200043468A1 · Willett · 2020 [cited by examiner]
US 20200105283A1 · Fischer · 2020 [cited by examiner]
US 20210089263A1 · Milne · 2021 [cited by examiner]
US 20240129687A1 · Neukam · 2024 [cited by examiner]
CN 1427987A · 2003 [cited by applicant]
CN 1703736A · 2005 [cited by applicant]
CN 105874820A · 2016 [cited by applicant]
CN 105900457A · 2016 [cited by applicant]
CN 106716978A · 2017 [cited by applicant]
CN 107770718A · 2018 [cited by applicant]
CN 109416585 · 2019 [cited by applicant]
CN 109564760 · 2019 [cited by applicant]
WO 2014146668 · 2014 [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2020/032274 mailed Aug. 11, 2020, 13 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2020/055774 mailed Jan. 22, 2021, 16 pages. [cited by applicant]
Hedau, Varsha, et al., “Thinking Inside the Box: Using Appearance Models and Context Based on Room Geometry,” European Conference on Computer Vision, ECCV 2010, Sep. 1, 2010, 14 pages. [cited by applicant]
Schissler, Carl, et al., “Interactive Sound Propagation and Rendering for Large Multi-Source Scenes,” ACM Trans. Graph. 36, 1, Article 2, Sep. 1, 2016, 12 pages. [cited by applicant]
Schissler, Carl, et al., “Interactive Sound Rendering on Mobile Devices using Ray-Parameterized Reverberation Filters,” Mar. 1, 2018, 20 pages. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2020/032274 mailed Nov. 25, 2021, 8 pages. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2020/055774 mailed May 12, 2022, 10 pages. [cited by applicant]
Examination Report under section 18(3) for United Kingdom Application No. 2112963.0 mailed Jun. 30, 2022, 8 pages. [cited by applicant]
Notice of Preliminary Rejection for Korean Application No. 10-2021-7031988 mailed Oct. 31, 2022, 10 pages. [cited by applicant]
Notification of the First Office Action and Search Report for Chinese Application No. 2020800194513 mailed Nov. 28, 2022, 17 pages. [cited by applicant]