IP Library Granted Patent US 12,283,289
Granted Patent B2
US 12,283,289 · App. 17/514,694 · Granted Apr 22, 2025

Separating and rendering voice and ambience signals by offsetting impact of device movements

Inventors: Jonathan D. Sheaffer (Santa Clara, CA); Joshua D. Atkins (Los Angeles, CA); Mehrez Souden (Los Angeles, CA); Symeon Delikaris Manias (Los Angeles, CA); Sean A. Ramprashad (Los Altos, CA)
Assignee: Apple Inc.
G10L25/78G06T7/248G10L21/0272H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,289
App. No.
17/514,694
Granted
Apr 22, 2025
Kind
B2
Abstract

Processing of ambience and speech can include extracting from audio signals, ambience and speech signals. One or more spatial parameters can be generated that define spatial characteristics of ambience sound in the one or more ambience audio signals. The primary speech signal, the one or more ambience audio signals, and the spatial parameters can be encoded into one or more encoded data streams. Other aspects are described and claimed.

Claims (46)

1. A method performed by a processor of a device having a plurality of microphones, comprising:

receiving a plurality of audio signals from the plurality of microphones, the plurality of microphones capturing a sound field;

processing the audio signals into a plurality of frequency domain signals;

extracting, from the frequency domain signals, a primary speech signal;

extracting, from the frequency domain signals, one or more ambience audio signals;

generating one or more spatial parameters defining spatial characteristics of an ambience sound in the one or more ambience audio signals, the one or more spatial parameters include a location of an ambience sound source;

detecting a location or an orientation of the device, as tracking data;

modifying the one or more spatial parameters based on the tracking data by offsetting a relative movement of the ambience sound source, the relative movement caused by a change in the location or orientation of the device, to maintain a location of the ambience sound source constant during playback; and

encoding the primary speech signal, the one or more ambience audio signals, and the as modified spatial parameters into one or more encoded data streams.

2. The method of claim 1 , wherein the tracking data is generated based on one or more sensors, the sensors including one or more of the following: a camera, a set of microphones, a gyroscope, an accelerometer, and a GPS receiver.

3. The method of claim 1 , wherein the tracking data is generated based on images captured by a camera, including comparing a first image with a second image and determining a change in the location or orientation of the device based on the comparison.

4. The method of claim 1 , wherein the tracking data is generated based on estimation of sound source locations in the sound field, and detected changes in the sound source locations in the sound field, indicating a change in the location or orientation of the device.

5. The method of claim 1 , further comprising encoding the tracking data in the encoded data streams.

6. The method of claim 1 , wherein the primary speech signal is encoded without corresponding spatial parameters and is to be played back by a playback device without spatialization.

7. The method of claim 1 , wherein the primary speech signal is identified as being the primary speech signal based on a detected location of a speaker relative to the device.

8. The method of claim 1 , further comprising transmitting the encoded data streams in real-time, to a playback device, wherein the encoded data streams further includes a stream of images in sync with the primary speech signal and the ambience audio signals.

9. A method performed by a playback device, for playback of sound captured by a capture device, comprising:

receiving from the capture device one or more encoded data streams;

decoding the one or more encoded data streams to extract a primary speech signal, one or more ambience audio signals representing an ambience sound source, and one or more spatial parameters of the one or more ambience audio signals wherein the one or more spatial parameters include a location of the ambience sound source;

modifying the one or more spatial parameters of the one or more ambience audio signals based on tracking data received in and decoded from the one or more encoded data streams, the tracking data including a location or orientation of the capture device, by offsetting a relative movement of the ambience sound source caused by a change in the location or orientation of the capture device, to maintain a location of the ambience sound source constant during playback in the playback device;

determining, based on the one or more spatial parameters, one or more impulse responses;

convolving each of the one or more ambience audio signals with the one or more impulse responses resulting in spatialized ambience audio signals;

processing the spatialized ambience audio signals and the primary speech signal to produce a plurality of time domain channel signals; and

driving a plurality of speakers based on the plurality of time domain channel signals.

10. The method of claim 9 , further comprising defining or modifying a playback level of the one or more ambience audio signals based on a user input.

11. The method of claim 10 , wherein the user input is received through a graphical user interface of the playback device.

12. The method of claim 9 , wherein the tracking data is generated based on one or more sensors of the capture device, the sensors including one or more of the following: a camera, a set of microphones, a gyroscope, an accelerometer, and a GPS receiver.

13. The method of claim 9 , further comprising defining or modifying a playback level of the one or more ambience audio signals based on a) a speech to noise ratio, b) a content type, or c) a detected noise in a playback environment.

14. The method of claim 9 , wherein the primary speech signal is played directly through the plurality of speakers without spatialization.

15. An article of manufacture comprising: a non-transitory machine readable medium having stored therein instructions that, when executed by a processor of an audio capture device, cause the audio capture device to perform the following:

receiving a plurality of audio signals from a plurality of microphones that capture a sound field;

processing the audio signals into a plurality of frequency domain signals;

extracting, from the frequency domain signals, a primary speech signal;

extracting, from the frequency domain signals, one or more ambience audio signals;

generating one or more spatial parameters defining spatial characteristics of ambience in the one or more ambience audio signals, wherein the one or more spatial parameters include a location of an ambience sound source;

detecting a location or an orientation of the audio capture device, as tracking data;

modifying the one or more spatial parameters based on the tracking data by offsetting a relative movement of the ambience sound source caused by a change in the location or orientation of the audio capture device, to maintain a location of the ambience sound source constant during playback; and

encoding the primary speech signal, the ambience audio signals, and the spatial parameters into one or more encoded data streams.

16. An article of manufacture comprising: a non-transitory machine readable medium having stored therein instructions that, when executed by a processor of a playback device, cause the playback device to perform the following:

receiving from a capture device one or more encoded data streams;

decoding the one or more encoded data streams to extract a primary speech signal, one or more ambience audio signals representing an ambience sound source, and one or more spatial parameters that describe a spatial characteristic of the one or more ambience audio signals including a location of the ambience sound source;

modifying the spatial parameters of the one or more ambience audio signals based on tracking data received in and decoded from the encoded data streams, the tracking data including a location or orientation of the capture device, by offsetting a relative movement of the ambience sound source caused by a change in the location or orientation of the capture device, to maintain a location of the ambience sound source constant during playback;

determining, based on the one or more spatial parameters, one or more impulse responses;

convolving each of the one or more ambience audio signals with the one or more impulse responses resulting in spatialized ambience audio signals;

processing the spatialized ambience audio signals and the primary speech signal to produce a plurality of time domain channel signals; and

driving a plurality of speakers based on the plurality of time domain channel signals.

Continuity (3)
Continuation PCTUS2020032273 · May 9, 2020
Provisional Application 62848368 · May 15, 2019
Related Publication 20220059123A1 · Feb 24, 2022
References Cited (27)
US 9445174B2 · Virolainen · 2016 [cited by applicant]
US 9712936B2 · Peters · 2017 [cited by applicant]
US 9933989B2 · Tsingos et al. · 2018 [cited by applicant]
US 20100241438A1 · Oh et al. · 2010 [cited by applicant]
US 20110178798A1 · Flaks et al. · 2011 [cited by applicant]
US 20130272527A1 · Oomen et al. · 2013 [cited by applicant]
US 20140247945A1 · Ramo et al. · 2014 [cited by applicant]
US 20150016641A1 · Ugur et al. · 2015 [cited by applicant]
US 20150223002A1 · Mehta et al. · 2015 [cited by applicant]
US 20150350804A1 · Crockett et al. · 2015 [cited by applicant]
US 20150373475A1 · Raghuvanshi et al. · 2015 [cited by applicant]
US 20160134988A1 · Gorzel et al. · 2016 [cited by applicant]
US 20160142854A1 · Fueg et al. · 2016 [cited by applicant]
US 20180005642A1 · Wang · 2018 [cited by applicant]
US 20180091920A1 · Family · 2018 [cited by applicant]
US 20180295463A1 · Eronen et al. · 2018 [cited by applicant]
DK 3477964A1 · 2018 [cited by examiner]
DK 3477964B1 · 2018 [cited by examiner]
EP 2360681 · 2011 [cited by applicant]
JP 2010541449A · 2008 [cited by examiner]
Luoting Fu, “Sound Rendering Technology and its Application in Virtual Environment”, Computer Science, Engineering, Published 1999, 4 pages. [cited by applicant]
Liu Hao, “Real-Time Multimedia Communication System Based on WebRTC Technology”, Nanjing University of Science & Technology, 1994-2023, 68 pages. [cited by applicant]
Delikaris-Manias et al., “Parametric spatial audio recording and reproduction based on playback-setup-defined beamformers”, received from https://www.researchgate.net/publication/308787346_Parametric_spatial_audio_recor… [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2020/032273 mailed Nov. 25, 2021, 10 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2020/032273 mailed Aug. 11, 2020, 30 pages. [cited by applicant]
“Spatial ambience and sound localization programs,” Eastman Computer Music Center User's Guide, Section 7, Nov. 2012, 28 pages. [cited by applicant]
He, Jianjun, “Spatial Audio Reproduction Using Primary Ambient Extraction,” Dissertation submitted to the Nanyang Technological University, School of Electrical & Electronic Engineering, 2016, 247 pages. [cited by applicant]