IP Library › Granted Patent US 11,546,692
Granted Patent B1
US 11,546,692 · App. 17/370,679 · Granted Jan 3, 2023

Audio renderer based on audiovisual information

Inventors: Symeon Delikaris Manias (Los Angeles, CA); Mehrez Souden (Los Angeles, CA); Ante Jukic (Culver City, CA); Matthew S. Connolly (San Jose, CA); Sabine Webel (San Francisco, CA); Ronald J. Guglielmone, Jr. (Redwood City, CA)
Assignee: APPLE INC.
H04R3/005H04R5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,546,692
App. No.
17/370,679
Granted
Jan 3, 2023
Kind
B1
Abstract

An audio renderer can have a machine learning model that jointly processes audio and visual information of an audiovisual recording. The audio renderer can generate output audio channels. Sounds captured in the audiovisual recording and present in the output audio channels are spatially mapped based on the joint processing of the audio and visual information by the machine learning model. Other aspects are described.

Claims (29)

1. A method comprising:

obtaining one or more microphone signals, a depth signal or a movement signal, and one or more video signals or features extracted from the depth signal or the movement signal;

processing the one or more microphone signals, at least one of: the depth signal or the movement signal or the features that are extracted from the depth signal or the movement signal, and the one or more video signals jointly with a machine learning model; and

generating, with the machine learning model, a plurality of output audio channels having one or more sounds that are represented in the one or more microphone signals that are spatially mapped to a target scene based on correlations between the one or more sounds that are represented in the one or more microphone signals, visual information represented in the one or more video signals, and the depth signal or the movement signal.

2. The method of claim 1 , wherein the output audio channels are associated with a target output audio format that is one of: a binaural audio format comprising a left audio channel and a right audio channel, a channel-based loudspeaker format, and a spherical surround sound format.

3. The method of claim 1 , wherein the target scene is visually the same as a recorded scene represented by the one or more video signals.

4. The method of claim 1 , wherein the target scene is different from a recorded scene represented by the one or more video signals, and the target scene includes one or more virtual representations of the one or more sounds.

5. The method of claim 1 , wherein the target scene is defined in metadata that is provided to the machine learning model, and generating the plurality of output audio channels includes mapping the one or more sounds that are represented in the one or more microphone signals to the target scene that is defined in the metadata.

6. The method of claim 1 , wherein the movement signal is obtained from an inertial measurement unit (IMU), an accelerometer, or a gyroscope.

7. The method of claim 1 , wherein the movement signal includes translational movement or rotational movement.

8. The method of claim 1 , wherein the machine learning model includes an object detection algorithm that is applied to at least one of: the one or more video signals or the depth signal, to determine a location and orientation of one or more sources of the one or more sounds, and the plurality of output audio channels are spatially mapped to the target scene based on the location and the orientation of the one or more sources.

9. A method comprising

obtaining one or more microphone signals, a depth signal or a movement signal, and one or more video signals or features extracted from the depth signal or the movement signal;

processing the one or more microphone signals, at least one of: the depth signal or the movement signal or features that are extracted from the depth signal or the movement signal, and the one or more video signals jointly with a machine learning model; and

generating, with the machine learning model, mapping parameters that, when applied to the one or more microphone signals or a separate audio source, generate a plurality of output audio channels that contain one or more sounds that are represented in the one or more microphone signals, wherein the one or more sounds are spatially mapped to a target scene in the plurality of output audio channels and the mapping parameters are generated based on correlations between the one or more microphone signals, the movement signal or the depth signal, and the one or more video signals.

10. The method of claim 9 , wherein the mapping parameters include at least one of: beamforming filters, direction of arrival estimation, diffuseness, inter-channel level difference, inter-channel time difference, direct-to-diffuse ratio, sound field energy, reverberation time, and frequency responses associated with each of the plurality of output audio channels.

11. The method of claim 9 , wherein the plurality of output audio channels are associated with a target output audio format that is one of: a binaural audio format comprising a left audio channel and a right audio channel, a channel-based loudspeaker format, and a spherical surround sound format.

12. The method of claim 9 , wherein the target scene is visually the same as a recorded scene represented by the one or more video signals.

13. The method of claim 9 , wherein the target scene is different from a recorded scene represented by the one or more video signals and the target scene contains one or more virtual representations of the one or more sounds.

14. The method of claim 9 , wherein the target scene is defined in metadata that is provided to the machine learning model, and generating the plurality of mapping parameters is based on mapping the one or more sounds that are represented in the one or more microphone signals to the target scene that is defined in the metadata.

15. The method of claim 9 , wherein the movement signal is obtained from an inertial measurement unit (IMU), an accelerometer, or a gyroscope.

16. The method of claim 9 , wherein the movement signal includes translational movement or rotational movement.

17. The method of claim 9 , wherein the machine learning model includes an object detection algorithm that is applied to at least one of: the one or more video signals or the depth signal, to determine a location and orientation of one or more sources of the one or more sounds, and the plurality of output audio channels are spatially mapped to the target scene based on the location and the orientation of the one or more sources.

18. A method comprising

obtaining one or more features extracted from a depth signal or a movement signal, one or more visual features extracted from one or more video signals, and one or more audio features extracted from the one or more microphone signals;

processing the one or more audio features extracted from the one or more microphone signals, at least one of: the depth signal or the movement signal or features that are extracted from the depth signal or the movement signal, and the one or more visual features extracted from one or more video signals jointly with a machine learning model; and

generating, with the machine learning model, a plurality of output audio channels having one or more sounds that are represented in the one or more microphone signals that are spatially mapped to a target scene based on correlations between the features that are extracted from the depth signal or the movement signal, the one or more visual features and the one or more audio features.

19. The method of claim 18 , wherein the output audio channels are associated with a target output audio format that is one of: a binaural audio format comprising a left audio channel and a right audio channel, a channel-based loudspeaker format, and a spherical surround sound format.

20. The method of claim 18 , wherein the target wherein the target scene is visually the same as a recorded scene represented by the one or more video signals.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2021
From: DELIKARIS MANIAS, SYMEON; SOUDEN, MEHREZ; JUKIC, ANTE; CONNOLLY, MATTHEW S.; WEBEL, SABINE; GUGLIELMONE, RONALD J., JR.
To: APPLE INC.
Reel/Frame 056795/0245 →
Continuity (2)
Provisional Application 63067735 · Aug 19, 2020
Provisional Application 63151515 · Feb 19, 2021