IP Library Granted Patent US 11,915,725
Granted Patent B2
US 11,915,725 · App. 17/441,190 · Granted Feb 27, 2024

Post-processing of audio recordings

Inventor: Peter Isberg (Lund, SE)
Assignee: Sony Group Corporation
G11B27/031G10H1/366G10H2210/155G10H2240/016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,915,725
App. No.
17/441,190
Granted
Feb 27, 2024
Kind
B2
Abstract

A method of post-processing an audio recording in an audio production equipment ( 101 ) includes receiving at least one audio track ( 91 ) of the audio recording, analyzing one or more characteristics ( 80 ) of the at least one audio track ( 91 ) to identify a timing of one or more points of interest ( 251 - 254 ) of a content ( 201 - 203, 269 ) of the at least one audio track ( 91 ), and adding, to the audio recording and at the timing of the one or more points of interest ( 251 - 254 ), one or more audience reaction effects ( 261 - 264 ).

Claims (51)

1. A method of post-processing an audio recording in an audio production equipment, comprising:

receiving at least one audio track of the audio recording,

analyzing one or more characteristics of the at least one audio track to identify a timing of one or more points of interest of a content of the at least one audio track,

adding, to the audio recording and at the timing of the one or more points of interest, one or more audience reaction effects;

detecting a first intensity of an audience reaction of the content of the at least one audio track;

generating the one or more audience reaction effects based on the audience reaction; and

adding the one or more audience reaction effects having a second intensity, the second intensity being larger than the first intensity.

2. The method of claim 1 ,

wherein the one or more points of interest of the content are associated with: end of song; end of solo performance; and artist-crowd interaction.

3. The method of claim 1 , further comprising:

receiving, via a human-machine-interface, control data associated with the one or more audience reaction effects, and

adding the one or more audience reaction effects in accordance with the control data.

4. The method of claim 1 , further comprising:

loading at least a part of the one or more audience reaction effects from a database.

5. The method of claim 1 ,

wherein the one or more characteristics are selected from the group comprising: dynamics of the content of the at least one audio track; contrast in audio level of the at least one audio track; contrast in spectral distribution of the at least one audio track; contrast in musical intensity of the content of the at least one audio track; contrast in musical tempo of the content of the at least one audio track; and/or key changes of the content of the at least one audio track.

6. The method of claim 1 ,

wherein the one or more characteristics of the at least one audio track are analyzed using a machine-learning algorithm.

7. The method of claim 1 ,

wherein the post-processing is performed in real-time.

8. The method of claim 1 ,

wherein the timing of the one or more points of interest is further identified based on a user-input received via a human-machine-interface.

9. The method of claim 1 ,

wherein the one or more audience reaction effects are selected from the group comprising: cheering; whistling; stadium ambience; club ambience; and applause.

10. The method of claim 1 ,

wherein the at least one track comprises a sum of multiple audio sources.

11. A method of post-processing an audio recording in an audio production equipment, comprising:

receiving at least one audio track of the audio recording,

analyzing one or more characteristics of the at least one audio track to identify a timing of one or more points of interest of a content of the at least one audio track,

adding, to the audio recording and at the timing of the one or more points of interest, one or more audience reaction effects

performing at least one of a pitch detection and a fricative detection on vocals of the content of the at least one audio track, and

generating crowd singing in accordance with the at least one of the pitch detection and the fricative detection, to obtain the one or more audience reaction effects.

12. An audio production equipment comprising at least one processor and a memory, wherein the at least one processor is configured to load program code from the memory and to execute the program code, wherein the at least one processor is configured to perform, upon executing the program code:

receive at least one audio track of the audio recording,

analyze one or more characteristics of the at least one audio track to identify a timing of one or more points of interest of a content of the at least one audio track,

add, to the audio recording and at the timing of the one or more points of interest, one or more audience reaction effects;

detecting a first intensity of an audience reaction of the content of the at least one audio track;

generating the one or more audience reaction effects based on the audience reaction; and

adding the one or more audience reaction effects having a second intensity, the second intensity being larger than the first intensity.

13. The audio production equipment of claim 12 ,

wherein the one or more points of interest of the content are associated with: end of song; end of solo performance; and artist-crowd interaction.

14. The audio production equipment of claim 12 , the at least one processor configured to perform:

receiving, via a human-machine-interface, control data associated with the one or more audience reaction effects, and

adding the one or more audience reaction effects in accordance with the control data.

15. The audio production equipment of claim 12 , the at least one processor configured to perform:

loading at least a part of the one or more audience reaction effects from a database.

16. An audio production equipment comprising at least one processor and a memory, wherein the at least one processor is configured to load program code from the memory and to execute the program code, wherein the at least one processor is configured to perform, upon executing the program code:

receive at least one audio track of the audio recording,

analyze one or more characteristics of the at least one audio track to identify a timing of one or more points of interest of a content of the at least one audio track, add, to the audio recording and at the timing of the one or more points of interest, one or more audience reaction effects,

perform at least one of a pitch detection and a fricative detection on vocals of the content of the at least one audio track, and

generate crowd singing in accordance with the at least one of the pitch detection and the fricative detection, to obtain the one or more audience reaction effects.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2024
From: ISBERG, PETER
To: SONY CORPORATION
Reel/Frame 066064/0448 →
CHANGE OF NAME Recorded Jan 9, 2024
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 066240/0027 →
Continuity (1)
Related Publication 20220172744A1 · Jun 2, 2022