IP Library Granted Patent US 11,735,199
Granted Patent B2
US 11,735,199 · App. 16/648,217 · Granted Aug 22, 2023

Method for modifying a style of an audio object, and corresponding electronic device, computer readable program products and computer readable storage medium

Inventors: Quang Khanh Ngoc Duong (Cesson-Sevigne, FR); Alexey Ozerov (Cesson-Sevigne, FR); Eric Grinstein (Rio de Janeiro, BR); Patrick Perez (Rennes, FR)
Assignee: INTERDIGITAL MADISON PATENT HOLDINGS, SAS
G10L21/003G10L21/013
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,735,199
App. No.
16/648,217
Granted
Aug 22, 2023
Kind
B2
Abstract

Method for modifying a style of an audio object, and corresponding electronic device, computer readable program products and computer readable storage medium The disclosure relates to a method for processing an input audio signal. According to an embodiment, the method includes obtaining a base audio signal being a copy of the input audio signal and generating an output audio signal from the base signal, the output audio signal having style features obtained by modifying the base signal so that a distance between base style features representative of a style of the base signal and a reference style feature decreases. The disclosure also relates to corresponding electronic device, computer readable program product and computer readable storage medium.

Claims (34)

1. An electronic device comprising at least one memory and one or several processors configured for:

obtaining at least one base audio signal; and

generating at least one output audio signal from said at least one base audio signal by iteratively modifying a same temporal portion of said at least one base audio signal to gradually transform said same temporal portion of said at least one base audio signal into a corresponding temporal portion of said at least one output audio signal such that a distance between at least one base style feature representative of a base style of said at least one base audio signal and at least one reference style feature representative of a reference style decreases, wherein said same temporal portion of said at least one base audio signal is iteratively modified until said distance reaches a value and wherein said at least one base audio signal comprises an audio content other than a speech content, the audio content being iteratively modified according to the reference style to be included in the at least one output audio signal.

2. The electronic device according to claim 1 , wherein said at least one base audio signal comprises a speech content.

3. The electronic device according to claim 1 , wherein said reference style is a style of at least one reference audio signal.

4. The electronic device according to claim 3 wherein said at least one reference audio signal comprises a speech content.

5. The electronic device according to claim 3 , wherein said at least one reference audio signal comprises an audio content other than a speech content.

6. The electronic device according to claim 3 , wherein at least one of said at least one reference style feature and said at least one base style feature is obtained by processing at least one of said at least one reference audio signal and said at least one base audio signal in at least one neural network.

7. The electronic device according to claim 3 , wherein obtaining said at least one reference style feature comprises at least one of:

subband filtering of said at least one reference audio signal;

obtaining an envelope of said at least one filtered reference audio signal; and

modulating said obtained envelope.

8. The electronic device according to claim 1 , wherein obtaining said at least one base style feature comprises at least one of:

subband filtering of said at least one base audio signal;

obtaining an envelope of said at least one filtered base audio signal; and

modulating said obtained envelope.

9. A method comprising:

obtaining at least one base audio signal; and

generating at least one output audio signal from said at least one base audio signal by iteratively modifying a same temporal portion of said at least one base audio signal to gradually transform said same temporal portion of said at least one base audio signal into a corresponding temporal portion of said at least one output audio signal such that a distance between at least one base style feature representative of a base style of said at least one base audio signal and at least one reference style feature representative of a reference style decreases, wherein said same temporal portion of said at least one base audio signal is iteratively modified until said distance reaches a value and wherein said at least one base audio signal comprises an audio content other than a speech content, the audio content being iteratively modified according to the reference style to be included in the at least one output audio signal.

10. The method according to claim 9 , wherein said reference style is a style of at least one reference audio signal.

11. The method according to claim 10 , wherein said at least one reference audio signal comprises a speech content.

12. The method according to claim 10 , wherein said at least one reference audio signal comprises an audio content other than a speech content.

13. The method according to claim 10 , wherein at least one of said at least one reference style feature and said at least one base style feature is obtained by processing at least one of said at least one reference audio signal and said at least one base audio signal in at least one neural network.

14. The method according to claim 10 , wherein obtaining said at least one reference style feature comprises at least one of:

subband filtering of said at least one reference audio signal;

obtaining an envelope of said at least one filtered reference audio signal; and

modulating said obtained envelope.

15. The method according to claim 9 , wherein obtaining said at least one base style feature comprises at least one of:

subband filtering of said at least one base audio signal;

obtaining an envelope of said at least one filtered base audio signal; and

modulating said obtained envelope.

16. A non-transitory computer readable storage medium, comprising program code instructions executable by a processor, for:

obtaining at least one base audio signal; and

generating at least one output audio signal from said at least one base audio signal by iteratively modifying a same temporal portion of said at least one base audio signal to gradually transform said same temporal portion of said at least one base audio signal into a corresponding temporal portion of said at least one output audio signal such that a distance between at least one base style feature representative of a base style of said at least one base audio signal and at least one reference style feature representative of a reference style decreases, wherein said same temporal portion of said at least one base audio signal is iteratively modified until said distance reaches a value and wherein said at least one base audio signal comprises an audio content other than a speech content, the audio content being iteratively modified according to the reference style to be included in the at least one output audio signal.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2022
From: INTERDIGITAL CE PATENT HOLDINGS, SAS
To: INTERDIGITAL MADISON PATENT HOLDINGS, SAS
Reel/Frame 060310/0350 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2020
From: DUONG, QUANG KHANH NGOC; GRINSTEIN, ERIC; OZEROV, ALEXEY; PEREZ, PATRICK
To: INTERDIGITAL CE PATENT HOLDINGS
Reel/Frame 052444/0305 →
Priority Claims (1)
EP 17306202 · Sep 18, 2017 · regional
Continuity (1)
Related Publication 20200286499A1 · Sep 10, 2020
Cited By (2)
US 12,505,820 US 12,646,504