IP Library › Granted Patent US 10,192,163
Granted Patent B2
US 10,192,163 · App. 15/725,419 · Granted Jan 29, 2019

Audio processing method and apparatus based on artificial intelligence

Inventor: Zhijian Wang (Beijing, CN)
Assignee: Baidu Online Network Technology (Beijing) Co., Ltd.
G06N3/10G06F17/3074G06F17/30244G06K9/00013H04N1/00129H04N1/00204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,192,163
App. No.
15/725,419
Granted
Jan 29, 2019
Kind
B2
Abstract

The present disclosure discloses an audio processing method and apparatus based on artificial intelligence. A specific embodiment of the method comprises: converting a to-be-processed audio to a to-be-processed picture; extracting a content characteristic of the to-be-processed picture; determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture, the style characteristic being obtained from a template picture converted from a template audio; and converting the target picture to a processed audio. The present embodiment achieves the processing effect that the processed audio takes a template audio style, improves the efficiency and the flexibility of audio processing, while without changing the content of the to-be-processed audio.

Claims (69)

1. An audio processing method based on artificial intelligence, comprising:

converting a to-be-processed audio to a to-be-processed picture;

extracting a content characteristic of the to-be-processed picture;

determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture, the style characteristic being obtained from a template picture converted from a template audio; and

converting the target picture to a processed audio,

wherein the extracting a content characteristic of the to-be-processed picture comprises:

inputting the to-be-processed picture into a pre-trained convolutional neural network, the convolutional neural network being used for extracting an image characteristic; and

determining a matrix output by at least one convolutional layer in the convolutional neural network as the content characteristic of the to-be-processed picture.

2. The method according to claim 1 , wherein the converting a to-be-processed audio to a to-be-processed picture comprises:

dividing the to-be-processed audio into audio clips at a preset interval; and

determining an audiogram, a spectrum, or a spectrogram of the audio clips as the to-be-processed picture.

3. The method according to claim 1 , wherein the style characteristic is determined through the following:

inputting the template picture into a pre-trained convolutional neural network, the convolutional neural network being used for extracting an image characteristic; and

determining a matrix output by at least one convolutional layer in the convolutional neural network as the style characteristic of the template picture.

4. The method according to claim 1 , wherein the determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture comprises:

importing the content characteristic of the to-be-processed picture to a preset style transfer model, and acquiring an output of the style transfer model as the target picture.

5. The method according to claim 1 , wherein the determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture comprises:

extracting a content characteristic and a style characteristic of an initial target picture;

determining a content loss function based on the content characteristic of the to-be-processed picture and the content characteristic of the initial target picture;

determining a style loss function based on the style characteristic of the template picture and the style characteristic of the initial target picture;

determining a total loss function based on the content loss function and the style loss function; and

obtaining the target picture by adjusting the initial target picture based on the total loss function.

6. The method according to claim 5 , wherein the content loss function is obtained based on a mean square error of the content characteristic of the to-be-processed picture and the content characteristic of the initial target picture.

7. The method according to claim 5 , wherein the style loss function is determined according to the following steps:

determining a Gram matrix of the template picture and a Gram matrix of the initial target picture respectively, based on the style characteristic of the template picture and the style characteristic of the initial target picture; and

determining the style loss function based on a mean square error of the Gram matrix of the template picture and the Gram matrix of the initial target picture.

8. The method according to claim 5 , wherein the total loss function is obtained based on a weighted sum of the content loss function and the style loss function.

9. The method according to claim 5 , wherein the obtaining the target picture by adjusting the initial target picture based on the total loss function further comprises:

obtaining a minimum value of the total loss function by adjusting the initial target picture based on a gradient descent method and the total loss function; and

determining the adjusted picture corresponding to the minimum value of the total loss function as the target picture.

10. An audio processing apparatus based on artificial intelligence, comprising:

at least one processor; and

a memory storing instructions, which when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

converting a to-be-processed audio to a to-be-processed picture;

extracting a content characteristic of the to-be-processed picture;

determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture, the style characteristic being obtained from a template picture converted from a template audio; and

converting the target picture to a processed audio,

wherein the extracting a content characteristic of the to-be-processed picture comprises:

inputting the to-be-processed picture into a retrained convolutional neural network, the convolutional neural network being used for extracting an image characteristic; and

determining a matrix output by at least one convolutional layer in the convolutional neural network as the content characteristic of the to-be-processed picture.

11. The apparatus according to claim 10 , wherein the converting a to-be-processed audio to a to-be-processed picture comprises:

dividing the to-be-processed audio into audio clips at a preset interval; and

determining an audiogram, a spectrum, or a spectrogram of the audio clips as the to-be-processed picture.

12. The apparatus according to claim 10 , wherein the style characteristic is determined through the following:

inputting the template picture into a pre-trained convolutional neural network, the convolutional neural network being used for extracting an image characteristic; and determining a matrix output by at least one convolutional layer in the convolutional neural network as the style characteristic of the template picture.

13. The apparatus according to claim 10 , wherein the determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture comprises:

importing the content characteristic of the to-be-processed picture to a preset style transfer model, and acquiring an output of the style transfer model as the target picture.

14. The apparatus according to claim 10 , wherein the determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture comprises:

extracting a content characteristic and a style characteristic of an initial target picture;

determining a content loss function, based on the content characteristic of the to-be-processed picture and the content characteristic of the initial target picture;

determining a style loss function based on the style characteristic of the template picture and the style characteristic of the initial target picture;

determining a total loss function based on the content loss function and the style loss function; and

obtaining the target picture by adjusting the initial target picture based on the total loss function.

15. The apparatus according to claim 14 , wherein the content loss function is obtained based on a mean square error of the content characteristic of the to-be-processed picture and the content characteristic of the initial target picture.

16. The apparatus according to claim 14 , wherein the style loss function is determined according to the following steps:

determining a Gram matrix of the template picture and a Gram matrix of the initial target picture respectively, based on the style characteristic of the template picture and the style characteristic of the initial target picture; and

determining the style loss function based on a mean square error of the Gram matrix of the template picture and the Gram matrix of the initial target picture.

17. The apparatus according to claim 14 , wherein the total loss function is obtained based on a weighted sum of the content loss function and the style loss function.

18. The apparatus according to claim 14 , wherein the obtaining the target picture by adjusting the initial target picture based on the total loss function further comprises:

obtaining a minimum value of the total loss function by adjusting the initial target picture based on a gradient descent method and the total loss function; and

determining the adjusted picture corresponding to the minimum value of the total loss function as the target picture.

19. A non-transitory computer storage medium storing a computer program, which when executed by one or more processors, cause the one or more processors to perform operations, the operations comprising:

converting a to-be-processed audio to a to-be-processed picture;

extracting a content characteristic of the to-be-processed picture;

determining a target picture based on a style characteristic and the content characteristic of the to-be-processed picture, the style characteristic being obtained from a template picture converted from a template audio; and

converting the target picture to a processed audio,

wherein the extracting a content characteristic of the to-be-processed picture comprises;

inputting the to-be-processed picture into a pre-trained convolutional neural network, the convolutional neural network being used for extracting an image characteristic; and

determining a matrix output by at least one convolutional layer in the convolutional neural network as the content characteristic of the to-be-processed picture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2017
From: WANG, ZHIJIAN
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 043792/0954 →
Priority Claims (1)
CN 2017 1 0031469 · Jan 17, 2017 · national
Continuity (1)
Related Publication 20180204121A1 · Jul 19, 2018
Cited By (1)
US 12,361,285