IP Library › Granted Patent US 11,887,619
Granted Patent B2
US 11,887,619 · App. 17/962,722 · Granted Jan 30, 2024

Method and apparatus for detecting similarity between multimedia information, electronic device, and storage medium

Inventors: Yurong Yang (Shenzhen, CN); Xuyuan Xu (Hongkong, CN); Guoping Gong (Shenzhen, CN); Yang Fang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G10L25/18G10L21/0308
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,887,619
App. No.
17/962,722
Granted
Jan 30, 2024
Kind
B2
Abstract

A multimedia information processing method includes: parsing multimedia information to separate an audio from the multimedia information; converting the audio to obtain a mel spectrogram corresponding to the audio; determining, according to the mel spectrogram corresponding to the audio, an audio feature vector corresponding to the audio; and determining, based on an audio feature vector corresponding to a source audio in source multimedia information and an audio feature vector corresponding to a target audio in target multimedia information, a similarity between the target multimedia information and the source multimedia information.

Claims (96)

1. A multimedia information processing method, performed by an electronic device, the method comprising:

determining a first audio feature vector for source multimedia information;

determining a second audio feature vector for target multimedia information;

determining a similarity between the source multimedia information and the target multimedia information by comparing the first audio vector and the second audio vector;

wherein an audio feature vector for multimedia information is determined by:

parsing the multimedia information to separate an audio from the multimedia information;

converting the audio to obtain a mel spectrogram corresponding to the audio;

determining input triplet samples based on the mel spectrogram;

performing cross-processing on the input triplet samples through a convolutional layer and a max pooling layer of a multimedia information processing model, to obtain downsampling results of the input triplet samples;

normalizing the downsampling results through a fully connected layer of the multimedia information processing model, to obtain normalized results; and

performing deep factorization processing on the normalized results through the multimedia information processing model, to obtain the audio feature vector matching the input triplet samples.

2. The method according to claim 1 , wherein parsing the multimedia information comprises:

parsing the multimedia information to obtain timing information of the multimedia information;

parsing, according to the timing information of the multimedia information, video parameters corresponding to the multimedia information, to obtain a playing duration parameter and an audio track information parameter that are corresponding to the multimedia information; and

extracting the audio from the multimedia information based on the playing duration parameter and the audio track information parameter that are corresponding to the multimedia information.

3. The method according to claim 1 , wherein converting the audio comprises:

performing channel conversion processing on the audio to obtain mono audio data;

performing short-time Fourier transform (STFT) on the mono audio data based on a windowing function, to obtain a corresponding spectrogram; and

processing the spectrogram according to a duration parameter, to obtain the mel spectrogram corresponding to the audio.

4. The method according to claim 1 , further comprising:

obtaining a first training sample set, the first training sample set comprising audio samples in acquired video information;

performing noise addition on the first training sample set, to obtain a corresponding second training sample set;

processing the second training sample set through the multimedia information processing model, to determine initial parameters of the multimedia information processing model;

processing, in response to the initial parameters of the multimedia information processing model, the second training sample set through the multimedia information processing model, to determine updated parameters of the multimedia information processing model; and

updating, according to the updated parameters of the multimedia information processing model, network parameters of the multimedia information processing model through the second training sample set.

5. The method according to claim 4 , wherein performing the noise addition comprises:

determining a dynamic noise type matching a use environment of the multimedia information processing model; and

performing noise addition on the first training sample set according to the dynamic noise type, to change at least one of background noise, volume, sampling rate, and sound quality of audio samples in the first training sample set, so as to obtain the corresponding second training sample set.

6. The method according to claim 4 , wherein processing the second training sample set comprises:

substituting different audio samples in the second training sample set into a loss function corresponding to a triplet-loss layer network of the multimedia information processing model;

determining parameters corresponding to the triplet-loss layer network in response to determining that the loss function satisfies a corresponding convergence condition; and

determining the parameters of the triplet-loss layer network as the updated parameters of the multimedia information processing model.

7. The method according to claim 4 , wherein updating the network parameters comprises:

determining a convergence condition matching a triplet-loss layer network in the multimedia information processing model; and

updating network parameters of the triplet-loss layer network.

8. The method according to claim 1 , wherein determining the similarity comprises:

determining a corresponding inter-frame similarity parameter set based on the audio feature vector corresponding to the source audio in the source multimedia information and the audio feature vector corresponding to the target audio in the target multimedia information;

determining a quantity of audio frames reaching a similarity threshold in the inter-frame similarity parameter set; and

determining the similarity between the target multimedia information and the source multimedia information based on the quantity of audio frames reaching the similarity threshold.

9. The method according to claim 1 , further comprising:

obtaining, in response to determining that the target multimedia information is similar to the source multimedia information, copyright information of the target multimedia information and copyright information of the source multimedia information;

determining legality of the target multimedia information according to the copyright information of the target multimedia information and the copyright information of the source multimedia information; and

transmitting warning information in response to determining that the copyright information of the target multimedia information is inconsistent with the copyright information of the source multimedia information.

10. The method according to claim 1 , further comprising:

adding, in response to determining that the target multimedia information is dissimilar to the source multimedia information, the target multimedia information to a multimedia information source;

sorting a recall order of to-be-recommended multimedia information in the multimedia information source; and

recommending multimedia information to a target user based on a sorting result of the recall order of the to-be-recommended multimedia information.

11. The method according to claim 1 , further comprising:

transmitting an identifier of the multimedia information, an audio feature vector corresponding to an audio in the multimedia information, and copyright information of the multimedia information to a blockchain network, so that

a node of the blockchain network fills the identifier of the multimedia information, the audio feature vector corresponding to the audio in the multimedia information, and the copyright information of the multimedia information into a new block, and appends the new block to an end of a block chain in response to determining that a consensus on the new block is unanimous.

12. The method according to claim 11 , further comprising:

receiving a data synchronization request of another node in the blockchain network;

performing verification on a permission of the another node in response to the data synchronization request; and

controlling, in response to determining that the verification on the permission of the another node succeeds, a current node to perform data synchronization with the another node, so that the another node obtains the identifier of the multimedia information, the audio feature vector corresponding to the audio in the multimedia information, and the copyright information of the multimedia information.

13. The method according to claim 11 , further comprising:

parsing, in response to a query request, the query request to obtain a corresponding object identifier;

obtaining permission information in a target block in the blockchain network according to the object identifier;

verifying a matchability between the permission information and the object identifier;

correspondingly obtaining, in response to determining that the permission information matches the object identifier, the identifier of the multimedia information, the audio feature vector corresponding to the audio in the multimedia information, and the copyright information of the multimedia information in the blockchain network; and

transmitting the obtained identifier of the multimedia information, the audio feature vector corresponding to the audio in the multimedia information, and the copyright information of the multimedia information to a corresponding client, so that the client obtains the identifier of the multimedia information, the audio feature vector corresponding to the audio in the multimedia information, and the copyright information of the multimedia information.

14. A multimedia information processing apparatus, comprising: a memory storing computer program instructions; and a processor coupled to the memory and configured to execute the computer program instructions and perform:

determining a first audio feature vector for source multimedia information;

determining a second audio feature vector for target multimedia information;

determining a similarity between the source multimedia information and the target multimedia information by comparing the first audio vector and the second audio vector;

wherein an audio feature vector for multimedia information is determined by:

parsing the multimedia information to separate an audio from the multimedia information;

converting the audio to obtain a mel spectrogram corresponding to the audio;

determining input triplet samples based on the mel spectrogram;

performing cross-processing on the input triplet samples through a convolutional layer and a max pooling layer of a multimedia information processing model, to obtain downsampling results of the input triplet samples;

normalizing the downsampling results through a fully connected layer of the multimedia information processing model, to obtain normalized results; and

performing deep factorization processing on the normalized results through the multimedia information processing model, to obtain the audio feature vector matching the input triplet samples.

15. The multimedia information processing apparatus according to claim 14 , wherein parsing the multimedia information includes:

parsing the multimedia information to obtain timing information of the multimedia information;

parsing, according to the timing information of the multimedia information, video parameters corresponding to the multimedia information, to obtain a playing duration parameter and an audio track information parameter that are corresponding to the multimedia information; and

extracting the audio from the multimedia information based on the playing duration parameter and the audio track information parameter that are corresponding to the multimedia information.

16. The multimedia information processing apparatus according to claim 14 , wherein converting the audio includes:

performing channel conversion processing on the audio to obtain mono audio data;

performing short-time Fourier transform (STFT) on the mono audio data based on a windowing function, to obtain a corresponding spectrogram; and

processing the spectrogram according to a duration parameter, to obtain the mel spectrogram corresponding to the audio.

17. The multimedia information processing apparatus according to claim 14 , wherein the processor is further configured to execute the computer program instructions and perform:

obtaining a first training sample set, the first training sample set comprising audio samples in acquired video information;

performing noise addition on the first training sample set, to obtain a corresponding second training sample set;

processing the second training sample set through the multimedia information processing model, to determine initial parameters of the multimedia information processing model;

processing, in response to the initial parameters of the multimedia information processing model, the second training sample set through the multimedia information processing model, to determine updated parameters of the multimedia information processing model; and

updating, according to the updated parameters of the multimedia information processing model, network parameters of the multimedia information processing model through the second training sample set.

18. A non-transitory computer-readable storage medium storing computer program instructions executable by at least one processor to perform:

determining a first audio feature vector for source multimedia information;

determining a second audio feature vector for target multimedia information;

determining a similarity between the source multimedia information and the target multimedia information by comparing the first audio vector and the second audio vector;

wherein an audio feature vector for multimedia information is determined by:

parsing the multimedia information to separate an audio from the multimedia information;

converting the audio to obtain a mel spectrogram corresponding to the audio;

determining input triplet samples based on the mel spectrogram;

performing cross-processing on the input triplet samples through a convolutional layer and a max pooling layer of a multimedia information processing model, to obtain downsampling results of the input triplet samples;

normalizing the downsampling results through a fully connected layer of the multimedia information processing model, to obtain normalized results; and

performing deep factorization processing on the normalized results through the multimedia information processing model, to obtain the audio feature vector matching the input triplet samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2022
From: YANG, YURONG; XU, XUYUAN; GONG, GUOPING; FANG, YANG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 061363/0046 →
Priority Claims (1)
CN 202010956391.X · Sep 11, 2020 · national
Continuity (2)
Continuation PCTCN2021107117 · Jul 19, 2021
Related Publication 20230031846A1 · Feb 2, 2023
Cited By (1)
US 12,368,652