IP Library Granted Patent US 10,997,236
Granted Patent B2
US 10,997,236 · App. 15/560,779 · Granted May 4, 2021

Audio content recognition method and device

Inventors: Sang-moon Lee (Seongnam-si, KR); In-woo Hwang (Suwon-si, KR); Byeong-seob Ko (Suwon-si, KR); Ki-beom Kim (Yongin-si, KR); Young-tae Kim (Seongnam-si, KR); Anant Baijal (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F16/683G06F16/61G10L19/02G10L25/03G10L25/54G10L25/45
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,997,236
App. No.
15/560,779
Granted
May 4, 2021
Kind
B2
Abstract

An audio contents recognition method includes receiving an audio signal; obtaining audio fingerprints (AFPs) based on a spectral shape of the received audio signal; generating hash codes for the obtained audio fingerprints; transmitting a matching query between the generated hash codes and hash codes stored in a database; and receiving a contents recognition result of the audio signal in response to the transmitting, wherein the generating of the hash codes includes: determining a frame interval delta_F of an audio fingerprint to generate the hash codes among the obtained audio fingerprints.

Claims (46)

1. An audio contents recognition method comprising:

receiving an audio signal;

obtaining audio fingerprints (AFPs) based on a spectral shape of the received audio signal;

generating hash codes for the obtained audio fingerprints;

transmitting a matching query about a match between the generated hash codes and hash codes stored in a database; and

receiving a contents recognition result of the audio signal in response to the transmitting,

wherein the generating of the hash codes comprises determining a frame interval of the obtained audio fingerprints based on a discrete cosine transform (DCT) coefficient difference change of reference frames among a plurality of frames of the obtained audio fingerprints and generating the hash codes for the obtained audio fingerprints based on the determined frame interval.

2. The audio contents recognition method of claim 1 , wherein the audio fingerprints are determined based on a frequency domain spectral shape of the received audio signal.

3. The audio contents recognition method of claim 2 , wherein the frame interval is generated based on a spectral size difference between adjacent frames of the obtained audio fingerprints.

4. The audio contents recognition method of claim 1 , wherein the generating of the hash codes comprises applying a weight determined based on frequency domain energy of the obtained audio fingerprints.

5. The audio contents recognition method of claim 1 , wherein the transmitting of the matching query comprises determining hash codes to transmit a matching query and a transmission priority of the hash codes to transmit the matching query among the generated hash codes based on a number of bit variations between hash codes corresponding to frames adjacent to each other.

6. The audio contents recognition method of claim 1 , wherein the contents recognition result is determined based on contents identifications (IDs) of the hash codes that transmitted the matching query and a frame concentration measure (FCM) of a frame domain.

7. The audio contents recognition method of claim 1 , wherein the audio signal comprises at least one of channel audio and object audio.

8. The audio contents recognition method of claim 1 , further comprising:

analyzing an audio scene feature of the received audio signal; and

setting a section to obtain an audio fingerprint based on the audio scene feature,

wherein the obtaining of the audio fingerprints comprises obtaining an audio fingerprint for the section.

9. The audio contents recognition method of claim 1 , further comprising:

receiving an audio contents recognition command and a matching query transmission command,

wherein the obtaining of the audio fingerprints comprises obtaining an audio fingerprint for a section from a time when the audio contents recognition command is received to a time when the matching query transmission command is received.

10. The audio contents recognition method of claim 1 , wherein the generating of the hash codes comprises, if audio fingerprints having the same value are present among the obtained audio fingerprints, deleting the audio fingerprints having the same value except for one.

11. An audio contents recognition method comprising:

receiving an audio signal;

obtaining audio fingerprints (AFPs) of the received audio signal;

generating hash codes for the obtained audio fingerprints;

matching the generated hash codes and hash codes stored in a database; and

recognizing contents of the audio signal based on a result of the matching,

wherein the generating of the hash codes comprises determining a frame interval of the obtained audio fingerprints based on a discrete cosine transform (DCT) coefficient difference change of reference frames among a plurality of frames of the obtained audio fingerprints and generating the hash codes for the obtained audio fingerprints based on the determined frame interval.

12. An audio contents recognition device comprising:

a transceiver configured to receive an audio signal;

a processor; and

a memory storing instructions executable by the processor,

wherein the processor is configured to:

obtain audio fingerprints (AFPs) of the received audio signal,

generate hash codes for the obtained audio fingerprints, transmit a matching query about a match between the generated hash codes and hash codes stored in a database, and receive a contents recognition result of the audio signal in response to the transmitting,

wherein the processor is further configured to determine a frame interval of the obtained audio fingerprints based on a discrete cosine transform (DCT) coefficient difference change of reference frames among a plurality of frames of the obtained audio fingerprints and to generate the hash codes for the obtained audio fingerprints based on the determined frame interval.

13. An audio contents recognition device comprising:

a transceiver configured to receive an audio signal;

a processor; and

a memory storing instructions executable by the processor,

wherein the processor is configured to:

obtain audio fingerprints (AFPs) of the received audio signal,

generate hash codes for the obtained audio fingerprints, and

match the generated hash codes and hash codes stored in a database and recognize contents of the audio signal based on a result of the matching,

wherein the processor is further configured to determine a frame interval of the obtained audio fingerprints based on a discrete cosine transform (DCT) coefficient difference change of reference frames among a plurality of frames of the obtained audio fingerprints and to generate the hash codes for the obtained audio fingerprints based on the determined frame interval.

14. A computer-readable recording medium having recorded thereon a computer program for implementing the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2017
From: LEE, SANG-MOON; HWANG, IN-WOO; KO, BYEONG-SEOB; KIM, KI-BEOM; KIM, YOUNG-TAE; BAIJAL, ANANT
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 043665/0724 →
Continuity (2)
Provisional Application 62153102 · Apr 27, 2015
Related Publication 20180060428A1 · Mar 1, 2018