IP Library Granted Patent US 12,451,136
Granted Patent B2
US 12,451,136 · App. 18/036,612 · Granted Oct 21, 2025

Audio caption generation method, audio caption generation apparatus, and program

Inventors: Yuma Koizumi (Tokyo, JP); Masahiro Yasuda (Tokyo, JP)
Assignee: NTT, Inc.
G10L15/26G06F16/634
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,451,136
App. No.
18/036,612
Granted
Oct 21, 2025
Kind
B2
Abstract

Even in a case where an amount of training data is small, a caption for an audio signal is generated with high accuracy. An audio caption generation apparatus ( 1 ) generates a caption for an input target audio. A training data storage ( 10 ) stores a training data set including a set of an audio signal and a caption corresponding thereto. An audio similarity calculation unit ( 11 ) calculates similarity between the target audio and each audio signal of training data. A guidance caption retrieval unit ( 12 ) acquires a plurality of captions corresponding to an audio signal similar to the target audio. A caption generation unit ( 13 ) generates a caption for the target audio by determining words in order from the head on the basis of the acquired captions.

Claims (9)

1. An audio caption generation method of generating a caption for a target audio, the audio caption generation method comprising:

causing a guidance caption retrieval circuitry to acquire a plurality of captions corresponding to an audio signal similar to the target audio; and

causing a caption generation circuitry to determine a current word of the caption for the target audio by using a feature obtained by integrating the acquired captions, a word string from the head of the caption for the target audio to an immediately preceding word of the current word, and an acoustic feature of the target audio.

2. The audio caption generation method according to claim 1 , wherein

the guidance caption retrieval circuitry is configured such that the more similar a caption for explaining a first audio signal and a caption for explaining a second audio signal is, the more likely that first audio signal and the second audio signal are determined to be similar.

3. A non-transitory computer-readable recording medium which stores a program for causing a computer to execute each step of the audio caption generation method according to claim 1 .

4. An audio caption generation apparatus that generates a caption for a target audio, the audio caption generation apparatus comprising:

a guidance caption retrieval circuitry that acquires a plurality of captions corresponding to an audio signal similar to the target audio; and

a caption generation circuitry that generates a current word of the caption for the target audio by using a feature obtained by integrating the acquired captions, a word string from the head of the caption for the target audio to an immediately preceding word of the current word, and an acoustic feature of the target audio.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2023
From: KOIZUMI, YUMA; YASUDA, MASAHIRO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 063617/0961 →
Continuity (1)
Related Publication 20240021201A1 · Jan 18, 2024
References Cited (7)
US 10999566B1 · Mahyar · 2021 [cited by examiner]
US 20200372066A1 · Saggi · 2020 [cited by examiner]
US 20220130408A1 · Younessian · 2022 [cited by examiner]
US 20220148614A1 · Block · 2022 [cited by examiner]
Koizumi et al. (2020) “A Transformer-based Audio Captioning Model with Keyword Estimation,” Interspeech, Oct. 25-29, 2020, Shanghai, China. [cited by applicant]
Takeuchi et al. (2020) “Effects of Word-Frequency Based Pre- and Post-Processings for Audio Captioning,” Detection and Classification of Acoustic Scenes and Events (DCASE) Workshop, Nov. 2-3, 2020, Tokyo, Japan. [cited by applicant]
Kim et al. (2019) “AudioCaps: Generating Captions for Audios in the Wild,” Proceedings of NAACL-HLT 2019, pp. 119-132, Minneapolis, Minnesota, Jun. 2-Jun. 7, 2019. [cited by applicant]