IP Library › Granted Patent US 12,494,222
Granted Patent B2
US 12,494,222 · App. 17/428,612 · Granted Dec 9, 2025

Sponsorship credit period identification apparatus, sponsorship credit period identification method and program

Inventors: Yasunori Oishi (Tokyo, JP); Takahito Kawanishi (Tokyo, JP); Kunio Kashino (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L25/57G06Q30/0246G06V20/46G10L15/02G10L15/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,494,222
App. No.
17/428,612
Granted
Dec 9, 2025
Kind
B2
Abstract

A credit segment identifying device includes an extracting unit. The extracting units extracts, from a first speech signal, a plurality of first partial speech signals. Each of the plurality of first partial speech signals is a part of the first speech signals and shifted from each other in time direction. An identifying unit identifies a credit segment in the first speech signal by determining whether each of the first partial speech signals includes a credit according to an association between each of second partial signals and the presence/absence of a credit. The each of second partial signals is extracted from a second speech signal.

Claims (35)

1 . A credit segment identifying device comprising:

an extractor configured to extract a plurality of first partial speech signals from a first speech signal, the first partial speech signals each being a part of the first speech signal and shifted from each other in time direction;

an identifier configured to identify a credit segment in the first speech signal by determining whether each of the first partial speech signals includes a credit according to a first set of partial speech signal comprising a previously set term and identifying a second set of speech signals not being in the same category as the first set of partial speech signals extracted from a second speech signal; wherein

identifying the credit segment further comprising:

comparing the each of the first partial speech signals to a predetermined threshold; and

generating a binary time-series signal on the each of the first partial speech signals that exceeds the predetermined threshold;

wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

2 . The credit segment identifying device according to claim 1 , wherein the identifier determines whether each of the first partial speech signals includes a credit using an identifier model which has learned each of the second partial speech signals and the presence/absence of a credit.

3 . The credit segment identifying device according to claim 2 , wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

4 . The credit segment identifying device according to claim 1 , wherein whether the second partial speech signal includes the term is determined according to speech recognition carried out to the second partial speech signal as a target.

5 . The credit segment identifying device according to claim 4 , wherein the identifier determines whether each of the first partial speech signals includes a credit using an identifier model which has learned each of the second partial speech signals and the presence/absence of a credit.

6 . The credit segment identifying device according to claim 4 , wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

7 . A method for identifying a credit segment, the method comprising:

extracting, by an extractor, from a first speech signal, a plurality of first partial speech signals each being a part of the first speech signal and shifted from each other in time direction;

identifying, by an identifier, a credit segment in the first speech signal by determining whether a credit is included in each of the partial speech signals according to a first set of partial speech signal comprising a previously set term and identifying a second set of speech signals not being in the same category as the first set of partial speech signals extracted from a second speech signal; wherein

identifying the credit segment further comprising:

comparing the each of the first partial speech signals to a predetermined threshold; and

generating a binary time-series signal on the each of the first partial speech signals that exceeds the predetermined threshold;

wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

8 . The method according to claim 7 , wherein the identifier determines whether each of the first partial speech signals includes a credit using an identifier model which has learned each of the second partial speech signals and the presence/absence of a credit.

9 . The method according to claim 8 , wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

10 . The method according to claim 7 , wherein the second partial speech signal is a speech signal including a previously set term, and whether the second partial speech signal includes the term is determined according to speech recognition carried out to the second partial speech signal as a target.

11 . The method according to claim 10 , wherein the identifier determines whether each of the first partial speech signals includes a credit using an identifier model which has learned each of the second partial speech signals and the presence/absence of a credit.

12 . The method according to claim 10 , wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

13 . A computer-readable non-transitory recording medium storing a computer-executable program instructions that when executed by a processor cause a computer system to:

extract by an extractor, from a first speech signal, a plurality of first partial speech signals each being a part of the first speech signal and shifted from each other in time direction; and

identify, by an identifier, a credit segment in the first speech signal by determining whether a credit is included in each of the partial speech signals according to a first set of partial speech signal comprising a previously set term and identifying a second set of speech signals not being in the same category as the first set of partial speech signals extracted from a second speech signal; wherein

identifying the credit segment further comprising:

comparing the each of the first partial speech signals to a predetermined threshold; and

generating a binary time-series signal on the each of the first partial speech signals that exceeds the predetermined threshold;

wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

14 . The computer-readable non-transitory recording medium of claim 13 , wherein the identifier determines whether each of the first partial speech signals includes a credit using an identifier model which has learned each of the second partial speech signals and the presence/absence of a credit.

15 . The computer-readable non-transitory recording medium of claim 13 , wherein the second partial speech signal is a speech signal including a previously set term, and whether the second partial speech signal includes the term is determined according to speech recognition carried out to the second partial speech signal as a target.

16 . The computer-readable non-transitory recording medium of claim 15 , wherein the identifier determines whether each of the first partial speech signals includes a credit using an identifier model which has learned each of the second partial speech signals and the presence/absence of a credit.

17 . The computer-readable non-transitory recording medium of claim 15 , wherein the extractor extracts a plurality of first still images corresponding to the first partial speech signals from a first video signal corresponding to the first speech signal, and the identifier identifies a credit segment in the first speech signal and the first video signal by determining whether each pair of the first partial speech signal and the first still image includes a credit according to each of the second partial speech signals and an association between a second still image extracted from a second video signal corresponding to the second speech signal and corresponding to each of the second partial speech signals and the presence/absence of a credit.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0675 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2021
From: OISHI, YASUNORI; KAWANISHI, TAKAHITO; KASHINO, KUNIO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 057084/0279 →
Priority Claims (1)
JP 2019-020322 · Feb 7, 2019 · national
Continuity (1)
Related Publication 20220115031A1 · Apr 14, 2022
References Cited (11)
US 20030028873A1 · Lemmons · 2003 [cited by examiner]
US 20040062520A1 · Gutta · 2004 [cited by examiner]
US 20090256972A1 · Ramaswamy · 2009 [cited by examiner]
US 20100306402A1 · Russell · 2010 [cited by examiner]
US 20110238495A1 · Kang · 2011 [cited by examiner]
US 20160073148A1 · Winograd · 2016 [cited by examiner]
US 20160219330A1 · Parker · 2016 [cited by examiner]
US 20170124048A1 · Campbell · 2017 [cited by examiner]
US 20180176645A1 · Reyes Sanchez · 2018 [cited by examiner]
WO 2008050718A1 · 2008 [cited by applicant]
Japan Post Production Association (2018) “CM metadata input support tool” Dec. 27, 2018 (Reading Day) [online] website: http://www.jppanet.or.jp/documents/video.html. [cited by applicant]