IP Library Granted Patent US 12,424,237
Granted Patent B2
US 12,424,237 · App. 18/657,584 · Granted Sep 23, 2025

Tag estimation device, tag estimation method, and program

Inventors: Ryo Masumura (Tokyo, JP); Tomohiro Tanaka (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L25/30G10L15/02G10L15/10G10L25/48
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,424,237
App. No.
18/657,584
Filed
May 7, 2024
Granted
Sep 23, 2025
Kind
B2
Art Unit
2658
USPC
704/231
Abstract

A tag estimation device capable of estimating, for an utterance made among several persons, a tag representing a result of analyzing the utterance is provided. The tag estimation device includes an utterance sequence information vector generation unit that adds a t-th utterance word feature vector and a t-th speaker vector to a (t-1)-th utterance sequence information vector u t-1 that includes an utterance word feature vector that precedes the t-th utterance word feature vector and a speaker vector that precedes the t-th speaker vector to generate a t-th utterance sequence information vector u t , where t is a natural number, and a tagging unit that determines a tag l t that represents a result of analyzing a t-th utterance from a model parameter set in advance and the t-th utterance sequence information vector u t .

Claims (39)

1. A tag estimation device comprising:

a hardware processor that:

generates, by a model, based on an input, a first utterance sequence information of an utterance in a sequence of utterances in a dialogue, wherein

the input comprises a combined utterance information of the utterance and a second utterance sequence information,

the combined utterance information of the utterance comprises added pieces of utterance word feature information of the utterance and speaker information of the utterance,

the second utterance sequence information comprises recursively added pieces of utterance word feature information of respective utterances in the sequence of utterances up to an immediately preceding utterance of the utterance in the sequence of utterances and pieces of speaker information of the respective utterances in the sequence of utterances up to the immediately preceding utterance of the utterance in the sequence of utterances,

the first utterance sequence information thereby represents recursively added pieces of utterance information of utterances up to the utterance in the sequence of utterances in the dialogue,

the utterance word feature information is associated with at least a word in the utterance spoken by a speaker, and

the speaker information of the utterance is associated with the speaker; and

determines a tag associated with the utterance, wherein the tag represents a result of analyzing the utterance from a predetermined model parameter and the first utterance sequence information.

2. The tag estimation device according to claim 1 , comprising:

the hardware processor that:

transforms a word sequence of the utterance into the utterance word feature information of the utterance in the sequence of utterances; and

transforms a speaker label of the utterance into the speaker information of the utterance in the sequence of utterances.

3. The tag estimation device according to claim 1 , comprising:

the hardware processor that:

transforms a word sequence of the utterance into the utterance word feature information of the utterance; and

transforms voice of the utterance into the speaker information of the utterance.

4. The tag estimation device according to claim 1 , comprising:

the hardware processor that:

transforms voice of the utterance into the utterance word feature information of the utterance; and

transforms a speaker label of the utterance into the speaker information of the utterance.

5. The tag estimation device according to claim 1 , comprising:

the hardware processor that:

transforms voice of the utterance into the utterance word feature information of the utterance; and

transforms the voice of a frame utterance of the utterance into the speaker information of the utterance.

6. The tag estimation device according to claim 1 , wherein

the tag comprises at least one of utterance scenes, an utterance type, an emotion of an utterance utterer, or utterance paralinguistic information.

7. The tag estimation device according to claim 1 , wherein the hardware processor uses an estimation model that is learned on a basis of teacher data including a set of the sequence of utterances and a sequence of labels that are correct tags corresponding to the sequence of utterances.

8. A non-transitory computer-readable storage medium storing a program for causing a computer to function as a tag estimation device according to claim 1 .

9. A tag estimation method to be executed by a tag estimation device, comprising:

generating, by a model, based on an input, a first utterance sequence information of an utterance in a sequence of utterances in a dialogue, wherein

the input comprises a combined utterance information of the utterance and a second utterance sequence information,

the combined utterance information of the utterance comprises added pieces of utterance word feature information of the utterance and speaker information of the utterance,

the second utterance sequence information comprises recursively added pieces of utterance word feature information of respective utterances in the sequence of utterances up to an immediately preceding utterance of the utterance in the sequence of utterances and speaker information of the respective utterances in the sequence of utterances up to the utterance in the sequence of utterances,

the first utterance sequence information thereby represents recursively added pieces of utterance information of utterances up to the utterance in the sequence of utterances in the dialogue,

the utterance word feature information is associated with at least a word in the utterance spoken by a speaker, and

the speaker information of the utterance is associated with the speaker; and

determining a tag associated with the utterance, wherein the tag represents a result of analyzing the utterance from a predetermined model parameter and the first utterance sequence information, wherein the tag specifies a scene of the dialogue.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0597 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: MASUMURA, RYO; TANAKA, TOMOHIRO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 067557/0975 →
Priority Claims (1)
JP 2018-180018 · Sep 26, 2018 · national
Continuity (2)
Continuation 17279009
Related Publication 20240290344A1 · Aug 29, 2024
References Cited (16)
US 10482885B1 · Moniz · 2019 [cited by examiner]
US 20140222423A1 · Cumani · 2014 [cited by examiner]
US 20150058019A1 · Chen · 2015 [cited by examiner]
US 20170084295A1 · Tsiartas · 2017 [cited by examiner]
US 20170372694A1 · Ushio · 2017 [cited by applicant]
US 20180046710A1 · Raanani · 2018 [cited by examiner]
US 20180308487A1 · Goel · 2018 [cited by examiner]
US 20190065464A1 · Finley · 2019 [cited by examiner]
US 20190279644A1 · Yamamoto · 2019 [cited by examiner]
US 20200211567A1 · Wang · 2020 [cited by examiner]
US 20220036912A1 · Masumura et al. · 2022 [cited by applicant]
JP 2016018229A · 2016 [cited by applicant]
JP 2017228160A · 2017 [cited by applicant]
Tsunoo et al. (2017) “Hierarchical recurrent neural network for story segmentation,” In Proc. Annual Conference of the International Speech Communication Association (Interspeech), Aug. 20-24, 2017, Stockholm, Sweden, p… [cited by applicant]
India et al. (2017) “LSTM Neural Network-Based Speaker Segmentation Using Acoustic and Language Modelling” Interspeech 2017: Aug. 20-24, 2017: Stockholm. International Speech Communication Association (ISCA). [cited by applicant]
Masumura et al. (2017) “Online End-of-Turn Detection from Speech Based on Stacked Time-Asynchronous Sequential Networks” Interspeech. vol. 2017. [cited by applicant]