IP Library Granted Patent US 11,386,901
Granted Patent B2
US 11,386,901 · App. 16/822,095 · Granted Jul 12, 2022

Audio confirmation system, audio confirmation method, and program via speech and text comparison

Inventors: Masaomi Nishidate (Tokyo, JP); Isamu Terasaka (Tokyo, JP); Norihiro Nagai (Kanagawa, JP)
Assignee: Sony Interactive Entertainment Inc.
G10L15/26G10L15/22H04N5/278G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,386,901
App. No.
16/822,095
Granted
Jul 12, 2022
Kind
B2
Abstract

An audio confirmation system includes a voice acquiring section configured to acquire a voice contained in a motion picture; a voice text producing section configured to produce a voice text based on the acquired voice; a determining section configured to determine whether or not the produced voice text and a caption text that is embedded in an image contained in the motion picture correspond to each other; and an outputting section configured to output a result of the determination of the determining section.

Claims (36)

1. An audio confirmation system comprising:

a voice acquiring section configured to acquire a voice contained in a motion picture;

a voice text producing section configured to produce a voice text based on the acquired voice;

a determining section configured to determine whether or not the produced voice text and a caption text that is embedded in an image contained in the motion picture correspond to each other;

an outputting section configured to output a result of the determination of the determining section; and

a database containing a plurality of respective segments of voice data and a plurality of respective segments of caption text data, each respective segment of caption text data being associated with, and being determined to correspond with, a respective one of the plurality of respective segments of voice data based on the determining section,

wherein the voice text producing section is configured to produce voice text of the acquired voice by:

(i) determining a feature amount of the acquired voice by extracting voice features from the acquired voice,

(ii) comparing the feature amount of the acquired voice with one or more feature amounts obtained from the plurality of respective segments of voice data in the database,

(iii) selecting one of the respective segments of voice data in the database based on a closeness of a match between the feature amount of the acquired voice and the feature amount of the selected one of the respective segments of voice data obtained from the database, and

(iv) using the associated caption text data of the selected one of the respective segments of voice data from the database as the produced voice text of the acquired voice.

2. The audio confirmation system according to claim 1 , further comprising: a character recognizing section configured to, from the image contained in the motion picture, recognize a caption text that is embedded in the image, based on a character recognition technique.

3. The audio confirmation system according to claim 2 , wherein the determining section determines whether or not a difference between a timing when a voice in which the voice text is recognized is output, and a timing when a caption text corresponding to the voice text is displayed while being embedded in the image is within an allowable range.

4. The audio confirmation system according to claim 1 , further comprising: a motion picture acquiring section configured to acquire the motion picture that is output when a user plays a game.

5. An audio confirmation method comprising:

acquiring a voice contained in a motion picture;

producing a voice text based on the acquired voice;

determining whether or not the produced voice text and a caption text that is embedded in an image contained in the motion picture correspond to each other;

outputting a determination result obtained by the determination; and

providing a database containing a plurality of respective segments of voice data and a plurality of respective segments of caption text data, each respective segment of caption text data being associated with, and being determined to correspond with, a respective one of the plurality of respective segments of voice data based on the determining section,

wherein the producing voice text includes producing voice text of the acquired voice by:

(i) determining a feature amount of the acquired voice by extracting voice features from the acquired voice,

(ii) comparing the feature amount of the acquired voice with one or more feature amounts obtained from the plurality of respective segments of voice data in the database,

(iii) selecting one of the respective segments of voice data in the database based on a closeness of a match between the feature amount of the acquired voice and the feature amount of the selected one of the respective segments of voice data obtained from the database, and

(iv) using the associated caption text data of the selected one of the respective segments of voice data from the database as the produced voice text of the acquired voice.

6. A non-transitory, computer readable storage medium containing a computer program, which when executed by a computer, causes the computer to carry out actions, comprising:

acquiring a voice contained in a motion picture;

producing voice text based on the acquired voice;

determining whether or not the produced voice text and a caption text that is embedded in an image contained in the motion picture correspond to each other;

outputting a determination result obtained by the determination; and

providing a database containing a plurality of respective segments of voice data and a plurality of respective segments of caption text data, each respective segment of caption text data being associated with, and being determined to correspond with, a respective one of the plurality of respective segments of voice data based on the determining section,

wherein the producing voice text includes producing voice text of the acquired voice by:

(i) determining a feature amount of the acquired voice by extracting voice features from the acquired voice,

(ii) comparing the feature amount of the acquired voice with one or more feature amounts obtained from the plurality of respective segments of voice data in the database,

(iii) selecting one of the respective segments of voice data in the database based on a closeness of a match between the feature amount of the acquired voice and the feature amount of the selected one of the respective segments of voice data obtained from the database, and

(iv) using the associated caption text data of the selected one of the respective segments of voice data from the database as the produced voice text of the acquired voice.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 18, 2020
From: NISHIDATE, MASAOMI; TERASAKA, ISAMU; NAGAI, NORIHIRO
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 052149/0268 →
Continuity (2)
Continuation PCTJP2019014094 · Mar 29, 2019
Related Publication 20200312328A1 · Oct 1, 2020
Cited By (1)
US 12,494,200