IP Library Patent Application 18683786
Patent Application
App. No. 18/683,786

SPEECH SYNTHESIS APPARATUS, SPEECH SYNTHESIS METHOD, AND SPEECH SYNTHESIS PROGRAM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/683,786
Abstract

A speech synthesis apparatus according to the present disclosure includes a memory and a processor coupled to the memory. The processor is configured to: obtain utterance information on subjects to be uttered, wherein the subjects to be uttered are texts contained in data on a book, obtain image information on images that are contained in the data on the book, obtain speech data corresponding to the subjects to be uttered; and generate, based on the obtained utterance information, the obtained image information, and the obtained speech data, a speech synthesis model for reading out a text associated with an image.

Claims (30)

1 . A speech synthesis apparatus comprising:

a memory; and

a processor coupled to the memory and configured to:

obtain utterance information on subjects to be uttered, wherein the subjects to be uttered are texts contained in data on a book,

obtain image information on images that M contained in the data on the book,

obtain speech data corresponding to the subjects to be uttered; and

generate, based on the obtained utterance information, the obtained image information, and the obtained speech data, speech synthesis model for reading out a text associated with an image.

2 . The speech synthesis apparatus of claim 1 , wherein the processor configured to obtain, as the image information, information on an image that is contained in a specific page of the first book and that is associated with a text contained in the specific page.

3 . The speech synthesis apparatus of claim 1 , wherein the processor configured to obtain, as the speech data, data of speech reading out a text that is contained in a specific page of the first book and that is associated with an image contained in the specific page.

4 . The speech synthesis apparatus of claim 1 , wherein the processor configured to obtain the utterance information presenting at least one of accents, parts of speech, and a time of start of a phonome or a time of end of a phonome of each of the subjects to be uttered.

5 . The speech synthesis apparatus of claim 1 , wherein the processor further configured to:

convert the utterance information into language vectors, wherein each language vector represents linguistic information on the corresponding subject to be uttered;

convert the image information into visual feature vectors, wherein each visual feature vector represents a visual feature of the corresponding image contained in the first book; and

generate the speech synthesis model using training data containing the speech data that is associated with the language vectors and the visual feature vectors.

6 . A speech synthesis method performed by a computer, the method comprising:

obtaining utterance information on subjects to be uttered, wherein the subjects to be uttered is-text are texts contained in data on a book,

obtaining image information on images that are contained in the data on the book,

obtaining speech data corresponding to the subjects to be uttered; and

generating, based on the obtained utterance information, the obtained image information, and the obtained speech data, a speech synthesis model for reading out a text associated with an image.

7 . A non-transitory computer readable storage medium having a speech synthesis program stored thereon that, when executed by a processor, causes the processor to perform operations comprising:

obtaining acquiring utterance information on subjects to be uttered, wherein the subjects to be uttered is-text are texts contained in data on a book,

obtaining image information on images that are contained in the data on the book,

obtaining speech data corresponding to the subjects to be uttered; and

generating, based on the obtained utterance information, the obtained image information, and the obtained speech data, a speech synthesis model for reading out a text associated with an image.

8 . A speech synthesis apparatus comprising:

a memory; and

a processor coupled to the memory and configured to:

obtain utterance information on a subject to be uttered, wherein the subject to be uttered is a text contained in data on a book;

obtain image information on an image, wherein the image information corresponds to the text contained in the data on the book;

acquire a synthesized speech corresponding to the subject to be uttered by inputting the obtained utterance information and the obtained image information to a speech synthesis model for reading out a text that is associated with an image.

Assignments (2)
CHANGE OF NAME Recorded Aug 20, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072556/0180 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2024
From: IJAMA, YUSUKE; KORIYAMA, TOMOKI; TAKAMICHI, SHINNOSUKE
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION; THE UNIVERSITY OF TOKYO
Reel/Frame 066472/0531 →