IP Library Granted Patent US 10,109,286
Granted Patent B2
US 10,109,286 · App. 15/704,051 · Granted Oct 23, 2018

Speech synthesizer, audio watermarking information detection apparatus, speech synthesizing method, audio watermarking information detection method, and computer program product

Inventors: Kentaro Tachibana (Kanagawa, JP); Takehiko Kagoshima (Kanagawa, JP); Masatsune Tamura (Kanagawa, JP); Masahiro Morita (Kanagawa, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G10L19/018G10L13/02G10L13/033G10L19/012
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,109,286
App. No.
15/704,051
Granted
Oct 23, 2018
Kind
B2
Abstract

According to an embodiment, a speech synthesizer includes a source generator, a phase modulator, and a vocal tract filter unit. The source generator generates a source signal by using a fundamental frequency sequence and a pulse signal. The phase modulator modulates, with respect to the source signal generated by the source generator, a phase of the pulse signal at each pitch mark based on audio watermarking information. The vocal tract filter unit generates a speech signal by using a spectrum parameter sequence with respect to the source signal in which the phase of the pulse signal is modulated by the phase modulator.

Claims (23)

1. An audio watermarking information detection apparatus comprising:

a memory; and

one or more processors configured to function as a pitch mark estimator, a phase extractor, a representative phase calculator and a determination unit, wherein

the pitch mark estimator estimates a pitch mark of a synthesized speech in which audio watermarking information is embedded and extracts a speech at each estimated pitch mark;

the phase extractor extracts a phase of the speech extracted by the pitch mark estimator;

the representative phase calculator calculates a representative phase to be a representative of a plurality of frequency bins from the phase extracted by the phase extractor; and

the determination unit determines, based on the representative phase, whether the audio watermarking information exists in the synthesized speech.

2. The audio watermarking information detection apparatus according to claim 1 , wherein the determination unit

calculates, in each frame which is a predetermined period, an inclination indicating a variation of the representative phase in elapse of time, and

determines, based on a frequency of the inclination, whether there is the audio watermarking information.

3. The audio watermarking information detection apparatus according to claim 1 , wherein the determination unit

calculates, in each frame which is a predetermined period, a correlation coefficient between the representative phase and a reference straight line which is assumed as an ideal value of a variation of the representative phase in elapse of time, and

determines that there is the audio watermarking information when the correlation coefficient exceeds a predetermined threshold.

4. An audio watermarking information detection method employed for an audio watermarking information detection apparatus including a memory and one or more processors configured to function as a pitch mark estimator, a phase extractor, a representative phase calculator and a determination unit, comprising:

estimating, by the itch mark estimator, a pitch mark of a synthesized speech in which audio watermarking information is embedded and extracting a speech at each estimated pitch mark;

extracting, by the phase extractor, a phase of the extracted speech;

calculating, by representative phase calculator, from the extracted phase, a representative phase to be a representative of a plurality of frequency bins; and

determining, by the determination unit, based on the representative phase, whether the audio watermarking information exists in the synthesized speech.

5. A computer program product comprising a non-transitory computer-readable medium that includes an audio watermarking information detection program to cause a computer to execute:

estimating a pitch mark of a synthesized speech in which audio watermarking information is embedded and extracting a speech at each estimated pitch mark,

extracting a phase of the extracted speech,

calculating, from the extracted phase, a representative phase to be a representative of a plurality of frequency bins, and

determining, based on the representative phase, whether the audio watermarking information exists in the synthesized speech.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
Continuity (3)
Division 14801152 · Jul 16, 2015
Continuation PCTJP2013050990 · Jan 18, 2013
Related Publication 20180005637A1 · Jan 4, 2018