IP Library Granted Patent US 8,027,837
Granted Patent B2
US 8,027,837 · App. 11/532,470 · Granted Sep 27, 2011

Using non-speech sounds during text-to-speech synthesis

Assignee: Apple Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,027,837
App. No.
11/532,470
Granted
Sep 27, 2011
Kind
B2
Abstract

Systems, apparatus, methods and computer program products are described for producing text-to-speech synthesis with non-speech sounds. In general, some of the pauses or silences that would otherwise be generated in synthesized speech are instead synthesized as non-speech sounds such as breaths. Non-speech sounds can be identified from pre-recorded speech that can include meta-data such as the grammatical and phrasal structure of words and sounds that precede and succeed non-speech sounds. A non-speech sound can be selected for use in synthesized speech based on the words, punctuation, grammatical and phrasal structure of text from which the speech is being synthesized, or other characteristics.

Claims (96)

1. A method, comprising:

parsing text into speech units and non-speech units at a first speech unit level;

attempting to match a non-speech unit with a first audio segment;

determining that there are unmatched non-speech units at the first speech unit level;

parsing speech units adjacent to unmatched non-speech units into speech units at a second speech unit level;

attempting to match an unmatched non-speech unit having an adjacent speech unit at the second speech unit level with a second audio segment; and

creating a portion of speech by synthesizing a portion of the text string containing speech units into speech and

augmenting the portion of synthesized speech with the first or second audio segment.

2. The method of claim 1 , where a non-speech sound includes the sound of one or more of:

inhalation;

exhalation;

mouth clicks;

lip smacks;

tongue flicks; and

salivation.

3. A computer-readable, non-transitory storage medium having instructions stored thereon, which, when executed by a processor, causes the processor to perform operations, comprising:

parsing text into speech units and non-speech units at a first speech unit level;

attempting to match a non-speech unit with a first audio segment;

determining that there are unmatched non-speech units at the first speech unit level;

parsing speech units adjacent to unmatched non-speech units into speech units at a second speech unit level;

attempting to match an unmatched non-speech unit having an adjacent speech unit at the second speech unit level with a second audio segment; and

creating a portion of speech by synthesizing a portion of the text string containing speech units into speech; and

augmenting the portion of synthesized speech with the first or second audio segment.

4. A system comprising:

a processor;

memory having instructions stored thereon, which, when executed by the processor, cause the processor to perform operations, comprising:

parsing text into speech units and non-speech units at a first speech unit level;

attempting to match a non-speech unit with a first audio segment;

determining that there are unmatched non-speech units at the first speech unit level;

parsing speech units adjacent to unmatched non-speech units into speech units at a second speech unit level;

attempting to match an unmatched non-speech unit having an adjacent speech unit at the second speech unit level with a second audio segment; and

creating a portion of speech by synthesizing a portion of the text string containing speech units into speech and

augmenting the portion of synthesized speech with the first or second audio segment.

5. A method comprising:

parsing a text string into phrase units and non-speech units;

attempting to match a non-speech unit to a first audio segment;

determining that there are unmatched non-speech units;

parsing phrase units adjacent to unmatched non-speech units into word units;

attempting to match an unmatched non-speech unit having an adjacent word unit to a second audio segment; and

creating a portion of speech by synthesizing a portion of the text string containing speech units into speech and

augmenting the portion of synthesized speech with the first or second audio segment.

6. The method of claim 5 , further comprising:

after attempting to match an unmatched non-speech unit having an adjacent word unit to a second audio segment, determining that there are unmatched non-speech units;

parsing word units adjacent to unmatched non-speech units into subword units;

attempting to match an unmatched non-speech unit having an adjacent subword unit to a third audio segment; and

augmenting the portion of synthesized speech with the third audio segment.

7. The method of claim 5 , where a non-speech sound includes the sound of one or more of:

inhalation;

exhalation;

mouth clicks;

lip smacks;

tongue flicks; and

salivation.

8. A computer-readable, non-transitory storage medium having instructions stored thereon, which, when executed by a processor, causes the processor to perform operations, comprising:

parsing a text string into phrase units and non-speech units;

attempting to match a non-speech unit to a first audio segment;

determining that there are unmatched non-speech units;

parsing phrase units adjacent to unmatched non-speech units into word units;

attempting to match an unmatched non-speech unit having an adjacent word unit to a second audio segment; and

creating a portion of speech by synthesizing a portion of the text string containing speech units into speech and

augmenting the portion of synthesized speech with the first or second audio segment.

9. The computer-readable, non-transitory storage medium of claim 8 , wherein the instructions include instructions which cause the processor to perform operations, comprising:

after attempting to match an unmatched non-speech unit having an adjacent word unit to a second audio segment, determining that there are unmatched non-speech units;

parsing word units adjacent to unmatched non-speech units into subword units;

attempting to match an unmatched non-speech unit having an adjacent subword unit to a third audio segment; and

augmenting the portion of synthesized speech with the third audio segment.

10. The computer-readable, non-transitory storage medium of claim 8 , where a non-speech sound includes the sound of one or more of:

inhalation;

exhalation;

mouth clicks;

lip smacks;

tongue flicks; and

salivation.

11. A system comprising:

a processor;

memory having instructions stored thereon, which, when executed by the processor,

cause the processor to perform operations, comprising:

parsing a text string into phrase units and non-speech units;

attempting to match a non-speech unit to a first audio segment;

determining that there are unmatched non-speech units;

parsing phrase units adjacent to unmatched non-speech units into word units;

attempting to match an unmatched non-speech unit having an adjacent word unit to a second audio segment; and

creating a portion of speech by synthesizing a portion of the text string containing speech units into speech and

augmenting the portion of synthesized speech with the first or second audio segment.

12. The system of claim 11 , wherein the instructions include instructions which cause the processor to perform operations, comprising:

after attempting to match an unmatched non-speech unit having an adjacent word unit to a second audio segment, determining that there are unmatched non-speech units;

parsing word units adjacent to unmatched non-speech units into subword units;

attempting to match an unmatched non-speech unit having an adjacent subword unit to a third audio segment; and

augmenting the portion of synthesized speech with the third audio segment.

13. The system of claim 11 , where a non-speech sound includes the sound of one or more of:

inhalation;

exhalation;

mouth clicks;

lip smacks;

tongue flicks; and

salivation.

Assignments (2)
CHANGE OF NAME Recorded Apr 10, 2007
From: APPLE COMPUTER, INC.
To: APPLE INC.
Reel/Frame 019142/0969 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2006
From: SILVERMAN, KIM E.A.; NEERACHER, MATTHIAS
To: APPLE COMPUTER, INC.
Reel/Frame 018292/0854 →
Continuity (1)
Related Publication 20080071529A1 · Mar 20, 2008