IP Library › Granted Patent US 10,854,219
Granted Patent B2
US 10,854,219 · App. 16/002,208 · Granted Dec 1, 2020

Voice interaction apparatus and voice interaction method

Inventor: Hiraku Kayama (Hamamatsu, JP)
Assignee: Yamaha Corporation
G10L25/90G10L13/00G10L13/10G10L15/02G10L15/1807G10L15/22G10L25/48G10L13/0335G10L21/013
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,854,219
App. No.
16/002,208
Granted
Dec 1, 2020
Kind
B2
Abstract

A voice interaction apparatus acquires a speech signal indicative of a speech sound, identifies a series of pitches of the speech sound from the speech signal, and causes a reproduction device to reproduce a response voice of pitches controlled in accordance with the lowest pitch of the pitches identified during a tailing section proximate to an end point within the speech sound.

Claims (56)

1. A voice interaction method comprising:

acquiring a speech signal indicative of a speech sound that is directed toward an interacting partner;

identifying a series of pitches of the speech sound from the speech signal;

identifying a lowest pitch among the series of pitches, wherein the series of pitches are pitches of a tailing section proximate to an end point within the speech sound; and

causing a reproduction device to reproduce a response voice of pitches controlled in accordance with the lowest pitch.

2. The voice interaction method according to claim 1 ,

wherein in the causing of the reproduction device to reproduce the response voice, an initial pitch of a final mora of the response voice is controlled in accordance with the lowest pitch of the tailing section within the speech sound.

3. The voice interaction method according to claim 1 ,

wherein a plurality of speech signals are acquired and for each of the plurality of acquired speech signals, the identifying of a series of pitches of a speech sound indicated by the speech signal and the causing of the reproduction device to reproduce a response voice are executed, and

wherein pitches of response voices corresponding to the plurality of speech signals are controlled differently for each of the plurality of speech signals.

4. The voice interaction method according to claim 1 ,

wherein the response voice has a prosody corresponding to transition of the pitches identified during the tailing section.

5. The voice interaction method according to claim 4 ,

wherein the response voice has a different prosody between a case where the identified pitches decrease and then increase within the tailing section and a case where the identified pitches decrease from a start point to an end point of the tailing section.

6. The voice interaction method according to claim 4 ,

wherein the causing of the reproduction device to reproduce the response voice includes comparing a first average pitch with a second average pitch, wherein the first average pitch is an average pitch in a first section within the tailing section and the second average pitch is an average pitch in a second section within the tailing section, the second section coming after the first section, and

wherein the response voice has a different prosody between a case where the first average pitch is lower than the second average pitch and a case where the first average pitch is higher than the second average pitch.

7. The voice interaction method according to claim 4 ,

wherein the causing of the reproduction device to reproduce the response voice includes:

acquiring a response signal indicative of the response voice from a storage device that stores a plurality of response signals indicative of response voices with different prosodies; and

outputting the acquired response signal to cause the reproduction device to reproduce the response voice.

8. The voice interaction method according to claim 4 ,

wherein the causing of the reproduction device to reproduce the response voice includes:

generating, from a response signal indicative of a response voice with a predetermined prosody, a response signal indicative of the response voice; and

outputting the generated response signal to cause the reproduction device to reproduce the response voice.

9. The voice interaction method according to claim 4 , wherein the prosody includes an identified prosody index value that is calculated for each speech sound.

10. The voice interaction method according to claim 1 ,

wherein in the causing of the reproduction device to reproduce the response voice, the response voice is selected from among a first response voice and a second response voice, wherein the first response voice represents an inquiry directed toward the speech sound and the second response voice represents a response other than an inquiry.

11. The voice interaction method according to claim 10 , further comprising identifying from the speech signal a prosody index value indicative of a prosody of the speech sound,

wherein the causing of the reproduction device to reproduce the response voice includes:

comparing the prosody index value of the speech sound with a threshold value; and

selecting either the first response voice or the second response voice as the response voice in accordance with a result of the comparison.

12. The voice interaction method according to claim 11 ,

wherein a plurality of speech signals are acquired and for each of the plurality of acquired speech signals, the identifying of a series of pitches, the identifying of a prosody index value, and the causing of the reproduction device to reproduce a response voice are executed, and

wherein the threshold value is set to a representative value of prosody index values identified from the plurality of speech signals.

13. The voice interaction method according to claim 11 ,

wherein in the causing of the reproduction device to reproduce the response voice, the first response voice is selected in a case where the prosody index value is a value outside a predetermined range that includes the threshold value, and the second response voice is selected in a case where the prosody index value is a value within the predetermined range.

14. The voice interaction method according to claim 10 ,

wherein a plurality of speech signals are acquired and for each of the plurality of acquired speech signals, the identifying of a series of pitches and the causing of the reproduction device to reproduce a response voice are executed, and

wherein in the causing of the reproduction device to reproduce the response voice, the first response voice is selected as the response voice directed toward a speech sound that is selected randomly from among a plurality of speech sounds indicated by the plurality of speech signals.

15. The voice interaction method according to claim 14 ,

wherein the causing of the reproduction device to reproduce the response voice includes setting a frequency for reproducing the first response voice as the response voice directed toward the plurality of speech sounds.

16. The voice interaction method according to claim 15 ,

wherein the frequency for reproducing the first response voice is set in accordance with a usage history of a voice interaction.

17. The voice interaction method according to claim 1 , the method further comprising generating a usage history of a voice interaction in which the response voice is reproduced in response to the speech sound,

wherein the response voice has a prosody corresponding to the usage history.

18. The voice interaction method according to claim 17 ,

wherein in the causing of the reproduction device to reproduce the response voice, an interval between the speech sound and the response voice is controlled as the prosody in accordance with the usage history.

19. A voice interaction apparatus comprising: a processor coupled to a memory storing instructions that, when executed by the processor, configure the processor to: acquire a speech signal indicative of a speech sound that is directed toward an interacting partner; identify a series of pitches of the speech sound from the speech signal; identify a lowest pitch among the series of pitches, wherein the series of pitches are pitches of a tailing section proximate to an end point within the speech sound; and cause a reproduction device to reproduce a response voice of pitches controlled in accordance with the lowest pitch.

20. The voice interaction apparatus according to claim 19 ,

wherein the response voice has a prosody corresponding to transition of the pitches identified during the tailing section.

21. The voice interaction apparatus according to claim 19 ,

wherein the response voice is selected from among a first response voice and a second response voice, wherein the first response voice represents an inquiry directed toward the speech sound and the second response voice represents a response other than an inquiry.

22. The voice interaction apparatus according to claim 19 ,

wherein the processor is further configured to generate a usage history of a voice interaction in which the response voice is reproduced in response to the speech sound; and

wherein the response voice has a prosody corresponding to the usage history.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2018
From: KAYAMA, HIRAKU
To: YAMAHA CORPORATION
Reel/Frame 046017/0699 →
Priority Claims (5)
JP 2015-238911 · Dec 7, 2015 · national
JP 2015-238912 · Dec 7, 2015 · national
JP 2015-238913 · Dec 7, 2015 · national
JP 2015-238914 · Dec 7, 2015 · national
JP 2016-088720 · Apr 27, 2016 · national
Continuity (2)
Continuation PCTJP2016085126 · Nov 28, 2016
Related Publication 20180294001A1 · Oct 11, 2018