IP Library Granted Patent US 12,027,165
Granted Patent B2
US 12,027,165 · App. 17/371,116 · Granted Jul 2, 2024

Computer program, server, terminal, and speech signal processing method

Inventor: Akihiko Shirai (Tokyo, JP)
Assignee: GREE, INC.
G10L15/22G06F3/14G10L15/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,027,165
App. No.
17/371,116
Granted
Jul 2, 2024
Kind
B2
Abstract

A non-transitory computer readable medium stores computer executable instructions which, when executed by at least one processor, cause the at least one processor to acquire a speech signal of speech of a user; perform a signal processing on the speech signal to acquire at least one feature of the speech of the user; and control display of information, related to each of one or more first candidate converters having a feature corresponding to the at least one feature, to present the one or more first candidate converters for selection by the user.

Claims (50)

1. A non-transitory computer readable medium storing computer executable instructions which, when executed by at least one processor, cause the at least one processor to:

acquire a speech signal of speech of a user;

perform a signal processing on the speech signal to acquire at least one feature of the speech of the user, wherein the at least one feature includes a first formant, a second formant, and/or a fundamental frequency; and

control display of information, related to each of one or more first candidate converters having a feature that is within a minimum and maximum value of the at least one feature, to present the one or more first candidate converters for selection by the user.

2. The non-transitory computer readable medium according to claim 1 , wherein

the at least one processor is further caused to control display of information related to each of the one or more first candidate converters to present the one or more first candidate converters for selection by the user, and

the one or more first candidate converters have a first formant, a second formant, and/or a fundamental frequency corresponding to the acquired first formant, the acquired second formant, and/or the acquired fundamental frequency of the user.

3. The non-transitory computer readable medium according to claim 2 , wherein the information related to each of the one or more first candidate converters includes information about a person and/or a character having a voice similar to a voice of the user.

4. The non-transitory computer readable medium according to claim 1 , wherein the at least one processor is further caused to control display of information related to each of one or more second candidate converters to present the one or more second candidate converters for selection by the user.

5. The non-transitory computer readable medium according to claim 4 , wherein the information related to the one or more second candidate converters is displayed due to the one or more second candidate converters being used by plural other devices at a high rate and/or at a high usage count.

6. The non-transitory computer readable medium according to claim 4 , wherein the information related to each of the one or more second candidate converters is displayed irrespective of whether the one or more second candidate converters has a feature corresponding to the at least one feature.

7. The non-transitory computer readable medium according to claim 4 , wherein the information related to each of the one or more second candidate converters includes information about a person, a character, and/or an intended voice.

8. The non-transitory computer readable medium according to claim 1 , wherein the at least one processor is further caused to

estimate emotion and/or personality of the user in accordance with the first formant, the second formant, and loudness indicated in a result of the signal processing of the speech signal, and

extract the one or more first candidate converters from among a plurality of prepared converters in accordance with information indicating the estimated emotion and/or personality.

9. The non-transitory computer readable medium according to claim 1 , wherein the at least one processor is further caused to

separately acquire a first speech signal of high-pitched speech of the user and a second speech signal of low-pitched speech of the user,

perform a signal processing on the first speech signal and the second speech signal to acquire a plurality of features of speech of the user, and

acquire, in accordance with the plurality of features, a converter that converts at least one of the plurality of features on an input speech signal to generate an output speech signal.

10. The non-transitory computer readable medium according to claim 9 , wherein the plurality of features includes a first formant, a second formant, and a fundamental frequency.

11. The non-transitory computer readable medium according to claim 10 , wherein the at least one processor is further caused to

acquire a third speech signal of a natural speech of the user, and

perform a signal processing on the first speech signal, the second speech signal, and the third speech signal to acquire the first formant, the second formant, and the fundamental frequency.

12. The non-transitory computer readable medium according to claim 10 , wherein the converter includes

a first parameter indicating a first frequency to which the first formant of the input speech signal is shifted,

a second parameter indicating a second frequency to which the second formant of the input speech signal is shifted, and

a third parameter indicating a third frequency to which the fundamental frequency of the input speech signal is shifted.

13. The non-transitory computer readable medium according to claim 10 , wherein the at least one processor is further caused to

acquire a frequency range of a voice of the user, obtained in accordance with a minimum value and a maximum value of each of the first formant, the second formant, and the fundamental frequency, and

acquire the converter that shifts a pitch while a number of bits to be allocated to a part of the input speech signal included in the frequency range is greater than a number of bits to be allocated to another part of the input speech signal not included in the frequency range.

14. The non-transitory computer readable medium according to claim 10 , wherein

the at least one processor is further caused to select and acquire the converter from among a plurality of converters prepared, and

each of the plurality of converters includes a first parameter indicating a frequency to which a first formant of an input speech signal is shifted, a second parameter indicating a frequency to which a second formant of the input speech signal is shifted, and a third parameter indicating a frequency to which a fundamental frequency of the input speech signal is shifted.

15. The non-transitory computer readable medium according to claim 14 , wherein the at least one processor is caused to select and acquire the converter having a first formant, a second formant, and/or a fundamental frequency corresponding to the acquired first formant, the acquired second formant, and/or the acquired fundamental frequency of the user from among the plurality of converters.

16. The non-transitory computer readable medium according to claim 14 , wherein the at least one processor is further caused to

acquire a fourth speech signal of speech the user speaks in imitation of a desired person or character,

perform a signal processing on the each of the first speech signal, the second speech signal, the third speech signal, and the fourth speech signal to acquire a first formant, a second formant and a fundamental frequency and

select and acquire another converter having a first formant, a second formant, and/or a fundamental frequency corresponding to the first formant, the second formant, and/or the fundamental frequency of the user, calculated by the signal processing of the fourth speech signal, from among the plurality of converters.

17. The non-transitory computer readable medium according to claim 9 , wherein the at least one processor is caused to

generate an output speech signal by shifting a pitch of the speech signal as an input speech signal with the converter, and

send the output speech signal to a server or a terminal.

18. A device, comprising:

processing circuitry configured to

acquire a speech signal of speech of a user;

perform a signal processing on the speech signal to acquire at least one feature of the speech signal, wherein the at least one feature includes a first formant, a second formant, and/or a fundamental frequency; and

control display of information, related to each of one or more first candidate converters having a feature that is within a minimum and maximum value of the at least one feature, to present the one or more first candidate converters for selection by the user.

19. A speech signal processing method, comprising:

acquiring, by processing circuitry, a speech signal of speech of a user;

performing, by the processing circuitry, a signal processing on the speech signal to acquire at least one feature of the speech signal, wherein the at least one feature includes a first formant, a second formant, and/or a fundamental frequency; and

controlling display of information, related to each of one or more first candidate converters having a feature that is within a minimum and maximum value of the at least one feature, to present the one or more first candidate converters for selection by the user.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY STREET ADDRESS AND ATTORNEY DOCKET NUMBER PREVIOUSLY RECORDED AT REEL: 71308 FRAME: 765. ASSIGNOR(S) HEREBY CONFIRMS THE CHANGE OF NAME. Recorded Jun 10, 2025
From: GREE, INC.
To: GREE HOLDINGS, INC.
Reel/Frame 071611/0252 →
CHANGE OF NAME Recorded May 16, 2025
From: GREE, INC.
To: GREE HOLDINGS, INC.
Reel/Frame 071308/0765 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2021
From: SHIRAI, AKIHIKO
To: GREE, INC.
Reel/Frame 056944/0047 →
Priority Claims (2)
JP 2019-002923 · Jan 10, 2019 · national
JP 2019-024354 · Feb 14, 2019 · national
Continuity (2)
Continuation PCTJP2020000497 · Jan 9, 2020
Related Publication 20210335364A1 · Oct 28, 2021