IP Library Granted Patent US 10,304,439
Granted Patent B2
US 10,304,439 · App. 15/388,053 · Granted May 28, 2019

Image processing device, animation display method and computer readable medium

Inventors: Shoichi Okaniwa (Fussa, JP); Hiroaki Negishi (Hamura, JP); Shigekatsu Moriya (Tokyo, JP); Hirokazu Kanda (Ome, JP)
Assignee: CASIO COMPUTER CO., LTD.
G10L13/08G06F17/279G06F17/2755G06F17/2775G06T13/40G06F17/2836G06T13/00G10L15/26G10L2021/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,304,439
App. No.
15/388,053
Granted
May 28, 2019
Kind
B2
Abstract

An image processing device includes a controller and a display. The controller adds an expression to a displayed face image in accordance with an audio when the audio is output. Further, the controller generates an animation in which a mouth contained in the face image with the expression moves in sync with the audio. The display displays the generated animation.

Claims (41)

1. An image processing device comprising:

a processor configured to:

detect words or phrases within text of a sentence or a clause, wherein the text corresponds to audio to be reproduced;

determine, for at least one of the words or phrases detected within the text of the sentence or the clause, a corresponding one of a plurality of word/phrase-expressions;

determine that at least one of the words or phrases within the text of the sentence or the clause is a context-dependent word/phrase-expression;

assign a most frequent one of the word/phrase-expression determined for the at least one of the words or phrases detected, while ignoring the context-dependent word/phrase-expression determined, as one of a plurality of sentence/clause-expressions to the text of the sentence or the clause; and

generate frames of animation of a face of increased expressiveness to be displayed in sync with a reproduction of the audio, by at least performing:

generate a mouth shape of a mouth of the face for each of the frames based on the words or phrases detected within the text; and

generate an emotional expression of the face for each of the frames based on the one of the plurality of sentence/clause-expressions assigned to the text of the sentence or the clause.

2. The image processing device according to claim 1 ,

wherein the processor is configured to:

determine whether to generate the animation in a different language from that of the audio; and

in response to determining to generate the animation in the different language from that of the audio:

generate a mouth shape of the mouth of the face for each of the frames based on words or phrases detected within a text of audio in the different language; and

generate the emotional expression of the face for each of the frames based on the one of the plurality of sentence/clause-expressions assigned.

3. A method comprising:

detecting words or phrases within text of a sentence or a clause, wherein the text corresponds to audio to be reproduced;

determining, for at least one of the words or phrases detected within the text of the sentence or the clause, a corresponding one of a plurality of word/phrase expressions;

determining that at least one of the words or phrases within the text of the sentence or the clause is a context-dependent word/phrase-expression;

assigning a most frequent one of the word/phrase-expression determined for the at least one of the words or phrases detected, while ignoring the context-dependent word/phrase-expression determined, as one of a plurality of sentence/clause-expressions to the text of the sentence or the clause; and

generating frames of animation of a face of increased expressiveness to be displayed in sync with a reproduction of the audio, by at least:

generating a mouth shape of a mouth of the face for each of the frames based on the words or phrases detected within the text; and

generating an emotional expression of the face for each of the frames based on the one of the plurality of sentence/clause-expressions assigned to the text of the sentence or the clause.

4. The method according to claim 3 , comprising:

determining whether to generate the animation in a different language from that of the audio; and

in response to determining to generate the animation in the different language from that of the audio:

generating a mouth shape of the mouth of the face for each of the frames based on words or phrases detected within a text of audio in the different language; and

generating the emotional expression of the face for each of the frames based on the one of the plurality of sentence/clause-expressions assigned.

5. A non-transitory computer readable storage medium storing a program to cause a computer to at least perform:

detecting words or phrases within text of a sentence or a clause, wherein the text corresponds to audio to be reproduced;

determining, for at least one of the words or phrases detected within the text of the sentence or the clause, a corresponding one of a plurality of word/phrase-expressions;

determining that at least one of the words or phrases within the text of the sentence or the clause is a context-dependent word/phrase-expression;

assigning a most frequent one of the word/phrase-expression determined for the at least one of the words or phrases detected, while ignoring the context-dependent word/phrase-expression determined, as one of a plurality of sentence/clause-expressions to the text of the sentence or the clause; and

generating frames of animation of a face of increased expressiveness to be displayed in sync with a reproduction of the audio, by at least:

generating a mouth shape of a mouth of the face for each of the frames based on the words or phrases detected within the text; and

generating an emotional expression of the face for each of the frames based on the one of the plurality of sentence/clause-expressions assigned to the text of the sentence or the clause.

6. The non-transitory computer readable storage medium according to claim 5 , wherein the program causes the computer to perform:

determining whether to generate the animation in a different language from that of the audio; and

in response to determining to generate the animation in the different language from that of the audio:

generating a mouth shape of the mouth of the face for each of the frames based on words or phrases detected within a text of audio in the different language; and

generating the emotional expression of the face for each of the frames based on the one of the plurality of sentence/clause-expressions assigned.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 22, 2016
From: OKANIWA, SHOICHI; NEGISHI, HIROAKI; MORIYA, SHIGEKATSU; KANDA, HIROKAZU
To: CASIO COMPUTER CO., LTD.
Reel/Frame 040745/0926 →
Priority Claims (1)
JP 2016-051932 · Mar 16, 2016 · national
Continuity (1)
Related Publication 20170270701A1 · Sep 21, 2017
Cited By (2)
US 12,597,175 US 12,602,849