IP Library Granted Patent US 8,131,551
Granted Patent B1
US 8,131,551 · App. 11/458,282 · Granted Mar 6, 2012

System and method of providing conversational visual prosody for talking heads

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,131,551
App. No.
11/458,282
Granted
Mar 6, 2012
Kind
B1
Abstract

A system and method of controlling the movement of a virtual agent while the agent is speaking to a human user during a conversation is disclosed. The method comprises receiving speech data to be spoken by the virtual agent, performing a prosodic analysis of the speech data, selecting matching prosody patterns from a speaking database and controlling the virtual agent movement according to the selected prosody patterns.

Claims (33)

1. A system comprising:

a processor;

a first module controlling the processor to perform a prosodic analysis and a syntactic analysis of speech data to be spoken by a virtual agent to a user, the prosodic analysis comprising analyzing speech intonations comprising loudness and accent, identifying prosodic phrase boundaries in the speech data, and identifying a type for each of the prosodic phrase boundaries;

a second module controlling the processor to determine a culture of the user based on an analysis of prosody associated with received speech from the user, the analysis being independent of an identity of the user; and

a third module controlling the processor to control movement of the virtual agent according to the prosodic analysis, the syntactic analysis, and the culture of the user and not based on a previously-stored template for controlling the movement, wherein the movement of the virtual agent at each of the prosodic phrase boundaries is selected based on the type identified for each of the prosodic phrase boundaries.

2. The system of claim 1 , wherein the movement of the virtual agent is controlled to be approximately simultaneous with the received speech that triggers the movement.

3. The system of claim 1 , further comprising:

a fourth module controlling the processor to receive the speech data to be spoken by the virtual agent to the user.

4. The system of claim 1 , wherein the movement of the virtual agent is at least one of: raising a head of the virtual agent, raising at least one eyebrow of the virtual agent, and nodding the head of the virtual agent.

5. The system of claim 1 , wherein the system is a client device that communicates over a network with a server.

6. The system of claim 5 , further comprising:

a fourth module controlling the processor, after the virtual agent finishes speaking to the user, to receive additional speech data from the user and transmit the additional speech data to the server for speech processing in order to generate a response by the virtual agent to the user.

7. The system of claim 6 , further comprising:

a fifth module controlling the processor, after the user finishes a speech segment, to receive responsive speech from the server for generating a virtual agent response to the user.

8. The system of claim 7 , wherein the third module controls the processor to control movement of the virtual agent to be approximately simultaneous with at least one of user speech data and virtual agent speech data that triggers the movement.

9. The system of claim 5 , wherein the movement of the virtual agent while the virtual agent speaks to the user is based on at least one of additional speech data received over the network from the server and virtual agent movement data received over the network from the server.

10. The system of claim 5 , wherein the network comprises at least one of the Internet, a packet network, a wireless network and an Internet Protocol network.

11. The system of claim 1 , wherein the prosodic analysis further comprises selecting segments of matching visual prosody patterns from an audio-visual database of recorded speech, where both audio and video are recorded of a person speaking, and wherein the third module controls the movement of the virtual agent according to matched visual prosody patterns.

12. A system for controlling movement of a virtual agent on a client device while the virtual agent is speaking to a user, the system comprising a server that:

transmits speech data to be spoken by the virtual agent to the client device over a network;

generates virtual agent movement data based on a prosodic analysis, a syntactic analysis of the speech data and a culture of the user determined based on an analysis of prosody associated with received speech from the user, independent of an identity of the user, an identification of phrase boundaries in each utterance defined in the speech data and a phrase boundary type for each of the phrase boundaries, and not based on a previously-stored template for controlling the movement of the virtual agent, wherein the virtual agent movement data is configured to synchronize the movement of the virtual agent with phrase boundaries and to reflect a pitch accent associated with the phrase boundary type associated with each of the phrase boundaries; and

transmits the virtual agent movement data to the client device over the network for controlling movement of the virtual agent while the virtual agent speaks to the user.

13. The system of claim 12 , wherein the network comprises at least one of: the Internet, a packet network, a wireless network and an Internet Protocol network.

14. A system for controlling movement of a virtual animated entity during a transition from speaking to listening, the system comprising:

a processor;

a first module controlling the processor, as the virtual animated entity is concluding a speaking segment, to select transition movement data based at least in part on a syntactic analysis of speech to be spoken by the virtual animated entity and further based on a user culture determined by an analysis of prosody associated with received speech from a user, the analysis being independent of an identity of the user and the transition movement data not based on a previously-stored template for controlling the movement of the virtual animated entity; and

a second module controlling the processor to control the movement of the virtual animated entity from a first time the virtual animated entity has approximately finished speaking and through a second time at which the virtual animated entity stops speaking based on the user culture, wherein after the virtual animated entity stops speaking the transition movement data continues to control movement of the virtual animated entity to signal the user to speak.

15. The system of claim 14 , wherein the transition movement data is selected from a transition movement database.

16. A system for controlling movement of a virtual animated entity during a transition from talking to listening, the system comprising:

a processor;

a first module controlling the processor, approximately at an end of the virtual animated entity talking, to select transition movement data based at least in part on a syntactic analysis of speech to be spoken by the virtual animated entity and further based on a user culture determined by an analysis of prosody associated with received speech from a user, the analysis being independent of an identity of the user and the transition movement data not based on a previously-stored template for controlling movement of the virtual animated entity; and

a second module controlling the processor to control the movement of the virtual animated entity to indicate that the virtual animated entity is approximately finished talking and will soon listen for speech data from the user based on the user culture, the movement including movement after the virtual animated entity finishes talking to signal the user to speak.

17. The system of claim 16 , wherein the transition movement data is selected from a transition database.

Assignments (16)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 034480/0215 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034482/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 034480/0260 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2014
From: COSATTO, ERIC; GRAF, HANS PETER; STROM, VOLKER FRANZ
To: AT&T CORP.
Reel/Frame 034480/0165 →