IP Library › Granted Patent US 7,620,547
Granted Patent B2
US 7,620,547 · App. 11/042,892 · Granted Nov 17, 2009

Spoken man-machine interface with speaker identification

Assignee: Sony Deutschland GmbH
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,620,547
App. No.
11/042,892
Granted
Nov 17, 2009
Kind
B2
Abstract

The present invention provides a method for operating and/or for controlling a man-machine interface unit (MMI) for a finite user group environment. Utterances out of a group of user are repeatedly received. A process of user identification is carried out based on said received utterances. The process of user identification comprises a set of clustering so as to enable an enrolment-free performance.

Claims (84)

1. A method for operating a man-machine interface unit included in at least one of a home network system, a home entertainment system, and a service robot, the method comprising:

receiving an utterance of a person;

identifying the person on the basis of a previously computed speaker model as one of an unknown person and a known member of a predetermined group restricted to a predetermined, finite first number of members that have not undergone an enrollment process by speaking an enrollment text;

determining, on the basis of a confidence measure measuring reliability of the identification, whether a clustering process is to be performed;

including, if the clustering process is to be performed, the received utterance into a garbage class including at most a predetermined second number of most recently received utterances, and clustering the garbage class in an unsupervised manner with each of the included utterances forming an initial cluster by repeatedly merging most similar clusters until the remaining most similar clusters are more dissimilar than a predetermined threshold;

computing a further speaker model from one of the clusters if the one of the clusters includes more than a predetermined third number of utterances, thereby deleting utterances of the one of the clusters from the garbage class;

storing the further speaker model for identifying another person when receiving an utterance of the another person;

associating a first submodel with a first user profile and a second submodel with a second user profile, the first and second user profiles including a user preference;

determining a distance between the first submodel and the second submodel based on an acoustic distance and differences between the first user profile and the second user profile;

splitting the speaker model into the first and the second submodel if the determined distance between the first and second submodel exceeds a predefined threshold; and

operating the at least one of the home network system, the home entertainment system, and the service robot.

2. The method according to claim 1 ,

wherein the utterances include speech input.

3. The method according to claim 2 ,

wherein said clustering is carried out with respect to the speech input and with respect to respective different voices.

4. The method according to claim 2 , further comprising:

storing the speaker model together with the speech input associated therewith.

5. The method according to claim 1 ,

wherein the identifying includes a process of multi-talker, multi-speaker, or multi-user detection.

6. The method according to claim 1 , further comprising:

classifying a noise.

7. The method according to claim 1 ,

wherein the identifying includes identifying or updating the first number of members.

8. The method according to claim 7 , further comprising:

determining characteristics of said different members and acoustic characteristics of voices of said different members.

9. The method according to claim 8 , further comprising:

classifying voices of said different members to different voice classes based on features of said voices and based on differences or similarities of said voices.

10. The method according to claim 9 ,

wherein the classifying the voices is based on a frequency of occurrence of the voices.

11. The method according to claim 10 , further comprising:

assigning voices having a frequency of occurrence below a given threshold to the garbage class.

12. The method according to claim 11 , further comprising:

an initial phase of operation; and

using said garbage class as an initial class in the initial phase of operation.

13. The method according to claim 9 , further comprising:

generating confidence measures describing the reliability of the assignment of a voice to an assigned voice class.

14. The method according to claim 13 , further comprising:

repeatedly or interactively improving an algorithm or a parameter in the identifying to modify speaker identification parameters until said confidence measures are robust.

15. The method according to claim 14 ,

wherein the improving the algorithm or the parameter in the identifying further comprises collecting speech input of different situations including far-field talking situations, close-talking situations, or various background noise situations.

16. The method according to claim 9 , further comprising:

assigning different rights to said different voice classes.

17. The method according to claim 16 ,

wherein the assigning different rights further comprises assigning a right to a non-garbage voice class to introduce a new voice class as a new non-garbage class, and the right pertains to a later acquisition, recognition, assignment, or an explicit verbal order.

18. The method according to claim 1 , further comprising:

adding the further member utterance to improve a speaker model for said identified member.

19. The method according to claim 1 , further comprising:

receiving a further speech input from the identified member; and

determining if a speaker cluster can be split up into distinct subclusters based on the further speech input.

20. The method according to claim 1 , further comprising:

determining an acoustic characteristic or a profile of the member based on a distinct submodel; and

obtaining differences between said distinct submodels based on the determined acoustic characteristic or the profile of the member.

21. A method for operating or controlling an entertainment robot, or a home network, for a group including a finite number of members, the method comprising:

operating a man-machine interface unit included in the entertainment robot or the home network, the operating comprising

receiving an utterance of a person;

identifying the person on the basis of a previously computed speaker model as one of an unknown person or a known member of a predetermined group restricted to a predetermined, finite first number of members that have not undergone an enrollment process by speaking an enrollment text;

determining, on the basis of a confidence measure measuring reliability of the identification, whether a clustering process is to be performed;

including, if the clustering process is to be performed, the received utterance into a garbage class including at most a predetermined second number of most recently received utterances, and clustering the garbage class in an unsupervised manner with each of the included utterances forming an initial cluster by repeatedly merging most similar clusters until the remaining most similar clusters are more dissimilar than a predetermined threshold;

computing a further speaker model from one of the clusters if the one of the clusters includes more than a predetermined third number of utterances, thereby deleting utterances of the one of the clusters from the garbage class;

storing the further speaker model for identifying another person when receiving an utterance of the another person;

associating a first submodel with a first user profile and a second submodel with a second user profile, the first and second user profiles including a user preference;

determining a distance between the first submodel and the second submodel based on an acoustic distance and differences between the first user profile and the second user profile; and

splitting the speaker model into the first and the second submodel if the determined distance between the first and second submodel exceeds a predefined threshold.

22. A system for operating a man-machine interface unit, the system comprising:

a receiver configured to receive an utterance of a person;

an identifying unit configured to identify the person on the basis of a previously computed speaker model as one of an unknown person and a known member of a predetermined group restricted to a predetermined, finite first number of members that have not undergone an enrollment process by speaking an enrollment text;

a determining unit configured to

determine, on the basis of a confidence measure measuring reliability of the identification, whether a clustering process is to be performed,

include, if the clustering process is to be performed, the received utterance into a garbage class including at most a predetermined second number of most recently received utterances, and to cluster the garbage class in an unsupervised manner with each of the included utterances forming an initial cluster by repeatedly merging most similar clusters until the remaining most similar clusters are more dissimilar than a predetermined threshold, and

compute a further speaker model from one of the clusters if the one of the clusters includes more than a predetermined third number of utterances, thereby deleting utterances of the one of the clusters from the garbage class,

associate a first submodel with a first user profile and a second submodel with a second user profile, the first and second user profiles including a user preference,

determine a distance between the first submodel and the second submodel based on an acoustic distance and differences between the first user profile and the second user profile, and

split the speaker model into the first and the second submodel if the determined distance between the first and second submodel exceeds a predefined threshold; and

a memory configured to store the further speaker model for identifying another person when receiving an utterance of the another person.

23. A computer memory, comprising a computer program, which when executed by a computer, performs a method for operating a man-machine interface unit, comprising:

receiving an utterance of a person;

identifying the person on the basis of a previously computed speaker model as one of an unknown person or a known member of a predetermined group restricted to a predetermined, finite first number of members that have not undergone an enrollment process by speaking an enrollment text;

determining, on the basis of a confidence measure measuring reliability of the identification, whether a clustering process is to be performed;

including, if the clustering process is to be performed, the received utterance into a garbage class including at most a predetermined second number of most recently received utterances, and clustering the garbage class in an unsupervised manner with each of the included utterances forming an initial cluster by repeatedly merging most similar clusters until the remaining most similar clusters are more dissimilar than a predetermined threshold;

computing a further speaker model from one of the clusters if the one of the clusters includes more than a predetermined third number of utterances, thereby deleting utterances of the one of the clusters from the garbage class;

storing the further speaker model for identifying another person when receiving an utterance of the another person;

associating a first submodel with a first user profile and a second submodel with a second user profile, the first and second user profiles including a user preference;

determining a distance between the first submodel and the second submodel based on an acoustic distance and differences between the first user profile and the second user profile; and

splitting the speaker model into the first and the second submodel if the determined distance between the first and second submodel exceeds a predefined threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2005
From: KOMPE, RALF; KEMP, THOMAS
To: SONY DEUTSCHLAND GMBH
Reel/Frame 016507/0572 →
Priority Claims (1)
EP 02016672 · Jul 25, 2002 · regional
Continuity (2)
Continuation PCTEP20030806800 · Jul 23, 2003
Related Publication 20050187770A1 · Aug 25, 2005