IP Library › Granted Patent US 8,804,918
Granted Patent B2
US 8,804,918 · App. 14/151,865 · Granted Aug 12, 2014

Method and system for using conversational biometrics and speaker identification/verification to filter voice streams

Inventors: Peeyush Jaiswal (Boca Raton, FL); Naveen Narayan (Flower Mound, TX)
Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,804,918
App. No.
14/151,865
Filed
Jan 10, 2014
Granted
Aug 12, 2014
Kind
B2
Examiner
HUYNH, VAN D
Art Unit
2653
USPC
379/88.02
Abstract

A method and system for using conversational biometrics and speaker identification and/or verification to filter voice streams during mixed mode communication. The method includes receiving an audio stream of a communication between participants. Additionally, the method includes filtering the audio stream of the communication into separate audio streams, one for each of the participants. Each of the separate audio streams contains portions of the communication attributable to a respective participant. Furthermore, the method includes outputting the separate audio streams to a storage system.

Claims (33)

1. A method comprising:

filtering by a computing device an audio stream of a communication into separate audio streams corresponding, respectively, to each of a plurality of participants in the communication, wherein each of the separate audio streams contains portions of the communication attributable to the corresponding one of the plurality of participants;

identifying by the computing device the plurality of participants by matching one or more of the portions of the communication in each of the separate audio streams to voice prints; and

subsequently in the communication, comparing by the computing device the separate audio streams to only a plurality of the voice prints corresponding to the identified participants within the communication.

2. The method of claim 1 , wherein the at least one component is further operable to perform a verification process for at least one of the plurality of participants.

3. The method of claim 1 , wherein the voice prints are seed phrases received for each of the plurality of participants.

4. The method of claim 1 , wherein:

providers of the voice prints are each associated with a role; and

each of the voice prints is stored in one of a plurality of discrete databases based on the role of the associated provider.

5. The method of claim 1 , wherein the filtering the audio stream comprises:

matching the one or more of the portions of the communication to the voice prints using at least one of conversational biometrics, speaker identification and speaker verification; and

assigning the one or more portions of the communication attributable to each of the plurality of participants to the separate audio streams corresponding to each of the plurality of participants.

6. The method of claim 1 , further comprising verifying at least one of the plurality of participants by authenticating a match between a voice sample of the at least one of the plurality of participants and a previously stored voice print of the at least one of the plurality of participants.

7. The method of claim 1 , receiving the audio stream of the communication via at least one of voice transmission technology, a wired telephone, a wireless telephone and a microphone.

8. The method of claim 1 , further comprising converting the communication from speech to text.

9. The method of claim 1 , wherein the communication is a call center communication and the plurality of participants includes at least two of: a caller, an agent and an interactive voice response (IVR) system.

10. The method of claim 9 , further collecting seed phrases for the agent and the IVR system prior to the receiving the audio stream.

11. The method of claim 1 , wherein the communication is a medical professional communication and the plurality of participants include at least a medical professional and a patient.

12. The method of claim 1 , wherein the filtering the audio stream of the communication into separate audio streams comprises using a voice print speaker identification for all but one of the plurality of participants and a process of elimination for the one of the plurality of participants.

13. The method of claim 1 , wherein the filtering the audio stream of the communication into the separate audio streams is performed in real-time.

14. The method of claim 1 , wherein the filtering the audio stream of the communication into the separate audio streams is performed in a batch.

15. The method of claim 1 , wherein the filtering the audio stream of the communication into the separate audio streams further comprises determining an origin of voice of at least one of the plurality of participants.

16. A system comprising:

one or more sensors sensor; and

a computing device including a conversational biometrics analysis and speaker identification-verification (CB/SIV) tool,

wherein the CB/SIV controls the computing device to:

filter an audio stream received via the one or more sensors into separate audio streams corresponding, respectively, to each of a plurality of participants in the communication, wherein each of the separate audio streams contains portions of the communication attributable to the corresponding one of the plurality of participants;

identify the plurality of participants by matching one or more of the portions of the communication in each of the separate audio streams to voice prints; and

compare the separate audio streams to only to a plurality of the voice prints corresponding to the identified participants within that conversation.

17. The system of claim 16 , wherein the filtering the audio stream of the communication into separate audio streams comprises using a voice print speaker identification for all but one of the plurality of participants and a process of elimination for the one of the plurality of participants.

18. The system of claim 16 , wherein the filtering the audio stream of the communication into the separate audio streams is performed in real-time.

19. The system of claim 16 , wherein the filtering the audio stream of the communication into the separate audio streams is performed in a batch.

20. The system of claim 16 , wherein the filtering the audio stream of the communication into the separate audio streams further comprises determining an origin of voice of at least one of the plurality of participants.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2014
From: JAISWAL, PEEYUSH; NARAYAN, NAVEEN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 031936/0503 →
Continuity (3)
Continuation 13901098 · May 23, 2013
Continuation 12246056 · Oct 6, 2008
Related Publication 20140119520A1 · May 1, 2014