IP Library Granted Patent US 11,138,978
Granted Patent B2
US 11,138,978 · App. 16/521,537 · Granted Oct 5, 2021

Topic mining based on interactionally defined activity sequences

Inventors: Margaret Helen Szymanski (Santa Clara, CA); Lei Huang (Mountain View, CA); Robert John Moore (San Jose, CA); Raphael Arar (San Jose, CA); Shun Jiang (San Jose, CA); Guangjie Ren (Belmont, CA); Eric Liu (Santa Clara, CA); Pawan Chowdhary (San Jose, CA); Chung-hao Tan (San Jose, CA); Sunhwan Lee (San Mateo, CA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G10L15/26G06N3/0445G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,138,978
App. No.
16/521,537
Granted
Oct 5, 2021
Kind
B2
Abstract

A method and system of automatically identifying topics of a conversation are provided. An electronic data package comprising a sequence of utterances between conversation entities is received by a computing device. Each utterance is classified to a corresponding social action. One or more utterances in the sequence are grouped into a segment based on a deep learning model. A similarity of topics between adjacent segments is determined. Upon determining that the similarity is above a predetermined threshold, the adjacent segments are grouped together. A transcript of the conversation including the grouping of the adjacent segments is stored in a memory.

Claims (48)

1. A computing device comprising:

a processor;

a network interface coupled to the processor to enable communication over a network;

an interaction engine configured to perform acts comprising:

receiving an electronic data package comprising a sequence of utterances between conversation entities, over the network;

classifying each utterance to a corresponding social action;

identifying transition boundaries based on the social actions;

grouping one or more utterances between transition boundaries in the sequence to a segment based on a deep learning model;

determining a similarity of topics between adjacent segments; and

upon determining that the similarity is above a predetermined threshold, grouping the adjacent segments together; and

storing a transcript of the conversation including the grouping of the adjacent segments in the memory.

2. The computing device of claim 1 , wherein the electronic data package comprises raw audio data.

3. The computing device of claim 2 , wherein the interaction engine is further configured to perform acts, comprising: converting a speech of the raw audio data to text by way of natural language processing (NLP).

4. The computing device of claim 2 , wherein the interaction engine is further configured to perform acts, comprising:

identifying a duration of silences between the utterances; and

including each duration in the transcript.

5. The computing device of claim 1 , wherein the classification of each utterance to a corresponding social action comprises using concept expansion to identify the social action.

6. The computing device of claim 1 , wherein each grouping of the one or more utterances into a segment is based on a common focused topic.

7. The computing device of claim 6 , wherein the topic is broadened for each iteration of merging the sequence of adjacent segments.

8. The computing device of claim 1 , wherein each segment comprises a sequence of adjacent utterances between transition boundaries.

9. The computing device of claim 1 , wherein the interaction engine is further configured to perform acts, comprising: removing out utterances that are not interactionally defined to reduce the computational complexity and storage requirements of the computing device.

10. The computing device of claim 1 , wherein the deep learning model is created by the computing device during a training phase, comprising:

receiving historical user interaction logs between conversation entities;

for each interaction log:

using concept expansion for a creation of a specialized dictionary for identifying social actions associated with each utterance in the interaction log;

classifying each utterance into a corresponding social action; and

grouping one or more utterances associated with social actions in a predetermined window in a sequence of the interaction log to a transition boundary.

11. The computing device of claim 10 , wherein the deep learning model is a sequential recurrent neural network (RNN) having a predetermined window.

12. The computing device of claim 1 , wherein the deep learning model is a many to one sequential recurrent neural network (RNN).

13. A non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, causes a computer device to carry out a method of determining topics of a conversation, the method comprising:

receiving an electronic data package comprising a sequence of utterances between conversation entities;

classifying each utterance to a corresponding social action;

identifying transition boundaries based on the social actions; grouping one or more utterances between transition boundaries in the sequence to a segment based on a deep learning model;

determining a similarity of topics between adjacent segments; and

upon determining that the similarity is above a predetermined threshold, grouping the adjacent segments together; and

storing a transcript of the conversation including the grouping of the adjacent segments in the memory.

14. The non-transitory computer readable storage medium of claim 13 , wherein the electronic data package comprises raw audio data.

15. The non-transitory computer readable storage medium of claim 14 , further comprising: converting a speech of the raw audio data to text by way of natural language processing (NLP).

16. The non-transitory computer readable storage medium of claim 13 , wherein the classification of each utterance to a corresponding social action comprises using concept expansion to identify the social action.

17. The non-transitory computer readable storage medium of claim 13 , wherein each grouping of the one or more utterances into a segment is based on a common focused topic.

18. The non-transitory computer readable storage medium of claim 17 , wherein the topic the predetermined threshold is replaced with has a more loose threshold for each iteration of merging the sequence of adjacent segments.

19. The non-transitory computer readable storage medium of claim 13 , wherein the deep learning model is created during a training phase, comprising:

receiving historical user interaction logs between conversation entities;

for each interaction log:

using concept expansion for a creation of a specialized dictionary for identifying social actions associated with each utterance in the interaction log;

classifying each utterance of the interaction log into a corresponding social action; and

grouping one or more utterances in a predetermined window in a sequence of the interaction log to a transition boundary.

20. The non-transitory computer readable storage medium of claim 19 , wherein the deep learning model is a many to one sequential recurrent neural network (RNN) having a predetermined window.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2019
From: SZYMANSKI, MARGARET HELEN; HUANG, LEI; MOORE, ROBERT JOHN; ARAR, RAPHAEL; JIANG, SHUN; REN, GUANGJIE; LIU, ERIC; CHOWDHARY, PAWAN; TAN, CHUNG-HAO; LEE, SUNHWAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 049853/0806 →
Continuity (1)
Related Publication 20210027783A1 · Jan 28, 2021
Cited By (1)
US 12,284,148