IP Library Granted Patent US 11,238,076
Granted Patent B2
US 11,238,076 · App. 16/852,483 · Granted Feb 1, 2022

Document enrichment with conversation texts, for enhanced information retrieval

Inventors: Haggai Roitman (Yoknea'm Elit, IL); Shai Erera (Gilon, IL); Doron Cohen (Gilon, IL); Yosi Mass (Ramat Gan, IL); Or Rivlin (Nesher, IL)
Assignee: International Business Machines Corporation
G06F16/3344G06F16/353G06F16/93H04L51/02H04L51/12H04L51/16H04L67/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,238,076
App. No.
16/852,483
Granted
Feb 1, 2022
Kind
B2
Abstract

A method including: Obtaining multiple conversation texts, one text per conversation, wherein each of the multiple conversation texts comprises: multiple messages authored by multiple parties, and a reference to an electronic document that provides resolution of a problem that is common to all the conversations. Calculating an importance score for each of the multiple messages of all the conversation texts. Clustering the multiple messages of all the conversation texts into multiple bins. Calculating an aggregated importance score for each of the multiple bins, based on the importance scores of the messages contained in the respective bin. Enriching (a) the electronic document, or (b) a record of the electronic document in an index of electronic documents, with at least some of the multiple bins and their aggregated importance scores, wherein the at least some of the multiple bins are added as fields to the electronic document or to the record.

Claims (56)

1. A method comprising operating at least one hardware processor to automatically:

obtain multiple conversation texts, one text per conversation, wherein each of the multiple conversation texts comprises:

multiple messages authored by multiple parties, and

a reference to an electronic document that provides resolution of a problem that is common to all the conversations;

calculate an importance score for each of the multiple messages of all the conversation texts;

cluster the multiple messages of all the conversation texts into multiple bins;

calculate an aggregated importance score for each of the multiple bins, based on the importance scores of the messages contained in the respective bin; and

enrich (a) the electronic document, or (b) a record of the electronic document in an index of electronic documents, with at least some of the multiple bins and their aggregated importance scores, wherein the at least some of the multiple bins are added as fields to the electronic document or to the record.

2. The method of claim 1 , further comprising, in an information retrieval system:

receiving a search query that comprises a weight for each of at least some of the fields;

ranking search results based on the weights comprised in the search query; and

returning the ranked search results.

3. The method of claim 2 , wherein:

the search query is received from a chatbot while the chatbot converses with a user;

the search query comprises one or more messages transmitted between the chatbot and the user; and

the returned search results comprise the reference to the electronic document, such that the chatbot is enabled to offer the reference to the electronic document to the user as a resolution of a problem presented by the user.

4. The method of claim 2 , wherein:

the search query is received from a chat system operated by a human agent, while the human agent converses with a user;

the search query comprises one or more messages transmitted between the human agent and the user; and

the returned search results comprise the reference to the electronic document, such that the human agent is enabled to offer the reference to the electronic document to the user as a resolution of a problem presented by the user.

5. The method of claim 1 , wherein the reference to the electronic document comprises a Uniform Resource Identifier (URI) of the electronic document.

6. The method of claim 1 , wherein the calculation of the importance score for each of the multiple messages is based on contents of each of the multiple messages.

7. The method of claim 1 , wherein the calculation of the importance score for each of the multiple messages is based on an order of the multiple messages inside the respective conversation text.

8. The method of claim 1 , wherein the clustering is based on contents of the multiple messages.

9. The method of claim 1 , wherein the enrichment of the electronic document or the record with the aggregated importance scores comprises:

implicitly adding the aggregated importance scores to the electronic document or the record, by—

naming,

numbering, or

orderly positioning

the added fields in a sequence reflecting their aggregated importance scores.

10. A system comprising:

(a) at least one hardware processor; and

(b) a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by said at least one hardware processor to:

obtain multiple conversation texts, one text per conversation, wherein each of the multiple conversation texts comprises:

multiple messages authored by multiple parties, and

a reference to an electronic document that provides resolution of a problem that is common to all the conversations;

calculate an importance score for each of the multiple messages of all the conversation texts;

cluster the multiple messages of all the conversation texts into multiple bins;

calculate an aggregated importance score for each of the multiple bins, based on the importance scores of the messages contained in the respective bin; and

enrich (a) the electronic document, or (b) a record of the electronic document in an index of electronic documents, with at least some of the multiple bins and their aggregated importance scores, wherein the at least some of the multiple bins are added as fields to the electronic document or to the record.

11. The system of claim 10 , wherein the reference to the electronic document comprises a Uniform Resource Identifier (URI) of the electronic document.

12. The system of claim 10 , wherein the calculation of the importance score for each of the multiple messages is based on contents of each of the multiple messages.

13. The system of claim 10 , wherein the calculation of the importance score for each of the multiple messages is based on an order of the multiple messages inside the respective conversation text.

14. The system of claim 10 , wherein the clustering is based on contents of the multiple messages.

15. A computer program product comprising a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by at least one hardware processor to:

obtain multiple conversation texts, one text per conversation, wherein each of the multiple conversation texts comprises:

multiple messages authored by multiple parties, and

a reference to an electronic document that provides resolution of a problem that is common to all the conversations;

calculate an importance score for each of the multiple messages of all the conversation texts;

cluster the multiple messages of all the conversation texts into multiple bins;

calculate an aggregated importance score for each of the multiple bins, based on the importance scores of the messages contained in the respective bin; and

enrich (a) the electronic document, or (b) a record of the electronic document in an index of electronic documents, with at least some of the multiple bins and their aggregated importance scores, wherein the at least some of the multiple bins are added as fields to the electronic document or to the record.

16. The computer program product of claim 15 , wherein the reference to the electronic document comprises a Uniform Resource Identifier (URI) of the electronic document.

17. The computer program product of claim 15 , wherein the calculation of the importance score for each of the multiple messages is based on contents of each of the multiple messages.

18. The computer program product of claim 15 , wherein the calculation of the importance score for each of the multiple messages is based on an order of the multiple messages inside the respective conversation text.

19. The computer program product of claim 15 , wherein the clustering is based on contents of the multiple messages.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2020
From: ROITMAN, HAGGAI; ERERA, SHAI; COHEN, DORON; MASS, YOSI; RIVLIN, OR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052435/0788 →
Continuity (1)
Related Publication 20210326369A1 · Oct 21, 2021
Cited By (2)
US 12,602,599 US 12,608,413