IP Library › Granted Patent US 10,585,991
Granted Patent B2
US 10,585,991 · App. 15/637,831 · Granted Mar 10, 2020

Virtual assistant for generating personalized responses within a communication session

Inventors: Adi Miller (Herzliya, IL); Shira Weinberg (Herzliya, IL); Haim Somech (Herzliya, IL); Hen Fitoussi (Ramat HaSharon, IL)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06F17/279G06Q10/04G06Q10/06G06Q50/01G10L15/183G10L15/22G10L15/26G06F2203/0381G10L15/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,585,991
App. No.
15/637,831
Granted
Mar 10, 2020
Kind
B2
Abstract

Intelligent agents (IA) for automatically generating responses to content within a communication session (CS) are disclosed. An IA is trained to target the responses to a user and the user's context within the CS. An IA receives CS content that includes natural language expressions encoding users' conversations and determines content features based on natural language models. The content features indicate intended semantics of the expressions. The IA identifies likely-relevant content to the targeted user, to generate a response for. Identifying such content includes determining a relevance of the content based on content features, a context of the CS, a user-interest model, and a content-relevance model. Identifying the likely-relevant content to respond to is based on the determined relevance of the content and relevance thresholds. Various responses to the identified portions of the content are automatically generated and provided based on a natural language response-generation model targeted to the user.

Claims (84)

1. A computerized system comprising:

one or more processors; and

computer storage memory having computer-executable instructions stored thereon which, when executed by the one or more processors, implement a method comprising:

receiving content that is exchanged within a communication session (CS), wherein the content includes one or more natural language expressions that encode a portion of a conversation carried out by a plurality of users participating in the CS;

determining a relevance of the content based on a user-interest model for a first user of the plurality of users and a content-relevance model for the first user;

identifying one or more likely-relevant portions of the content based on the relevance of the content and one or more relevance thresholds, wherein the one or more identified likely-relevant portions of the content are likely-relevant to the first user;

generating a response to the one or more likely-relevant portions of the content based on a response-generation model for the first user, the generated response configured for participation in the conversation in place of the first user; and

causing the generated response that is configured for participation in the conversation in place of the first user to be presented.

2. The system of claim 1 , wherein the method further comprises:

determining one or more content features based on the content and one or more natural language models, wherein the one or more content features indicate one or more intended semantics of the one or more natural language expressions;

determining the relevance of the content further based on the content features; and

generating the response to one or more likely-relevant portions further based on the one or more content features.

3. The system of claim 1 , wherein the method further comprises:

receiving metadata associated with the CS;

determining one or more contextual features of the CS based on the received metadata and a CS context model, wherein the one or more contextual features indicate a context of the conversation for the first user; and

generating the response to the one or more likely-relevant portions of the content further based on the one or more contextual features of the CS.

4. The system of claim 1 , wherein the method further comprises:

receiving user feedback based on the response to the one or more likely-relevant portions of the content; and

updating the response-generation model based on the user feedback.

5. The system of claim 4 , wherein the method further comprises:

determining one or more content-substance features based on the content and a content-substance model included in the one or more natural language models, wherein the one or more content-substance features indicate one or more topics discussed in the conversation; and

determining one or more content-style features based on the content and a content-style model included in the one or more natural language models, wherein the one or more content-style features indicate an emotion of at least one of the plurality of the users; and

generating the response to the one or more likely-relevant portions of the content further based on the one or more content-substance features and the one or more content-style features of the content.

6. The system of claim 1 , wherein the method further comprises:

determining one or more content-substance features to encode in the response based on other content-substance features encoded in the likely-relevant portions of the content;

determining one or more content-style features to encode in the response based on other content-style features encoded in the likely-relevant portions of the content; and

generating the response to the one or more likely-relevant portions of the content such that the response encodes the one or more content-substance features and the one or more content-style features.

7. The system of claim 1 , wherein the method further comprises:

when the system is operated in a semi-autonomous mode, providing the response to the one or more likely-relevant portions of the content to the first user; and

when the system is operated in an autonomous mode, providing the response to the one or more likely-relevant portions of the content to the CS.

8. A method comprising:

receiving content that is exchanged within a communication session (CS), wherein the content includes one or more natural language expressions that encode a portion of a conversation carried out by a plurality of users participating in the CS;

determining a relevance of the content based on a user-interest model for a first user of the plurality of users and a content-relevance model for the first user;

identifying one or more likely-relevant portions of the content based on the relevance of the content and one or more relevance thresholds, wherein the one or more identified likely-relevant portions of the content are likely-relevant to the first user;

generating a response to the one or more likely-relevant portions of the content based on a response-generation model for the first user, the generated response configured for participation in the conversation in place of the first user; and

causing the generated response that is configured for participation in the conversation in place of the first user to be presented.

9. The method of claim 8 , further comprising:

determining one or more content features based on the content and one or more natural language models, wherein the one or more content features indicate one or more intended semantics of the one or more natural language expressions;

determining the relevance of the content further based on the content features; and

generating the response to one or more likely-relevant portions further based on the one or more content features.

10. The method of claim 8 , further comprising:

receiving metadata associated with the CS;

determining one or more contextual features of the CS based on the received metadata and a CS context model, wherein the one or more contextual features indicate a context of the conversation for the first user; and

generating the response to the one or more likely-relevant portions of the content further based on the one or more contextual features of the CS.

11. The method of claim 8 , further comprising:

identifying a sub-portion of the one or more likely-relevant portions of the content based on the relevance of the content and an additional relevance threshold, wherein the identified sub-portion of the one or more portions of the content is highly-relevant to the first user;

generating a response to the highly-relevant content based on the response-generation model for the first user; and

providing a real-time notification of the identified highly-relevant content and the response to the highly-relevant content to the first user.

12. The method of claim 8 , further comprising:

determining one or more content-substance features based on the content and a content-substance model included in the one or more natural language models, wherein the one or more content-substance features indicate one or more topics discussed in the conversation; and

determining one or more content-style features based on the content and a content-style model included in the one or more natural language models, wherein the one or more content-style features indicate an emotion of at least one of the plurality of the users; and

generating the response to the one or more likely-relevant portions of the content further based on the one or more content-substance features and the one or more content-style features of the content.

13. The method of claim 12 , further comprising:

receiving user feedback based on the response to the one or more likely-relevant portions of the content; and

updating at least one of the response-generation model, the content-substance model, or the content-style model based on the user feedback.

14. The method of claim 8 , further comprising:

when operated in a semi-autonomous mode, providing the response to the one or more likely-relevant portions of the content to the first user; and

when operated in an autonomous mode, providing the response to the one or more likely-relevant portions of the content to the CS.

15. One or more computer-storage media having instructions stored thereon, wherein the instructions, when executed by a processor of a computing device, cause the computing device to perform actions including:

receiving content that is exchanged within a communication session (CS), wherein the content includes one or more natural language expressions that encode a portion of a conversation carried out by a plurality of users participating in the CS;

determining a relevance of the content based on a user-interest model for a first user of the plurality of users and a content-relevance model for the first user;

identifying one or more likely-relevant portions of the content based on the relevance of the content and one or more relevance thresholds, wherein the one or more identified likely-relevant portions of the content are likely-relevant to the first user;

generating a response to the one or more likely-relevant portions of the content based on a response-generation model for the first user, the generated response configured for participation in the conversation in place of the first user; and

causing the generated response that is configured for participation in the conversation in place of the first user to be presented.

16. The media of claim 15 , the actions further comprising:

receiving metadata associated with the CS;

determining one or more contextual features of the CS based on the received metadata and a CS context model, wherein the one or more contextual features indicate a context of the conversation for the first user; and

generating the response to the one or more likely-relevant portions of the content further based on the one or more contextual features of the CS.

17. The media of claim 15 , wherein the actions further comprise:

identifying a sub-portion of the one or more likely-relevant portions of the content based on the relevance of the content and an additional relevance threshold, wherein the identified sub-portion of the one or more portions of the content is highly-relevant to the first user;

generating a response to the highly-relevant content based on the response-generation model for the first user; and

providing a real-time notification of the identified highly-relevant content and the response to the highly-relevant content to the first user.

18. The media of claim 15 , wherein the actions further comprise:

determining one or more content-substance features based on the content and a content-substance model included in the one or more natural language models, wherein the one or more content-substance features indicate one or more topics discussed in the conversation; and

determining one or more content-style features based on the content and a content-style model included in the one or more natural language models, wherein the one or more content-style features indicate an emotion of at least one of the plurality of the users; and

generating the response to the one or more likely-relevant portions of the content further based on the one or more content-substance features and the one or more content-style features of the content.

19. The media of claim 15 , wherein the actions further comprise:

determining one or more content-substance features to encode in the response based on other content-substance features encoded in the likely-relevant portions of the content;

determining one or more content-style features to encode in the response based on other content-style features encoded in the likely-relevant portions of the content; and

generating the response to the one or more likely-relevant portions of the content such that the response encodes the one or more content-substance features and the one or more content-style features.

20. The actions of claim 15 , wherein the actions further comprise:

providing the response to the one more likely-relevant portions of the content to one or more other users that are separate from the first user and participating in the CS;

receiving user feedback, from at least one of the one or more other users, based on the response to the one or more likely-relevant portions of the content; and

updating the response-generation model based on the user feedback received from the at least one of the one or more other users.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2017
From: MILLER, ADI; WEINBERG, SHIRA; SOMECH, HAIM; FITOUSSI, HEN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 043075/0254 →
Continuity (1)
Related Publication 20190005021A1 · Jan 3, 2019
Cited By (4)
US 12,354,603 US 12,380,475 US 12,718,435 US 12,737,557