IP Library › Granted Patent US 11,189,283
Granted Patent B2
US 11,189,283 · App. 16/572,554 · Granted Nov 30, 2021

Freeform conversation writing assistant

Inventors: Tracy ThuyDuyen Tran (Seattle, WA); Daniel Parish (Seattle, WA); Ajitesh Kishore (Kirkland, WA); James Patrick Spotanski (Seattle, WA); Keri Diane Talbot (Sammamish, WA); Kiruthika Selvamani (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/22G06F40/56G09B7/00G10L2015/223G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,189,283
App. No.
16/572,554
Granted
Nov 30, 2021
Kind
B2
Abstract

In a device including a processor and a memory, the memory includes executable instructions that cause the processor to control the device to perform functions of receiving a user input to initiate a conversation session for generating an outline for a writing; generating a voice output asking a question regarding the writing; receiving a voice input from the user responding to the voice output; identifying, based on the received voice input, content of the voice input responding to the voice output; repeating, until a predetermined condition is met, the steps of generating the voice output, receiving the voice input, and identifying the content of the voice input, wherein the question asked via each voice output is generated in response to the content of the voice input responding to the preceding voice output; and generating, based on the content of the voice inputs, the outline for the writing.

Claims (68)

1. A system for generating an outline for a writing via a voice conversation session, comprising:

a processor; and

a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor, cause the processor to control the system to perform functions of:

receiving, from a user, a user input requesting to initiate a voice conversation session for generating an outline for a writing;

in response to receiving the user input, generating a voice output asking a question regarding the writing;

capturing a voice input from the user responding to the voice output;

inputting, to a deep learning (DL) engine, the captured voice input from the user, the DL engine trained with data from a plurality of previous voice conversation sessions to perform:

determine a relevancy of content of the captured voice input to the question asked in the voice output and an accuracy of the content of the captured voice input; and

generate, based on the determined relevancy and accuracy of the content of the voice input, a follow-up question regarding the writing;

extracting, from the DL engine, the follow-up question regarding the writing;

generating a follow-up voice output asking the follow-up question extracted from the DL engine;

repeating, until a predetermined condition is met, the functions of capturing the voice input, inputting the captured voice input to the DL engine, extracting the follow-up question from the DL engine, and generating the follow-up voice output; and

in response to determining that the predetermined condition is met, generating, based on the content of the captured voice inputs, the outline for the writing.

2. The system of claim 1 , wherein the instructions, when executed by the processor, further cause the processor to control the system to perform functions of:

extracting, from the DL engine, the determined relevancy indicating that the content of the captured voice input is not relevant to the question in the preceding voice output; and

generating another voice output indicating that the voice input is not relevant to the preceding voice output.

3. The system of claim 1 , wherein the instructions, when executed by the processor, further cause the processor to control the system to perform functions of:

extracting, from the DL engine, the determined accuracy indicating that the content of the captured voice input is not accurate; and

generating another voice output indicating that the voice input is not accurate.

4. The system of claim 1 , wherein:

for generating the voice output, the instructions, when executed by the processor, further cause the processor to control the system to perform a function of generating a first voice output asking a first question regarding a type of the writing, and

the DL engine is trained to determine, based on the captured voice input responding to the first voice output, the type of the writing.

5. The system of claim 4 , wherein, for generating the follow-up voice output, the instructions, when executed by the processor, further cause the processor to control the system to perform a function of generating, in response to determining the type of writing, a second voice output asking a second question regarding a length of the writing.

6. The system of claim 5 , wherein:

the instructions, when executed by the processor, further cause the processor to control the system to perform functions of capturing another voice input responding to the second voice output and related to the length of the writing, and

the DL engine is trained to perform:

determining, based on the captured voice input responding to the second voice output, the predetermined condition comprising an amount of information to be collected from the user; and

determining, based on the determined amount of information, whether the repeated functions of capturing the voice input, inputting the captured voice input to the DL engine, extracting the follow-up question from the DL engine and generating the follow-up voice output have met the predetermined condition.

7. The system of claim 1 , wherein for generating the voice output, the instructions, when executed by the processor, further cause the processor to control the system to perform a function of generating a first voice output asking a first question related to a subject of the writing, a reason for selecting the subject of the writing, or an example of the reason.

8. A non-transitory computer readable medium containing instructions which, when executed by a processor, cause a system to perform functions of operating a writing assistant bot for generating an outline for a writing via a voice conversation session, the functions comprising:

receiving, from a user, a user input requesting to initiate a voice conversation session for generating an outline for a writing;

in response to receiving the user input, generating a voice output asking a question regarding the writing;

capturing a voice input from the user responding to the voice output;

inputting, to a deep learning (DL) engine, the captured voice input from the user, the DL engine trained with data from a plurality of previous voice conversation sessions to perform:

determine a relevancy of content of the captured voice input to the question asked in the voice output and an accuracy of the content of the captured voice input; and

generate, based on the determined relevancy and accuracy of the content of the voice input, a follow-up question regarding the writing;

extracting, from the DL engine, the follow-up question regarding the writing;

generating a follow-up voice output asking the follow-up question extracted from the DL engine;

repeating, until a predetermined condition is met, the functions of capturing the voice input, inputting the captured voice input to the DL engine, extracting the follow-up question from the DL engine, and generating the follow-up voice output; and

in response to determining that the predetermined condition is met, generating, based on the content of the captured voice inputs, the outline for the writing.

9. A method of operating a system for generating an outline for a writing via a voice conversation session, the method comprising:

receiving, from a user, a user input requesting to initiate a voice conversation session for generating an outline for a writing;

in response to receiving the user input, generating a voice output asking a question regarding the writing;

capturing receiving a voice input from the user responding to the voice output;

inputting, to a deep learning (DL) engine, the captured voice input from the user, the DL engine trained with data from a plurality of previous voice conversation sessions to perform:

determine a relevancy of content of the captured voice input to the question asked in the voice output and an accuracy of the content of the captured voice input; and

generate, based on the determined relevancy and accuracy of the content of the voice input, a follow-up question regarding the writing;

extracting, from the DL engine, the follow-up question regarding the writing;

generating a follow-up voice output asking the follow-up question extracted from the DL engine;

repeating, until a predetermined condition is met, the steps of capturing the voice input, inputting the captured voice input to the DL engine, extracting the follow-up question from the DL engine, and generating the follow-up voice output; and

in response to determining that the predetermined condition is met, generating, based on the content of the captured voice inputs, the outline for the writing.

10. The method of claim 9 , further comprising:

extracting, from the DL engine, the determined relevancy indicating that the content of the captured voice input is not relevant to the question in the preceding voice output; and

generating another voice output indicating that the voice input is not relevant to the preceding voice output.

11. The method of claim 9 , further comprising:

extracting, from the DL engine, the determined accuracy indicating that the content of the captured voice input is not accurate; and

generating another voice output indicating that the voice input is not accurate.

12. The method of claim 9 , wherein:

generating the voice output comprises generating a first voice output asking a first question regarding a type of the writing, and

the DL engine is trained to determine, based on the captured voice input responding to the first voice output, the type of the writing.

13. The method of claim 12 , wherein generating the follow-up voice output comprises generating, in response to determining the type of writing, a second voice output asking a second question regarding a length of the writing.

14. The method of claim 13 , further comprising:

capturing another voice input responding to the second voice output and related to the length of the writing,

wherein the DL engine is trained to perform:

determining, based on the captured voice input responding to the second voice output, the predetermined condition comprising an amount of information to be collected from the user; and

determining, based on the determined amount of information, whether the repeated steps of capturing the voice input, inputting the captured voice input to the DL engine, extracting the follow-up question from the DL engine and generating the follow-up voice output have met the predetermined condition.

15. The method of claim 9 , wherein generating the voice output comprises generating a first voice output asking a first question related to a subject of the writing, a reason for selecting the subject of the writing, or an example of the reason.

16. The method of claim 9 , further comprising displaying, on display, the content of the captured voice input.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2019
From: TRAN, TRACY THUYDUYEN; PARISH, DANIEL; KISHORE, AJITESH; SPOTANSKI, JAMES PATRICK; TALBOT, KERI DIANE; SELVAMANI, KIRUTHIKA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 050406/0648 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2019
From: TRAN, TRACY THUYDUYEN; PARISH, DANIEL; KISHORE, AJITESH; SPOTANSKI, JAMES PATRICK; TALBOT, KERI DIANE; SELVAMANI, KIRUTHIKA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 050393/0979 →
Continuity (1)
Related Publication 20210082419A1 · Mar 18, 2021
Cited By (1)
US 12,277,384