IP Library › Granted Patent US 12,406,668
Granted Patent B2
US 12,406,668 · App. 18/213,525 · Granted Sep 2, 2025

Network-based communication session copilot

Inventors: Xiao Yan Lu (Bellevue, WA); Amir Kantor (Haifa, IL); Ido Priness (Kiryat Ono, IL); Shiraz Jitendra Cupala (Snohomish, WA); Kevin Michael Carter (Atlanta, GA); Adi Miller (Ramat Hasharon, IL); Kumud Ranjan (Redmond, WA); Shyam Gupta (Surrey, CA); Gautam Jain (Surrey, CA); Yasemin Cenberoglu (Montreal, CA); Shai Ifrach (Yavne, IL); Shlomi Maliah (Rosh Hayin, IL); Gilad Gildin (Tel Aviv, IL); Ofek David (Tel-Aviv, IL); Eleonora Shtotland (Herzliya, IL); Jaime Teevan (Bellevue, WA); Matthew Jonathan Gardner (Irvine, CA); Lan Ye (Issaquah, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/22G10L15/063G10L2015/221G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,668
App. No.
18/213,525
Granted
Sep 2, 2025
Kind
B2
Abstract

A system for providing a personalized assistant within a network-based communication session includes a processor and a memory storage device storing instructions. The system determines when a first communication session participant joins the network-based communication session after a threshold duration of time subsequent to the start time of the session. Upon determining the first participant has joined, the system obtains content associated with the session and creates request data for a pre-trained generative language model. The request data includes an instruction requesting a predetermined number of suggested utterances not present in the content, each utterance relating to one or more topics corresponding to the content. The system transforms the request data to a command based on a command template and provides the command to the generative language model. The system receives a response from the model, including the predetermined number of suggested utterances, and presents them to the communication session participant in a graphical user interface while the session is in session.

Claims (91)

1. A system providing a personalized assistant within a network-based communication session, the system comprising:

a processor; and

a memory storage device storing instructions thereon, which, when executed by the processor, cause the system to perform operations comprising:

determining a first communication session participant has joined a network-based communication session after a threshold duration of time subsequent to the start time of the network-based communication session; and

responsive to determining the first communication session participant has joined the network-based communication session after the threshold duration of time:

obtaining content associated with the network-based communication session, the content originating during a window of time spanning the start time of the network-based communication session and the time at which the communication session participant joined the network-based communication session;

creating request data for a pre-trained generative language model based upon a portion of the content associated with the network-based communication session and an instruction requesting, as output, a predetermined number of suggested utterances not present in the portion of the content, each utterance relating to one or more topics corresponding to the portion of the content;

transforming the request data to a command based upon a command template;

providing the command to a pre-trained generative language model with the request data;

receiving a response from the generative language model, the response including the predetermined number of suggested utterances; and

causing one or more of the predetermined number of suggested utterances to be presented to the communication session participant in a graphical user interface of the network-based communication session, while the network-based communication session is in session.

2. The system of claim 1 , wherein the instruction requesting, as output, the predetermined number of suggested utterances not present in the portion of the content includes a request to generate questions that have not already been asked by another communication session participant, or a request to generate questions that exclude questions that have already been asked by another communication session participant.

3. The system of claim 1 , wherein the content associated with the network-based communication session is structured as a communication session transcript having a plurality of chronologically ordered content items, wherein each content item in the plurality of chronologically ordered content items represents a communication made by a communication session participant and includes data indicating the name of the communication session participant who made the communication, wherein the instruction to generate the predetermined number of suggested utterances not present in the portion of the content includes a request to include the name of a communication session participant to whom a question should be directed.

4. The system of claim 1 , wherein the memory storage device is storing instructions, which, when executed by the one or more processors, cause the system to perform additional operations comprising:

providing as a second input to the pre-trained generative language model an instruction to generate a summary description of the network-based communication session, based on the content;

receiving as output from the generative language model a summary description of the network-based communication session; and

causing the summary description of the network-based communication session to be presented to the communication session participant in a graphical user interface of the network-based communication session.

5. The system of claim 1 , wherein the memory storage device is storing instructions, which, when executed by the one or more processors, cause the system to perform additional operations comprising:

segmenting the content into a plurality of segments, each segment in the plurality of segments having a size that is based on a maximum input size requirement of the pre-trained generative language model;

using the generative language model to generate a summary description of the network-based communication session, by:

for each segment of the plurality of segments, providing as input to the pre-trained generative language model content from the segment, and ii) an instruction to generate a summary description of the network-based communication session, based on the content;

receiving as output from the generative language model a summary description of the network-based communication session, for each segment of the plurality of segments;

providing to the pre-trained generative language model a final input, the final input including i) the summary description of the network-based communication session, as output by the generative language model for each segment of the plurality of segments and ii) an instruction to generate an overall summary description of the network-based communication session, based on the summary description of the network-based communication session as output by the generative language model for each segment of the plurality of segments;

responsive to providing the final input to the generative language model, receiving as output from the generative language model an overall summary description of the network-based communication session, based on the summary description of the network-based communication session as output by the generative language model for each segment of the plurality of segments; and

causing the overall summary description of the network-based communication session to be presented to the communication session participant in a graphical user interface of the network-based communication session.

6. The system of claim 1 , wherein the memory storage device is storing instructions, which, when executed by the one or more processors, cause the system to perform additional operations comprising:

segmenting the content into a plurality of segments, each segment in the plurality of segments having a size that is based on a maximum input size requirement of the pre-trained generative language model;

using the generative language model to generate a summary description of the network-based communication session, by:

for a first segment in the plurality of segments, providing as input to the pre-trained generative language model i) content from the first segment, and ii) an instruction to generate a summary description of the network-based communication session, based on the content from the first segment;

for each segment in the plurality of segments subsequent to the first segment, providing as input to the pre-trained generative language model i) the summary description of the network-based communication session output by the pre-trained generative language model based on a prior segment and content from the segment, and ii) an instruction to generate a summary description of the network-based communication session;

receiving as output from the generative language model a final summary description of the network-based communication session, based on the generative language model processing a final prompt for a last segment in the plurality of segments; and

causing the final summary description of the network-based communication session to be presented to the communication session participant in a graphical user interface of the network-based communication session.

7. The system of claim 1 , wherein the content associated with the in-session network-based communication session is structured as a communication session transcript having a plurality of chronologically ordered content items, wherein each content item in the plurality of chronologically ordered content items represents a communication made by a communication session participant and includes data indicating the name of the communication session participant who made the communication, wherein the memory storage device is storing instructions, which, when executed by the one or more processors, cause the system to perform additional operations comprising:

determine existence of a specific type of relationship between the first communication session participant and a second communication session participant;

extracting from the content one or more content items representing a communication made by the second communication session participant;

providing as input to a pre-trained generative language model, i) the one or more extracted content items, and ii) an instruction to generate a summary description of communications made by the second communication session participant;

receiving as output from the generative language model the summary description of communications made by the second communication session participant; and

causing the summary description of communications made by the second communication session participant to be presented to the first communication session participant in a graphical user interface of the network-based communication session.

8. The system of claim 1 , wherein the pre-trained generative language model has been fine-tuned, using a supervised learning technique, to generate suggested utterances not present in a portion of content based on a conversation of communication session participants as expressed in content, wherein a training dataset used in fine-tuning the pre-trained generative language model includes a plurality of instances of a communication session transcript having a plurality of chronologically ordered content items, each content item in the plurality of chronologically ordered content items representing a communication made by a communication session participant, wherein one or more communications have been labeled as questions, statements or opinions.

9. A method for providing a personalized assistant within a network-based communication session, the method comprising:

determining a first communication session participant has joined a network-based communication session after a threshold duration of time subsequent to the start time of the network-based communication session; and

responsive to determining the first communication session participant has joined the network-based communication session after the threshold duration of time:

obtaining content associated with the network-based communication session, the content originating during a window of time spanning the start time of the network-based communication session and the time at which the communication session participant joined the network-based communication session;

creating request data for a pre-trained generative language model based upon a portion of the content associated with the network-based communication session and an instruction requesting, as output, a predetermined number of suggested utterances not present in the portion of the content, each utterance relating to one or more topics corresponding to the portion of the content;

transforming the request data to a command based upon a command template;

providing the command to a pre-trained generative language model with the request data;

receiving a response from the generative language model, the response including the predetermined number of suggested utterances; and

causing one or more of the predetermined number of suggested utterances to be presented to the communication session participant in a graphical user interface of the network-based communication session, while the network-based communication session is in session.

10. The method of claim 9 , wherein the instruction requesting, as output, the predetermined number of suggested utterances not present in the portion of the content includes a request to generate questions that have not already been asked by another communication session participant, or a request to generate questions that exclude questions that have already been asked by another communication session participant.

11. The method of claim 9 , wherein the content associated with the network-based communication session is structured as a communication session transcript having a plurality of chronologically ordered content items, wherein each content item in the plurality of chronologically ordered content items represents a communication made by a communication session participant and includes data indicating the name of the communication session participant who made the communication, wherein the instruction to generate the predetermined number of suggested utterances not present in the portion of the content includes a request to include the name of a communication session participant to whom a question should be directed.

12. The method of claim 9 , further comprising:

providing as a second input to the pre-trained generative language model an instruction to generate a summary description of the network-based communication session, based on the content;

receiving as output from the generative language model a summary description of the network-based communication session; and

causing the summary description of the network-based communication session to be presented to the communication session participant in a graphical user interface of the network-based communication session.

13. The method of claim 9 , wherein the memory storage device is storing instructions, which, when executed by the one or more processors, cause the system to perform additional operations comprising:

segmenting the content into a plurality of segments, each segment in the plurality of segments having a size that is based on a maximum input size requirement of the pre-trained generative language model;

using the generative language model to generate a summary description of the network-based communication session, by:

for each segment of the plurality of segments, providing as input to the pre-trained generative language model content from the segment, and ii) an instruction to generate a summary description of the network-based communication session, based on the content;

receiving as output from the generative language model a summary description of the network-based communication session, for each segment of the plurality of segments;

providing to the pre-trained generative language model a final input, the final input including i) the summary description of the network-based communication session, as output by the generative language model for each segment of the plurality of segments and ii) an instruction to generate an overall summary description of the network-based communication session, based on the summary description of the network-based communication session as output by the generative language model for each segment of the plurality of segments;

responsive to providing the final input to the generative language model, receiving as output from the generative language model an overall summary description of the network-based communication session, based on the summary description of the network-based communication session as output by the generative language model for each segment of the plurality of segments; and

causing the overall summary description of the network-based communication session to be presented to the communication session participant in a graphical user interface of the network-based communication session.

14. The system of claim 9 , further comprising:

segmenting the content into a plurality of segments, each segment in the plurality of segments having a size that is based on a maximum input size requirement of the pre-trained generative language model:

using the generative language model to generate a summary description of the network-based communication session, by:

for a first segment in the plurality of segments, providing as input to the pre-trained generative language model i) content from the first segment, and ii) an instruction to generate a summary description of the network-based communication session, based on the content from the first segment;

for each segment in the plurality of segments subsequent to the first segment, providing as input to the pre-trained generative language model i) the summary description of the network-based communication session output by the pre-trained generative language model based on a prior segment and content from the segment, and ii) an instruction to generate a summary description of the network-based communication session;

receiving as output from the generative language model a final summary description of the network-based communication session, based on the generative language model processing a final prompt for a last segment in the plurality of segments; and

causing the final summary description of the network-based communication session to be presented to the communication session participant in a graphical user interface of the network-based communication session.

15. The method of claim 9 , wherein the content associated with the in-session network-based communication session is structured as a communication session transcript having a plurality of chronologically ordered content items, wherein each content item in the plurality of chronologically ordered content items represents a communication made by a communication session participant and includes data indicating the name of the communication session participant who made the communication, wherein the method further comprises:

determining existence of a specific type of relationship between the first communication session participant and a second communication session participant;

extracting from the content one or more content items representing a communication made by the second communication session participant;

providing as input to a pre-trained generative language model, i) the one or more extracted content items, and ii) an instruction to generate a summary description of communications made by the second communication session participant;

receiving as output from the generative language model the summary description of communications made by the second communication session participant; and

causing the summary description of communications made by the second communication session participant to be presented to the first communication session participant in a graphical user interface of the network-based communication session.

16. The method of claim 9 , wherein the pre-trained generative language model has been fine-tuned, using a supervised learning technique, to generate suggested utterances not present in a portion of content based on a conversation of communication session participants as expressed in content, wherein a training dataset used in fine-tuning the pre-trained generative language model includes a plurality of instances of a communication session transcript having a plurality of chronologically ordered content items, each content item in the plurality of chronologically ordered content items representing a communication made by a communication session participant, wherein one or more communications have been labeled as questions, statements or opinions.

17. A system providing a personalized assistant within a network-based communication session, the system comprising:

means for determining a first communication session participant has joined a network-based communication session after a threshold duration of time subsequent to the start time of the network-based communication session; and

responsive to determining the first communication session participant has joined the network-based communication session after the threshold duration of time:

means for obtaining content associated with the network-based communication session, the content originating during a window of time spanning the start time of the network-based communication session and the time at which the communication session participant joined the network-based communication session;

means for creating request data for a pre-trained generative language model based upon a portion of the content associated with the network-based communication session and an instruction requesting, as output, a predetermined number of suggested utterances not present in the portion of the content, each utterance relating to one or more topics corresponding to the portion of the content;

means for transforming the request data to a command based upon a command template;

means for providing the command to a pre-trained generative language model with the request data;

means for receiving a response from the generative language model, the response including the predetermined number of suggested utterances; and

means for causing one or more of the predetermined number of suggested utterances to be presented to the communication session participant in a graphical user interface of the network-based communication session, while the network-based communication session is in session.

18. The system of claim 17 , wherein the instruction requesting, as output, the predetermined number of suggested utterances not present in the portion of the content includes a request to generate questions that have not already been asked by another communication session participant, or a request to generate questions that exclude questions that have already been asked by another communication session participant.

19. The system of claim 17 , wherein the content associated with the network-based communication session is structured as a communication session transcript having a plurality of chronologically ordered content items, wherein each content item in the plurality of chronologically ordered content items represents a communication made by a communication session participant and includes data indicating the name of the communication session participant who made the communication, wherein the instruction to generate the predetermined number of suggested utterances not present in the portion of the content includes a request to include the name of a communication session participant to whom a question should be directed.

20. The system of claim 17 , further comprising:

means for providing as a second input to the pre-trained generative language model an instruction to generate a summary description of the network-based communication session, based on the content;

means for receiving as output from the generative language model a summary description of the network-based communication session; and

means for causing the summary description of the network-based communication session to be presented to the communication session participant in a graphical user interface of the network-based communication session.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2024
From: LU, XIAO YAN; KANTOR, AMIR; PRINESS, IDO; CUPALA, SHIRAZ JITENDRA; CARTER, KEVIN MICHAEL; MILLER, ADI; RANJAN, KUMUD; GUPTA, SHYAM; JAIN, GAUTAM; CENBEROGLU, YASEMIN; IFRACH, SHAI; MALIAH, SHLOMI; GILDIN, GILAD; DAVID,, OFEK; SHTOTLAND, ELEONORA; TEEVAN, JAIME; GARDNER, MATTHEW JONATHAN; YE, LAN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 066611/0848 →
Continuity (2)
Provisional Application 63448624 · Feb 27, 2023
Related Publication 20240290331A1 · Aug 29, 2024
References Cited (73)
US 7886012B2 · Bedi et al. · 2011 [cited by applicant]
US 8429146B2 · Shen et al. · 2013 [cited by applicant]
US 8560562B2 · Kanefsky et al. · 2013 [cited by applicant]
US 8700653B2 · Hansson et al. · 2014 [cited by applicant]
US 9646612B2 · Sharifi · 2017 [cited by examiner]
US 9756286B1 · Faulkner · 2017 [cited by applicant]
US 10540971B2 · Kumar · 2020 [cited by applicant]
US 10594757B1 · Shevchenko et al. · 2020 [cited by applicant]
US 10810897B2 · Dechu et al. · 2020 [cited by applicant]
US 10878033B2 · Ahmed et al. · 2020 [cited by applicant]
US 10878805B2 · Agarwal et al. · 2020 [cited by applicant]
US 10911718B2 · Nassar · 2021 [cited by applicant]
US 11017778B1 · Thomson et al. · 2021 [cited by applicant]
US 11042538B2 · Braundmeier · 2021 [cited by applicant]
US 11483273B2 · Mahmoud et al. · 2022 [cited by applicant]
US 11537674B2 · Galimovich · 2022 [cited by applicant]
US 11729009B1 · Religa · 2023 [cited by applicant]
US 11736423B2 · Wang · 2023 [cited by applicant]
US 11908477B2 · Embar · 2024 [cited by examiner]
US 12165646B2 · Springer · 2024 [cited by examiner]
US 20070288559A1 · Parsadayan et al. · 2007 [cited by applicant]
US 20090165000A1 · Gyorfi · 2009 [cited by examiner]
US 20130198159A1 · Hendry · 2013 [cited by applicant]
US 20160006981A1 · Bauman et al. · 2016 [cited by applicant]
US 20160210602A1 · Siddique · 2016 [cited by applicant]
US 20180165583A1 · Guiver et al. · 2018 [cited by applicant]
US 20180308473A1 · Scholar · 2018 [cited by applicant]
US 20190132265A1 · Nowak-przygodzki et al. · 2019 [cited by applicant]
US 20190341050A1 · Diamant et al. · 2019 [cited by applicant]
US 20200005784A1 · Vadackupurath Mani et al. · 2020 [cited by applicant]
US 20200184956A1 · Agarwal et al. · 2020 [cited by applicant]
US 20200186482A1 · Johnson, III et al. · 2020 [cited by applicant]
US 20200243095A1 · Adlersberg et al. · 2020 [cited by applicant]
US 20200356237A1 · Moran et al. · 2020 [cited by applicant]
US 20200410998A1 · Bar-on et al. · 2020 [cited by applicant]
US 20210056860A1 · Fahrendorff · 2021 [cited by examiner]
US 20210058263A1 · Fahrendorff · 2021 [cited by examiner]
US 20210224336A1 · Bright · 2021 [cited by applicant]
US 20210249009A1 · Manjunath et al. · 2021 [cited by applicant]
US 20210358188A1 · Lebaredian et al. · 2021 [cited by applicant]
US 20210365500A1 · Gunaselara et al. · 2021 [cited by applicant]
US 20210390144A1 · B M S et al. · 2021 [cited by applicant]
US 20220036153A1 · O'malia et al. · 2022 [cited by applicant]
US 20220109585A1 · Asthana · 2022 [cited by applicant]
US 20220271963A1 · Shah et al. · 2022 [cited by applicant]
US 20220310080A1 · Qiu et al. · 2022 [cited by applicant]
US 20220414320A1 · Dolan · 2022 [cited by applicant]
US 20230033104A1 · Geddes · 2023 [cited by applicant]
US 20230058470A1 · Chandrashekar et al. · 2023 [cited by applicant]
US 20230074406A1 · Baeuml et al. · 2023 [cited by applicant]
US 20230091949A1 · Erasmus · 2023 [cited by applicant]
US 20230135179A1 · Mielke · 2023 [cited by applicant]
US 20230135703A1 · Weil · 2023 [cited by examiner]
US 20230154453A1 · Erdenee · 2023 [cited by applicant]
US 20230177015A1 · Madisetti · 2023 [cited by examiner]
US 20240048513A1 · Tessler · 2024 [cited by examiner]
US 20240273345A1 · Bharadwaj · 2024 [cited by applicant]
EP 1828934B1 · 2019 [cited by applicant]
JP 2022180282A · 2022 [cited by applicant]
WO 2020139865A1 · 2020 [cited by applicant]
“4 Ways You Can Use ChatGPT for Your Next Meeting”, Tactiq, Retrieved from the Internet URL: https://web.archive.org/web/20230204100721/https://tactiq.io/learn/4-ways-you-can-use-chatgpt-for-your-next-meeting, Feb. 4, 2… [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/016336, Jul. 11, 2024, 13 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/016338, Jul. 15, 2024, 13 pages. [cited by applicant]
Abdollahpour, et al., “Zero-shot Question Answering with Large Language Models in Python”, Retrieved from: https://medium.com/nlplanet/zero-shot-question-answering-with-large-language-models-in-python-9964c55c3b38, Aug.… [cited by applicant]
Arabzadeh, et al., “PREME: Preference-based Meeting Exploration through an Interactive Questionnaire”, In Repository of arXiv:2205.02370v1, Apr. 27, 2023, 12 Pages. [cited by applicant]
Asi, et al., “An End-to-End Dialogue Summarization System for Sales Calls”, In Proceedings of NAACL-HLT: Industry Track Papers, Jul. 10, 2022, pp. 45-53. [cited by applicant]
Briggs, James, “Fixing YouTube Search with OpenAI's Whisper”, Retrieved from: https://www.pinecone.io/learn/openai-whisper/, Retrieved Date: Mar. 31, 2023, 25 Pages. [cited by applicant]
Cavanell, et al., “Use Natural Language & Prompts with AI Models | Azure OpenAI Service”, Retrieved from: https://techcommunity.microsoft.com/t5/microsoft-mechanics-blog/use-natural-language-amp-prompts-with-ai-models-a… [cited by applicant]
Greyling, Cobus, “Prompt Engineering, Text Generation & Large Language Models”, Retrieved from: https://cobusgreyling.medium.com/prompt-engineering-text-generation-large-language-models-3d90c527c6d5, Sep. 5, 2022, 18 Pa… [cited by applicant]
Kashyap, Prerna, “Coling 2022 Shared Task: LED Finteuning and Recursive Summary Generation for Automatic Summarization of Chapters from Novels”, In Proceedings of the 58th Annual Meeting of the Association for Computati… [cited by applicant]
Kim, Sung, “How to Get Around OpenAI GPT-3 Token Limits”, Retrieved from: https://blog.devgenius.io/how-to-get-around-openai-gpt-3-token-limits-b11583691b32, Retrieved Date: Mar. 31, 2023, 25 Pages. [cited by applicant]
Sb, et al., “Automatic Follow-up Question Generation for Asynchronous Interviews”, In Proceedings of the Workshop on Intelligent Information Processing and Natural Language Generation, Sep. 2020, 11 Pages. [cited by applicant]
Notice of Allowance mailed on Jul. 22, 2025, in U.S. Appl. No. 17/213,511, 18 pages. [cited by applicant]