IP Library Granted Patent US 12,596,889
Granted Patent B2
US 12,596,889 · App. 18/675,840 · Granted Apr 7, 2026

Generation of natural language (NL) based summaries using a large language model (LLM) and subsequent modification thereof for attribution

Inventors: Shrestha Basu Mallick (San Francisco, CA); Owen Lewis (Stanford, CA); Jaclyn Konzelmann (Mountain View, CA); Christina Yang Choi (San Jose, CA); James Freedman (Los Angeles, CA); Jonathan Malmaud (Campbell, CA); Xin Xie (San Jose, CA); Brian Carver (Emeryville, CA)
Assignee: GOOGLE LLC
G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,889
App. No.
18/675,840
Granted
Apr 7, 2026
Kind
B2
Abstract

Implementations described herein relate to attribution of a natural language (NL) based summary generated using a large language model (LLM). Processor(s) of a system can: receive NL based input associated with a client device, generate the NL based summary using the LLM, and process the NL based summary to determine whether a NL based summary segment of the NL based summary matches a dataset segment of a dataset that was utilized to initially train the LLM and/or to fine-tune the LLM. Further, the processor(s) can, in response to determining that the NL based summary segment matches the dataset segment, modify the NL based summary segment of the NL based summary to generate a modified NL based summary. Moreover, the processor(s) can cause the modified NL based summary to be rendered at the client device. The attribution of the NL based summary can be provided as a service to various third-parties.

Claims (64)

1 . A method implemented by one or more processors, the method comprising:

receiving, from a third-party, a third-party dataset;

fine-tuning, based on the third-party dataset, a large language model (LLM) that was initially trained on a LLM dataset;

receiving natural language (NL) based input associated with a client device of the third-party;

generating, based on processing the NL based input using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset, a NL based response that is responsive to the NL based input;

processing the NL based response to determine whether a NL based response segment, of the NL based response, matches a third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM, based on comparing a plurality of NL based response segments, of the NL based response, to a plurality of third-party dataset segments, of the third-party dataset that was utilized to fine-tune the LLM; and

in response to determining that a NL based response segment of the NL based response matches a third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM:

causing, to be rendered at the client device:

an indication of the NL based response segment of the NL based response that matches the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM; and

a hyperlink, included in the NL based response segment of the NL based response, to a document that includes the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM, wherein the hyperlink, when selected, causes the client device to navigate to the document that includes the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM.

2 . The method of claim 1 , wherein the third-party dataset comprises a corpus of access-restricted data that is specific to the third-party.

3 . The method of claim 1 , further comprising:

prior to receiving the NL based input:

storing, in association with each of the third-party dataset segments and in a third-party index, corresponding third-party metadata that indicates a corresponding document that includes one or more corresponding third-party dataset segments.

4 . The method of claim 3 , further comprising:

receiving, from the third-party, a third-party token that is specific to the third-party and that enables access to the third-party index in response to receiving the NL based input.

5 . The method of claim 1 , wherein the plurality of third-party dataset segments are stored in a third-party index.

6 . The method of claim 5 , wherein determining that a NL based response segment, of the NL based response, matches a third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM, comprises:

determining that given NL based response alphanumeric characters, of a given NL based response segment, match given third-party dataset alphanumeric characters, of a given third-party dataset segment.

7 . The method of claim 1 , wherein processing the NL based response to determine whether a NL based response segment, of the NL based response, matches a third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM, is restricted to processing the NL based response to determine whether the NL based response segment, of the NL based response, matches the third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM.

8 . The method of claim 1 , wherein generating the NL based response that is responsive to the NL based input based on processing the NL based input using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset comprises:

processing, using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset, the NL based input to generate LLM output; and

generating, based on the LLM output, the NL based response that is responsive to the NL based input.

9 . The method of claim 8 , further comprising:

processing, using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset, and along with the NL based input, additional third-party data received from the third-party to prime the LLM prior to generating the NL based response that is responsive to the NL based input.

10 . The method of claim 9 , wherein the additional third-party data received from the third-party that is utilized to prime the LLM prior to generating the NL based response that is responsive to the NL based input comprises: recently accessed third-party documents or recently contacted third-party contacts.

11 . A system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the at least one processor to be operable to:

receive, from a third-party, a third-party dataset;

fine-tune, based on the third-party dataset, a large language model (LLM) that was initially trained on a LLM dataset;

receive natural language (NL) based input associated with a client device of the third-party;

generate, based on processing the NL based input using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset, a NL based response that is responsive to the NL based input;

process the NL based response to determine whether a NL based response segment, of the NL based response, matches a third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM, based on comparing a plurality of NL based response segments, of the NL based response, to a plurality of third-party dataset segments, of the third-party dataset that was utilized to fine-tune the LLM; and

in response to determining that a NL based response segment of the NL based response matches a third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM:

cause, to be rendered at the client device:

an indication of the NL based response segment of the NL based response that matches the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM; and

a hyperlink, included in the NL based response segment of the NL based response, to a document that includes the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM, wherein the hyperlink, when selected, causes the client device to navigate to the document that includes the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM.

12 . The system of claim 11 , wherein the third-party dataset comprises a corpus of access-restricted data that is specific to the third-party.

13 . The system of claim 11 , wherein the instructions further cause the at least one processor to be operable to:

prior to receiving the NL based input:

store, in association with each of the third-party dataset segments and in a third-party index, corresponding third-party metadata that indicates a corresponding document that includes one or more corresponding third-party dataset segments.

14 . The system of claim 13 , wherein the instructions further cause the at least one processor to be operable to:

receive, from the third-party, a third-party token that is specific to the third-party and that enables access to the third-party index in response to receiving the NL based input.

15 . The system of claim 11 , wherein the plurality of third-party dataset segments are stored in a third-party index.

16 . The system of claim 15 , wherein the instructions to determine that a NL based response segment, of the NL based response, matches a third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM, comprise instructions to:

determine that given NL based response alphanumeric characters, of a given NL based response segment, match given third-party dataset alphanumeric characters, of a given third-party dataset segment.

17 . The system of claim 11 , wherein processing the NL based response to determine whether a NL based response segment, of the NL based response, matches a third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM, is restricted to processing the NL based response to determine whether the NL based response segment, of the NL based response, matches the third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM.

18 . The system of claim 11 , wherein the instructions to generate the NL based response that is responsive to the NL based input based on processing the NL based input using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset comprise instructions to:

process, using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset, the NL based input to generate LLM output; and

generate, based on the LLM output, the NL based response that is responsive to the NL based input.

19 . The system of claim 18 , wherein the instructions further cause the at least one processor to be operable to:

process, using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset, and along with the NL based input, additional third-party data received from the third-party to prime the LLM prior to generating the NL based response that is responsive to the NL based input,

wherein the additional third-party data received from the third-party that is utilized to prime the LLM prior to generating the NL based response that is responsive to the NL based input comprises: recently accessed third-party documents or recently contacted third-party contacts.

20 . A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations comprising:

receiving, from a third-party, a third-party dataset;

fine-tuning, based on the third-party dataset, a large language model (LLM) that was initially trained on a LLM dataset;

receiving natural language (NL) based input associated with a client device of the third-party;

generating, based on processing the NL based input using the LLM that was initially trained on the LLM dataset and that was fine-tuned on the third-party dataset, a NL based response that is responsive to the NL based input;

processing the NL based response to determine whether a NL based response segment, of the NL based response, matches a third-party dataset segment, of the third-party dataset that was utilized to fine-tune the LLM, based on comparing a plurality of NL based response segments, of the NL based response, to a plurality of third-party dataset segments, of the third-party dataset that was utilized to fine-tune the LLM; and

in response to determining that a NL based response segment of the NL based response matches a third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM:

causing, to be rendered at the client device:

an indication of the NL based response segment of the NL based response that matches the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM; and

a hyperlink, included in the NL based response segment of the NL based response, to a document that includes the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM, wherein the hyperlink, when selected, causes the client device to navigate to the document that includes the third-party dataset segment of the third-party dataset that was utilized to fine-tune the LLM.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2024
From: MALLICK, SHRESTHA BASU; LEWIS, OWEN; KONZELMANN, JACLYN; CHOI, CHRISTINA YANG; FREEDMAN, JAMES; MALMAUD, JONATHAN; XIE, XIN; CARVER, BRIAN
To: GOOGLE LLC
Reel/Frame 067693/0435 →
Continuity (3)
Continuation 18241731 · Sep 1, 2023
Provisional Application 63447234 · Feb 21, 2023
Related Publication 20240320445A1 · Sep 26, 2024
References Cited (16)
US 11704351B1 · Soleimani · 2023 [cited by applicant]
US 11763069B2 · Makino · 2023 [cited by applicant]
US 12046155B1 · Ghosh · 2024 [cited by examiner]
US 20180113922A1 · Gulwani · 2018 [cited by applicant]
US 20190325066A1 · Krishna · 2019 [cited by applicant]
US 20210295822A1 · Tomkins · 2021 [cited by examiner]
US 20210374338A1 · Shrivastava · 2021 [cited by examiner]
US 20220269853A1 · Itani · 2022 [cited by examiner]
US 20220318522A1 · Wolf · 2022 [cited by examiner]
US 20230020886A1 · Mahapatra · 2023 [cited by applicant]
US 20240185001A1 · Nagaraju · 2024 [cited by examiner]
US 20240281668A1 · Lev · 2024 [cited by examiner]
Gao, L. et al., “RARR: Researching and Revising What Language Models Say, Using Language Models”; Proceedings of the 61st Annual Meeting of The Association for Computational Linguistics (vol. 1: Long Papers); retrieved … [cited by applicant]
Dunn, A. et al., “Structured information extraction from complex scientific text with fine-tuned large language models”; arXiv.org., Cornell University; arXiv.2212.05238v1; 21 pages; dated Dec. 10, 2022. [cited by applicant]
Chung, H.W. et al., “Scaling Instruction-Finetuned Language Models”; arXiv.org, Cornell University; arXiv:2210.11416v5; 54 pages; dated Dec. 6, 2022. [cited by applicant]
European Patent Office; International Search Report and Written Opinion issued in Application No. PCT/US2023/034690; 13 pages; dated Feb. 5, 2024. [cited by applicant]