IP Library Granted Patent US 12,664,377
Granted Patent B2
US 12,664,377 · App. 16/940,095 · Granted Jun 23, 2026

Text string summarization

Inventors: Neilesh Chorakhalikar (San Jose, CA); Manas Joshi (Pickering, CA); Arunachalam Arunachalam (Austin, TX); Baber M. Shaikh (Redmond, WA)
Assignee: NVIDIA Corporation
G06F40/58G06F9/454G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,377
App. No.
16/940,095
Filed
Jul 27, 2020
Granted
Jun 23, 2026
Kind
B2
Art Unit
2657
USPC
704/2
Abstract

Apparatuses, systems, and techniques to summarize a plurality of text strings. In at least one embodiment, one or more machine learning structures are used to generate a summary of a plurality of text strings based, at least in part, on an association between one or more words in a text string and one or more words with similar meaning.

Claims (38)

1 . One or more processors, comprising:

circuitry to:

use one or more neural networks to generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of the plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

2 . The one or more processors of claim 1 , wherein the circuitry is to normalize the plurality of text strings by replacing the one or more words in the plurality of text strings with the one or more keywords indicative of a topic, and to map the normalized plurality of text strings, using the one or more neural networks, to one or more points in a multidimensional space, wherein distance between points in the multidimensional space is indicative of similarity between text strings mapped to the points.

3 . The one or more processors of claim 2 , wherein the circuitry is to calculate a median of the cluster of encodings, and to generate the summary based, at least in part, on proximity of an encoding of the one or more encodings to the median.

4 . The one or more processors of claim 3 , wherein the circuitry is to evaluate, for suitability as the summary of the plurality of text strings, a text string mapped to the center point based, at least in part, on a conciseness criteria.

5 . The one or more processors of claim 4 , wherein the circuitry is to determine that no text string mapped to a point near a geometric median conforms to the conciseness criteria and, in response to the determination, generate the summary based, at least in part, on a frequency of keywords in the plurality of text strings.

6 . The one or more processors of claim 2 , wherein the one or more neural networks map the normalized plurality of text strings to the center point in the multidimensional space.

7 . The one or more processors of claim 1 , wherein the plurality of text strings comprise user-provided feedback and the one or more words with similar meaning are obtained from a list of keywords related to a subject to which the feedback is predicted to pertain.

8 . The one or more processors of claim 1 , wherein the circuitry is further to evaluate the plurality of text strings based, at least in part, on one or more text strings of the plurality of text strings having fewer than a threshold number of words.

9 . The one or more processors of claim 1 , wherein the summary is normalized by replacing at least a portion of words of the summary with similar words.

10 . A system, comprising:

one or more processors to generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of a plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

11 . The system of claim 10 , the one or more processors to cause the system to generate the keywords indicative of a topic based, at least in part, on a number of key words and a total number of words in one or more of the plurality of text strings.

12 . The system of claim 10 , the one or more processors to map the plurality of text strings to one or more points in a multidimensional space, wherein distance between points in the multidimensional space is indicative of a degree of similarity between semantic meanings of the text strings mapped to the points.

13 . The system of claim 12 , the one or more processors to generate the summary based at least in part on a median of a cluster of the encodings.

14 . The system of claim 13 , the one or more processors to evaluate one or more statements near the median based, at least in part, on a conciseness criteria.

15 . The system of claim 13 , wherein one of the plurality of text strings is selected as a summary, based at least in part, on proximity to the median and a conciseness criteria.

16 . The system of claim 13 , the one or more processors to determine that no point near the median conforms to a conciseness criteria and, in response to the determination, generate the summary based, at least in part, on keyword frequency in the text string.

17 . The system of claim 10 , wherein the plurality of text strings comprise user-provided feedback and the one or more words with similar meaning are obtained from a list of keywords related to a subject to which the feedback pertains.

18 . A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to use one or more machine learning structures to at least generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of the plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

19 . The machine-readable medium of claim 18 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to at least normalize the plurality of text strings by replacing the one or more words in the text strings with the one or more words with similar meaning, and to map the normalized plurality of text strings, using one or more neural networks, to one or more points in a multidimensional space.

20 . The machine-readable medium of claim 19 , wherein distance between points in the multidimensional space is indicative of similarity between text strings mapped to the points.

21 . The machine-readable medium of claim 19 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to calculate a median of the points in the multidimensional space.

22 . The machine-readable medium of claim 21 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to generate the summary of the plurality of text strings by at least evaluating one or more points nearest to the median according to a conciseness criteria.

23 . The machine-readable medium of claim 21 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to determine that no point near the median conforms to a conciseness criteria and, in response to the determination, generate the summary based, at least in part, on counting a frequency of keywords identified in the text strings.

24 . A system, comprising:

one or more computing devices that receive a plurality of text strings and use one or more machine learning structures to generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of the plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

25 . The system of claim 24 , the one or more computing devices to normalize the plurality of text strings based, at least in part, on the one or more words with similar meaning.

26 . The system of claim 25 , the one or more computing devices to map the normalized plurality of text strings to one or more points in a multidimensional space.

27 . The system of claim 26 , the one or more computing devices to generate the summary based, at least in part, on a geometric median of a cluster of the points and conciseness of points near the geometric median.

28 . The system of claim 26 , the one or more computing devices to determine that no text string mapped to a point near a geometric median of the one or more points conforms to a conciseness criteria and, in response to the determination, generate the summary based, at least in part, on frequency of keywords in the text string.

29 . The system of claim 24 , wherein the plurality of text strings comprise user-provided feedback.

30 . The system of claim 24 , wherein the one or more words with similar meaning are obtained from a list of keywords related to a service provided by the one or more computing devices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2020
From: CHORAKHALIKAR, NEILESH; JOSHI, MANAS; ARUNACHALAM, ARUNACHALAM; SHAIKH, BABER M.
To: NVIDIA CORPORATION
Reel/Frame 053360/0296 →
Continuity (1)
Related Publication 20220027578A1 · Jan 27, 2022
References Cited (95)
US 6061675A · Wical · 2000 [cited by examiner]
US 6675159B1 · Lin · 2004 [cited by examiner]
US 7376644B2 · Chen · 2008 [cited by examiner]
US 9454524B1 · Modani · 2016 [cited by examiner]
US 10311087B1 · Kayyoor · 2019 [cited by examiner]
US 10445745B1 · Chopra · 2019 [cited by examiner]
US 11010666B1 · Terilla · 2021 [cited by examiner]
US 11263523B1 · Duchon · 2022 [cited by examiner]
US 11281860B2 · Yue · 2022 [cited by examiner]
US 11481417B2 · Tiwari · 2022 [cited by examiner]
US 20090177463A1 · Gallagher · 2009 [cited by examiner]
US 20090327878A1 · Grandison · 2009 [cited by examiner]
US 20100185689A1 · Hu · 2010 [cited by examiner]
US 20120185466A1 · Yamasaki · 2012 [cited by examiner]
US 20120209871A1 · Lai · 2012 [cited by examiner]
US 20130021346A1 · Terman · 2013 [cited by examiner]
US 20150262069A1 · Gabriel · 2015 [cited by examiner]
US 20150339288A1 · Baker · 2015 [cited by examiner]
US 20160034512A1 · Singhal · 2016 [cited by examiner]
US 20160162456A1 · Munro · 2016 [cited by examiner]
US 20160224662A1 · King · 2016 [cited by examiner]
US 20170060999A1 · Son · 2017 [cited by examiner]
US 20170075877A1 · Lepeltier · 2017 [cited by examiner]
US 20180011920A1 · Simske · 2018 [cited by examiner]
US 20180143980A1 · Tanikella · 2018 [cited by examiner]
US 20180225368A1 · Grond · 2018 [cited by examiner]
US 20180267997A1 · Lin · 2018 [cited by examiner]
US 20180268053A1 · Tata et al. · 2018 [cited by applicant]
US 20180300400A1 · Paulus · 2018 [cited by examiner]
US 20180349693A1 · Watanabe · 2018 [cited by examiner]
US 20190103091A1 · Chen · 2019 [cited by applicant]
US 20190155877A1 · Sharma · 2019 [cited by examiner]
US 20190244119A1 · Farri · 2019 [cited by examiner]
US 20190251150A1 · Vinay · 2019 [cited by examiner]
US 20190354583A1 · Ralhan · 2019 [cited by examiner]
US 20200034366A1 · Kivatinos · 2020 [cited by examiner]
US 20200050638A1 · Hancock · 2020 [cited by examiner]
US 20200057807A1 · Kapur · 2020 [cited by examiner]
US 20200065387A1 · Matthews · 2020 [cited by examiner]
US 20200110916A1 · Chang · 2020 [cited by examiner]
US 20200202171A1 · Hughes · 2020 [cited by examiner]
US 20200242299A1 · Ekmekci · 2020 [cited by examiner]
US 20200272692A1 · Maan · 2020 [cited by examiner]
US 20200327151A1 · Coquard · 2020 [cited by examiner]
US 20200357408A1 · Boekweg · 2020 [cited by examiner]
US 20200380022A1 · Ramakrishna · 2020 [cited by examiner]
US 20200387531A1 · Agnihotram · 2020 [cited by examiner]
US 20200401767A1 · Hirao · 2020 [cited by examiner]
US 20210034964A1 · Chu · 2021 [cited by examiner]
US 20210042467A1 · Liu · 2021 [cited by examiner]
US 20210141822A1 · McLeod · 2021 [cited by examiner]
US 20210165969A1 · Galitsky · 2021 [cited by examiner]
US 20210174016A1 · Fox · 2021 [cited by examiner]
US 20210232943A1 · Abishek Kumar · 2021 [cited by examiner]
US 20210248322A1 · Modani · 2021 [cited by examiner]
US 20210248323A1 · Maheshwari · 2021 [cited by examiner]
US 20210248326A1 · Han · 2021 [cited by examiner]
US 20210256391A1 · Karlinsky · 2021 [cited by examiner]
US 20210286951A1 · Song · 2021 [cited by examiner]
US 20210295822A1 · Tomkins · 2021 [cited by examiner]
US 20210311973A1 · Radhakrishnan · 2021 [cited by examiner]
US 20210326098A1 · Neckermann · 2021 [cited by examiner]
US 20210328888A1 · Rath · 2021 [cited by examiner]
US 20210350202A1 · Zachariah · 2021 [cited by examiner]
US 20210365773A1 · Subramanian · 2021 [cited by examiner]
US 20210374338A1 · Shrivastava · 2021 [cited by examiner]
US 20210406460A1 · Chen · 2021 [cited by examiner]
US 20210406735A1 · Nahamoo · 2021 [cited by examiner]
US 20220012268A1 · Ghoshal · 2022 [cited by examiner]
US 20220222437A1 · Lauber · 2022 [cited by examiner]
US 20220237230A1 · Zovic · 2022 [cited by examiner]
US 20220293107A1 · Leaman · 2022 [cited by examiner]
US 20220318522A1 · Wolf · 2022 [cited by examiner]
US 20250298838A1 · Agley · 2025 [cited by examiner]
CN 102184028A · 2011 [cited by applicant]
CN 102662987A · 2012 [cited by applicant]
CN 108009135A · 2018 [cited by applicant]
CN 108153864A · 2018 [cited by applicant]
CN 110069631A · 2019 [cited by applicant]
CN 110191096A · 2019 [cited by applicant]
CN 111316274A · 2020 [cited by applicant]
Alguliev et al., title={DESAMC+ DocSum: Differential evolution with self-adaptive mutation and crossover parameters for multi-document summarization}, journal={Knowledge-Based Systems}, vol. ={36}, pp. ={21-38}, 2012 (Y… [cited by examiner]
Alguliev et al., title={pSum-SaDE: A Modified p-Median Problem and Self-Adaptive Differential Evolution Algorithm for Text Summarization}, journal={Applied Computational Intelligence and Soft Computing}, vol. ={2011}, N… [cited by examiner]
Chu et al., “MeanSum: A Model for Unsupervised Summarization,” Jan. 29, 2019, 22 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2021/043249, mailed Oct. 27, 2021, filed Jul. 26, 2021, 16 pages. [cited by applicant]
Juvekar et al., “Comparing the Performance of Neural and Statistical Sentence Embeddings on Summarization and Word Sense Disambiguation,” International Conference on Advances in Computing, Communications and Informatics… [cited by applicant]
Liu et al., “Text Summarization with Pretrained Encoders,” Sep. 5, 2019, 11 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Office Action for Chinese Application No. 202180007860.6, mailed May 30, 2025, 21 pages. [cited by applicant]
Office Action for Chinese Application No. 202180007860.6, mailed Oct. 1, 2025, 8 pages. [cited by applicant]
Notice of Decision to Grant for Chinese Application No. 202180007860.6, mailed Jan. 15, 2026, 12 pages. [cited by applicant]
Fang et al., “Multi-Directional Text in Natural Scene,” International Conference on Computational Systems and Communications, May 16, 2018, 5 pages. [cited by applicant]
Jose et al., “Lexical Normalization Model for Noisy SMS Text,” First Internatioanl Conference on Computational Systems and Communications, Dec. 2014, 6 pages. [cited by applicant]