IP Library Granted Patent US 12664377
Granted Patent B2
US 12664377 · App. 16/940,095 · Granted Jun 23, 2026

Text string summarization

Inventors: Neilesh Chorakhalikar (San Jose, CA); Manas Joshi (Pickering, CA); Arunachalam Arunachalam (Austin, TX); Baber M. Shaikh (Redmond, WA)
Assignee: NVIDIA Corporation
G06F40/58G06F9/454G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664377
App. No.
16/940,095
Granted
Jun 23, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to summarize a plurality of text strings. In at least one embodiment, one or more machine learning structures are used to generate a summary of a plurality of text strings based, at least in part, on an association between one or more words in a text string and one or more words with similar meaning.

Claims (38)

1 . One or more processors, comprising:

circuitry to:

use one or more neural networks to generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of the plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

2 . The one or more processors of claim 1 , wherein the circuitry is to normalize the plurality of text strings by replacing the one or more words in the plurality of text strings with the one or more keywords indicative of a topic, and to map the normalized plurality of text strings, using the one or more neural networks, to one or more points in a multidimensional space, wherein distance between points in the multidimensional space is indicative of similarity between text strings mapped to the points.

3 . The one or more processors of claim 2 , wherein the circuitry is to calculate a median of the cluster of encodings, and to generate the summary based, at least in part, on proximity of an encoding of the one or more encodings to the median.

4 . The one or more processors of claim 3 , wherein the circuitry is to evaluate, for suitability as the summary of the plurality of text strings, a text string mapped to the center point based, at least in part, on a conciseness criteria.

5 . The one or more processors of claim 4 , wherein the circuitry is to determine that no text string mapped to a point near a geometric median conforms to the conciseness criteria and, in response to the determination, generate the summary based, at least in part, on a frequency of keywords in the plurality of text strings.

6 . The one or more processors of claim 2 , wherein the one or more neural networks map the normalized plurality of text strings to the center point in the multidimensional space.

7 . The one or more processors of claim 1 , wherein the plurality of text strings comprise user-provided feedback and the one or more words with similar meaning are obtained from a list of keywords related to a subject to which the feedback is predicted to pertain.

8 . The one or more processors of claim 1 , wherein the circuitry is further to evaluate the plurality of text strings based, at least in part, on one or more text strings of the plurality of text strings having fewer than a threshold number of words.

9 . The one or more processors of claim 1 , wherein the summary is normalized by replacing at least a portion of words of the summary with similar words.

10 . A system, comprising:

one or more processors to generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of a plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

11 . The system of claim 10 , the one or more processors to cause the system to generate the keywords indicative of a topic based, at least in part, on a number of key words and a total number of words in one or more of the plurality of text strings.

12 . The system of claim 10 , the one or more processors to map the plurality of text strings to one or more points in a multidimensional space, wherein distance between points in the multidimensional space is indicative of a degree of similarity between semantic meanings of the text strings mapped to the points.

13 . The system of claim 12 , the one or more processors to generate the summary based at least in part on a median of a cluster of the encodings.

14 . The system of claim 13 , the one or more processors to evaluate one or more statements near the median based, at least in part, on a conciseness criteria.

15 . The system of claim 13 , wherein one of the plurality of text strings is selected as a summary, based at least in part, on proximity to the median and a conciseness criteria.

16 . The system of claim 13 , the one or more processors to determine that no point near the median conforms to a conciseness criteria and, in response to the determination, generate the summary based, at least in part, on keyword frequency in the text string.

17 . The system of claim 10 , wherein the plurality of text strings comprise user-provided feedback and the one or more words with similar meaning are obtained from a list of keywords related to a subject to which the feedback pertains.

18 . A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to use one or more machine learning structures to at least generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of the plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

19 . The machine-readable medium of claim 18 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to at least normalize the plurality of text strings by replacing the one or more words in the text strings with the one or more words with similar meaning, and to map the normalized plurality of text strings, using one or more neural networks, to one or more points in a multidimensional space.

20 . The machine-readable medium of claim 19 , wherein distance between points in the multidimensional space is indicative of similarity between text strings mapped to the points.

21 . The machine-readable medium of claim 19 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to calculate a median of the points in the multidimensional space.

22 . The machine-readable medium of claim 21 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to generate the summary of the plurality of text strings by at least evaluating one or more points nearest to the median according to a conciseness criteria.

23 . The machine-readable medium of claim 21 , having stored thereon a set of instructions, which if performed by one or more processors, causes the one or more processors to determine that no point near the median conforms to a conciseness criteria and, in response to the determination, generate the summary based, at least in part, on counting a frequency of keywords identified in the text strings.

24 . A system, comprising:

one or more computing devices that receive a plurality of text strings and use one or more machine learning structures to generate a cluster of encodings of a plurality of text strings, wherein one or more words of the plurality of text strings are replaced with one or more keywords indicative of a topic to cause a representative statistic of the cluster of the encodings to be biased towards the topic, and wherein the one or more keywords are obtained from a plurality of keywords selected based, at least in part, on an expected frequency of respective keywords of the plurality of keywords appearing in the plurality of text strings; and

generate a summary of the plurality of text strings based, at least in part, on identification of a center point of the cluster to select one or more text strings representative of the cluster of the encodings.

25 . The system of claim 24 , the one or more computing devices to normalize the plurality of text strings based, at least in part, on the one or more words with similar meaning.

26 . The system of claim 25 , the one or more computing devices to map the normalized plurality of text strings to one or more points in a multidimensional space.

27 . The system of claim 26 , the one or more computing devices to generate the summary based, at least in part, on a geometric median of a cluster of the points and conciseness of points near the geometric median.

28 . The system of claim 26 , the one or more computing devices to determine that no text string mapped to a point near a geometric median of the one or more points conforms to a conciseness criteria and, in response to the determination, generate the summary based, at least in part, on frequency of keywords in the text string.

29 . The system of claim 24 , wherein the plurality of text strings comprise user-provided feedback.

30 . The system of claim 24 , wherein the one or more words with similar meaning are obtained from a list of keywords related to a service provided by the one or more computing devices.