IP Library › Granted Patent US 11,875,585
Granted Patent B2
US 11,875,585 · App. 18/082,386 · Granted Jan 16, 2024

Semantic cluster formation in deep learning intelligent assistants

Inventors: Balaji Vasan Srinivasan (Bangalore, IN); Sujith Sai Venna (San Jose, CA); Kuldeep Kulkarni (San Jose, CA); Durga Prasad Maram (San Jose, CA); Dasireddy Sai Shritishma Reddy (San Jose, CA)
Assignee: ADOBE INC.
G06V30/274G06F16/3329G06F16/355G06F17/18G06N3/08G06V10/763G06V10/82G06V30/153G06V30/19173G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,875,585
App. No.
18/082,386
Filed
Dec 15, 2022
Granted
Jan 16, 2024
Kind
B2
Art Unit
2167
USPC
707/759
Abstract

Enhanced techniques and circuitry are presented herein for providing responses to user questions from among digital documentation sources spanning various documentation formats, versions, and types. One example includes a method comprising receiving a user question directed to subject having a documentation corpus, determining a set of passages of the documentation corpus related to the user question, ranking the set of passages according to relevance to the user question, forming semantic clusters comprising sentences extracted from ranked ones of the set of passages according to sentence similarity, and providing a response to the user question based at least on a selected semantic cluster.

Claims (39)

1. A method comprising:

in response to a question directed to a subject of a documentation corpus that describes operations or features of a software platform or service, extracting a textual answer comprising one or more sentences from one or more top ranked passages of the documentation corpus;

extracting images comprising screenshots of user interfaces of the software platform or service from one or more pages of the documentation corpus from which the textual answer was extracted;

identifying, from the images, one or more selected images within a threshold similarity to the textual answer; and

providing the textual answer and the or more selected images as a multimodal answer to the question.

2. The method of claim 1 , wherein the screenshots of the user interfaces of the software platform or service extracted from the documentation corpus visually represent user interface tabs or dialog boxes, and wherein identifying the one or more selected images identifies one or more of the screenshots of the user interface tabs or the dialog boxes within the threshold similarity to the textual answer.

3. The method of claim 1 , wherein identifying the one or more selected images comprises quantifying similarity between the textual answer and one or more of a heading, caption, title, or optical character recognition (OCR) results associated with at least one image of the images.

4. The method of claim 1 , wherein identifying the one or more selected images comprises quantifying similarity between the textual answer and a sentence encoding of a heading or title associated with least one image of the images.

5. The method of claim 1 , wherein identifying the one or more selected images comprises, for least one image of the images, using optical character recognition (OCR) to generate representations of extracted words in the at least one image, and quantifying similarity between the textual answer and the representations of the extracted words.

6. The method of claim 1 , wherein identifying the one or more selected images comprises, for least one image of the images, quantifying similarity between the textual answer and a weighted average of embeddings of words that are optically extracted from the at least one image and weighted by corresponding inverse term frequency in the documentation corpus.

7. The method of claim 1 , wherein identifying the one or more selected images comprises, for least one image of the images:

generating a plurality of scores quantifying similarity to the textual answer;

generating a combined score based on the plurality of scores; and

determining to whether to include the least one image in the one or more selected images based on the combined score.

8. One or more computer readable storage media storing computer-useable instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising:

extracting, in response to a question and from one or more pages of a documentation corpus that describes operations or features of a software platform or service, a textual answer comprising one or more sentences from one or more top ranked passages of the documentation corpus;

extracting images comprising screenshots of user interfaces of the software platform or service from the one or more pages from which the textual answer was extracted;

identifying, from the images, one or more selected images within a threshold similarity to the textual answer; and

triggering presentation the textual answer and the or more selected images as a multimodal answer to the question.

9. The one or more computer readable storage media of claim 8 , wherein the screenshots of the user interfaces of the software platform or service extracted from the documentation corpus visually represent user interface tabs or dialog boxes, and wherein identifying the one or more selected images identifies one or more of the screenshots of the user interface tabs or the dialog boxes within the threshold similarity to the textual answer.

10. The one or more computer readable storage media of claim 8 , wherein identifying the one or more selected images comprises quantifying similarity between the textual answer and one or more of a heading, caption, title, or optical character recognition (OCR) results associated with least one image of the images.

11. The one or more computer readable storage media of claim 8 , wherein identifying the one or more selected images comprises quantifying similarity between the textual answer and a sentence encoding of a heading or title associated with least one image of the images.

12. The one or more computer readable storage media of claim 8 , wherein identifying the one or more selected images comprises, for least one image of the images, using optical character recognition (OCR) to generate representations of extracted words in the least one image, and quantifying similarity between the textual answer and the representations of the extracted words.

13. The one or more computer readable storage media of claim 8 , wherein identifying the one or more selected images comprises, for at least one image of the images, quantifying similarity between the textual answer and a weighted average of embeddings of words that are optically extracted from the at least one image and weighted by corresponding inverse term frequency in the documentation corpus.

14. The one or more computer readable storage media of claim 8 , wherein identifying the one or more selected images comprises, for least one image of the images:

generating a plurality of scores quantifying similarity to the textual answer;

generating a combined score based on the plurality of scores; and

determining to whether to include the least one image in the one or more selected images based on the combined score.

15. A computer system comprising one or more processors and memory configured to provide computer program instructions to the one or more processors, the computer program instructions comprising:

a response generator configured to extract, in response to a question directed to a subject of a documentation corpus that describes operations or features of a software platform or service, a textual answer comprising one or more sentences from one or more top ranked passages of the documentation corpus; and

an image embellishment element configured to:

extract images comprising screenshots of user interfaces of the software platform or service from one or more pages of the documentation corpus from which the textual answer was extracted; and

identify, from the images, one or more selected images within a threshold similarity to the textual answer,

wherein the response generator is configured to provide the textual answer and the or more selected images as a multimodal answer to the question.

16. The computer system of claim 15 , wherein the screenshots of the user interfaces of the software platform or service extracted from the documentation corpus visually represent user interface tabs or dialog boxes, and wherein identifying the one or more selected images identifies one or more of the screenshots of the user interface tabs or the dialog boxes within the threshold similarity to the textual answer.

17. The computer system of claim 15 , wherein the image embellishment element is configured to identify the one or more selected images based on quantifying similarity between the textual answer and one or more of a heading, caption, title, or optical character recognition (OCR) results associated with least one image of the images.

18. The computer system of claim 15 , wherein the image embellishment element is configured to identify the one or more selected images based on quantifying similarity between the textual answer and a sentence encoding of a heading or title associated with least one image of the images.

19. The computer system of claim 15 , wherein the image embellishment element is configured to identify the one or more selected images based on, for least one image of the images, using optical character recognition (OCR) to generate representations of extracted words in the least one image, and quantifying similarity between the textual answer and the representations of the extracted words.

20. The computer system of claim 15 , wherein the image embellishment element is configured to identify the one or more selected images based on, for at least one image of the images, quantifying similarity between the textual answer and a weighted average of embeddings of words that are optically extracted from the at least one image and weighted by corresponding inverse term frequency in the documentation corpus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2022
From: SRINIVASAN, BALAJI VASAN; VENNA, SUJITH SAI; KULKARNI, KULDEEP; MARAM, DURGA PRASAD; REDDY, DASIREDDY SAI SHRITISHMA
To: ADOBE INC.
Reel/Frame 062143/0066 →
Continuity (2)
Continuation 16888082 · May 29, 2020
Related Publication 20230121355A1 · Apr 20, 2023