IP Library › Granted Patent US 10,891,943
Granted Patent B2
US 10,891,943 · App. 15/874,119 · Granted Jan 12, 2021

Intelligent short text information retrieve based on deep learning

Inventors: Jinren Zhang (Nanjing, CN); Ke Xu (Nanjing, CN); Zhen Fan (Nanjing, CN); Bo Chen (Nanjing, CN)
Assignee: Citrix Systems, Inc.
G10L15/16G06F16/483G06F16/93G06F17/16G06F40/30G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,891,943
App. No.
15/874,119
Granted
Jan 12, 2021
Kind
B2
Abstract

Text based searching can return results based on the system determining the searched text includes keywords or search terms. The present solution can return results based on a semantic analysis. The solutions described herein can provide high accuracy compared against the full-text or keyword-based retrieval algorithms. The solution can sort the results by semantic relevance based on the user's input search request. The present solution can provide meaningful results to the user even when the search text does not include the exact search keywords or phrases entered by the user.

Claims (41)

1. A method to retrieve content based on text input, comprising:

receiving, by a data processing system, a request comprising a plurality of terms;

determining, by a vector generator executed by the data processing system, an average of a plurality of word vectors, the plurality of word vectors including a word vector retrieved for each term of the plurality of terms of the request, the plurality of word vectors generated by multiplying an encoded vector for a respective term of the plurality of terms by a matrix of weights provided by at least one intermediate layer of a neural network;

generating, by the vector generator using the average of the plurality of word vectors, a sentence vector to map the request to a first vector space;

retrieving, from a database by the vector generator, a plurality of trained sentence vectors corresponding to a plurality of candidate electronic documents, wherein each of the plurality of trained sentence vectors map a respective sentence of each of the plurality of candidate electronic documents to the first vector space;

determining, by a scoring engine executed by the data processing system, a distance in the first vector space between the sentence vector and each trained sentence vector of the plurality of trained sentence vectors;

generating, by the scoring engine, a similarity score for each of the plurality of trained sentence vectors based on the respective one of the plurality of trained sentence vectors and the sentence vector and the distance in the first vector space between the sentence vector and each trained sentence vector of the plurality of trained sentence vectors;

selecting, by the scoring engine, an electronic document from the plurality of candidate electronic documents based on a ranking of the similarity score of each of the plurality of trained sentence vectors; and

providing, by the data processing system, the electronic document.

2. The method of claim 1 , further comprising generating, by the vector generator, a word vector for each of the plurality of terms, wherein the word vector maps a respective term of the plurality of terms to a second vector space.

3. The method of claim 2 , wherein the word vector for each of the plurality of terms comprise a vector of weights indicating a probability of one of the plurality of terms occurring.

4. The method of claim 2 , further comprising generating, by the vector generator, the word vector for each of the plurality of terms with one of a Continuous Bag-of-Words neural network model or a Skip-Gram neural network model.

5. The method of claim 1 , further comprising generating, by the vector generator, a trained sentence vector based on an average of candidate word vectors of terms in a sentence.

6. The method of claim 1 , further comprising generating, by the scoring engine, the similarity score for each of the plurality of trained sentence vectors using a Pearson Similarity Calculation.

7. The method of claim 1 , further comprising:

generating, by the scoring engine, a return list comprising a subset of the plurality of candidate electronic documents corresponding to one of the plurality of trained sentence vectors having the similarity score above a predetermined threshold; and

providing, by the data processing system, the return list.

8. The method of claim 1 , further comprising:

calculating, by the vector generator, the sentence vector based on a difference between an inner product of each of a plurality of word vectors in a sentence and a common sentence vector.

9. The method of claim 8 , further comprising calculating, by the vector generator, a common sentence vector by averaging each of the plurality of trained sentence vectors.

10. The method of claim 1 , wherein the plurality of candidate electronic documents comprise web pages, text files, log files, forum questions, or forum answers.

11. The method of claim 1 , further comprising one hot encoding, by the vector generator, each of the plurality of terms to generate a binary array for each of the plurality of terms.

12. A system to retrieve content based on text input, the system comprising a memory storing processor executable instructions and one or more processors to:

receive a request comprising a plurality of terms;

determine, by a vector generator executed by the data processing system, an average of a plurality of word vectors, the plurality of word vectors including a word vector retrieved for each term of the plurality of terms of the request, the plurality of word vectors generated by multiplying an encoded vector for a respective term of the plurality of terms by a matrix of weights provided by at least one intermediate layer of a neural network;

generate, by the vector generator using the average of the plurality of word vectors, a sentence vector to map the request to a first vector space;

retrieve, from a database by the vector generator, a plurality of trained sentence vectors corresponding to a plurality of candidate electronic documents, wherein each of the plurality of trained sentence vectors map a respective sentence of each of the plurality of candidate electronic documents to the first vector space;

determine, by a scoring engine executed by the data processing system, a distance in the first vector space between the sentence vector and each trained sentence vector of the plurality of trained sentence vectors;

generate, by the scoring engine, a similarity score for each of the plurality of trained sentence vectors based on the respective one of the plurality of trained sentence vectors and the sentence vector and the distance in the first vector space between the sentence vector and each trained sentence vector of the plurality of trained sentence vectors;

select, by the scoring engine, an electronic document from the plurality of candidate electronic documents based on a ranking of the similarity score of each of the plurality of trained sentence vectors; and

provide the electronic document.

13. The system of claim 12 , further comprising the one or more processors to generate, by the vector generator, a word vector for each of the plurality of terms, wherein the word vector maps a respective term of the plurality of terms to a second vector space.

14. The system of claim 13 , wherein word vector for each of the plurality of terms comprises a vector of weights indicating a probability of one of the plurality of terms occurring.

15. The system of claim 13 , further comprising the one or more processors to generate, by the vector generator, the word vector for each of the plurality of terms with one of a Continuous Bag-of-Words neural network model or a Skip-Gram neural network model.

16. The system of claim 12 , further comprising the one or more processors to generate, by the vector generator, a trained sentence vector based on an average of candidate word vectors of terms in a sentence.

17. The system of claim 12 , further comprising the one or more processors to generate, by the scoring engine, the similarity score for each of the plurality of trained sentence vectors using a Pearson Similarity Calculation.

18. The system of claim 12 , further comprising the one or more processors to:

generate, by the scoring engine, a return list comprising a subset of the plurality of candidate electronic documents corresponding to one of the plurality of trained sentence vectors having the similarity score above a predetermined threshold; and

provide the return list.

19. The system of claim 12 , further comprising the one or more processors to calculate, by the vector generator, a common sentence vector by averaging each of the plurality of trained sentence vectors.

20. The system of claim 12 , wherein the plurality of candidate electronic documents comprise web pages, text files, log files, forum questions, or forum answers.

Assignments (9)
PATENT SECURITY AGREEMENT Recorded Aug 15, 2025
From: CLOUD SOFTWARE GROUP, INC.; CITRIX SYSTEMS, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 072488/0172 →
SECURITY INTEREST Recorded May 24, 2024
From: CLOUD SOFTWARE GROUP, INC. (F/K/A TIBCO SOFTWARE INC.); CITRIX SYSTEMS, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 067662/0568 →
PATENT SECURITY AGREEMENT Recorded Apr 14, 2023
From: CLOUD SOFTWARE GROUP, INC. (F/K/A TIBCO SOFTWARE INC.); CITRIX SYSTEMS, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 063340/0164 →
RELEASE AND REASSIGNMENT OF SECURITY INTEREST IN PATENT (REEL/FRAME 062113/0001) Recorded Apr 14, 2023
From: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
To: CITRIX SYSTEMS, INC.; CLOUD SOFTWARE GROUP, INC. (F/K/A TIBCO SOFTWARE INC.)
Reel/Frame 063339/0525 →
PATENT SECURITY AGREEMENT Recorded Oct 7, 2022
From: TIBCO SOFTWARE INC.; CITRIX SYSTEMS, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 062112/0262 →
PATENT SECURITY AGREEMENT Recorded Oct 7, 2022
From: TIBCO SOFTWARE INC.; CITRIX SYSTEMS, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 062113/0470 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Oct 7, 2022
From: TIBCO SOFTWARE INC.; CITRIX SYSTEMS, INC.
To: GOLDMAN SACHS BANK USA, AS COLLATERAL AGENT
Reel/Frame 062113/0001 →
SECURITY INTEREST Recorded Sep 30, 2022
From: CITRIX SYSTEMS, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 062079/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2018
From: ZHANG, JINREN; XU, KE; FAN, ZHEN; CHEN, BO
To: CITRIX SYSTEMS, INC.
Reel/Frame 044669/0761 →
Continuity (1)
Related Publication 20190221204A1 · Jul 18, 2019