IP Library › Granted Patent US 11,567,981
Granted Patent B2
US 11,567,981 · App. 16/849,885 · Granted Jan 31, 2023

Model-based semantic text searching

Inventors: Trung Bui (San Jose, CA); Yu Gong (San Jose, CA); Tushar Dublish (Uttar Pradesh, IN); Sasha Spala (Arlington, MA); Sachin Soni (New Delhi, IN); Nicholas Miller (Boston, MA); Joon Kim (Dublin, CA); Franck Dernoncourt (Sunnyvale, CA); Carl Dockhorn (San Jose, CA); Ajinkya Kale (San Jose, CA)
Assignee: Adobe Inc.
G06F16/3347G06F40/30G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,567,981
App. No.
16/849,885
Granted
Jan 31, 2023
Kind
B2
Abstract

Techniques and systems are described for performing semantic text searches. A semantic text-searching solution uses a machine learning system (such as a deep learning system) to determine associations between the semantic meanings of words. These associations are not limited by the spelling, syntax, grammar, or even definition of words. Instead, the associations can be based on the context in which characters, words, and/or phrases are used in relation to one another. In response to detecting a request to locate text within an electronic document associated with a keyword, the semantic text-searching solution can return strings within the document that have matching and/or related semantic meanings or contexts, in addition to exact matches (e.g., string matches) within the document. The semantic text-searching solution can then output an indication of the matching strings.

Claims (43)

1. A method for model-based semantic text searches, comprising:

detecting input corresponding to a request to locate text within an electronic document that is associated with a keyword included in the input, wherein the input includes a number of characters of the keyword;

generating a set of tokens from the electronic document, each token corresponding to one or more strings within the electronic document;

sending the keyword and the set of tokens to a machine learning system configured for semantic mapping in response to determining that input corresponding to an additional character of the keyword has not been detected within a threshold period of time following detection of input corresponding to a most recently provided character of the keyword, wherein the machine learning system generates a representation of the keyword, a representation of each token within the set of tokens, and a representation of a string based on contextual usage of the string in relation to other strings in training data;

receiving, from the machine learning system, at least one string within the electronic document that is associated with the keyword, the at least one string being associated with the keyword based on the representation of the keyword having at least a threshold similarity to a representation of a token corresponding to the at least one string; and

outputting an indication of the at least one string that is associated with the keyword.

2. The method of claim 1 , further comprising detecting input prompting an application displaying the electronic document to open a word search user interface, wherein the input including the keyword is received from the word search user interface.

3. The method of claim 2 , further comprising generating the set of tokens of the electronic document in response to detecting the input prompting the application to open the word search user interface.

4. The method of claim 1 , wherein the input corresponding to the request to locate text within the electronic document that is associated with the keyword includes a number of characters of the keyword, and further comprising:

sending the keyword and the set of tokens to the machine learning system in response to determining that the number of characters of the keyword exceeds a threshold number.

5. The method of claim 4 , further comprising:

detecting input corresponding to at least one additional character of the keyword;

sending, to the machine learning system, the keyword that includes the additional character; and

receiving, from the machine learning system based on the keyword that includes the additional character, at least one additional string within the electronic document that is associated with the keyword.

6. The method of claim 1 , wherein representations generated by the machine learning system include feature vectors of words determined within a vector space, and wherein feature vectors of words that have similar contextual usage in the training data are closer together within the vector space than feature vectors of words that have dissimilar contextual usage in the training data.

7. The method of claim 6 , wherein the machine learning system generates the feature vectors using a word embedding model trained to map semantic meanings of words to feature vectors.

8. The method of claim 1 , wherein:

generating the set of tokens of the electronic document includes generating the set of tokens on an end-user device displaying the electronic document; and

sending the set of tokens to the machine learning system includes forwarding the set of tokens to a server external to the end-user device that implements the machine learning system.

9. The method of claim 1 , wherein sending the set of tokens to the machine learning system includes:

forwarding a partial set of tokens of the electronic document before a complete set of tokens of the electronic document is generated, wherein the machine learning system determines at least one initial word within the electronic document that is associated with the keyword based on the partial set of tokens; and

forwarding the complete set of tokens once the complete set of tokens is generated, wherein the machine learning system determines at least one additional word within the electronic document that is associated with the keyword based on the complete set of tokens.

10. The method of claim 1 , wherein outputting the indication of the least one string that is associated with the keyword includes highlighting each instance of the at least one string within the electronic document.

11. The method of claim 1 , wherein outputting the indication of the at least one string that is associated with the keyword includes displaying the at least one string within a user interface via which a user provided input corresponding to the keyword.

12. The method of claim 1 , wherein the at least one string that is associated with the keyword does not include a string corresponding to the keyword.

13. The method of claim 12 , further comprising:

determining that at least one additional string within the electronic document includes the string corresponding to the keyword; and

outputting an additional indication of the at least one additional string.

14. An end-user device configured for model-based semantic text searches, comprising:

one or more processors; and

memory accessible to the one or more processors, the memory storing instructions, which upon execution by the one or more processors, cause the one or more processors to:

detect, on the end-user device, input corresponding to a request to locate text within an electronic document that is associated with a keyword included in the input, wherein the input includes a number of characters of the keyword;

generate, on the end-user device, a set of tokens from the electronic document, each token corresponding to one or more strings within the electronic document;

send in response to determining that input corresponding to an additional character of the keyword has not been detected within a threshold period of time following detection of input corresponding to a most recently provided character of the keyword, from the end-user device to a machine learning system configured for semantic mapping hosted on a server external to the client device, a request to determine one or more tokens within the electronic document that are associated with the keyword based on determining tokens with representations that have at least a threshold similarity to a representation of the keyword;

receive, at the end-user device from the machine learning system based on the request, at least one string within the electronic document that is associated with the keyword; and

display, on the end-user device, an indication of the at least one string that is associated with the keyword.

15. The end-user device of claim 14 , wherein:

the input corresponding to the request to locate text within the electronic document that is associated with the keyword includes a portion of the keyword;

sending, from the end-user device to the machine learning system, the request to determine the one or more tokens within the electronic document that are associated with the keyword includes sending a request to determine one or more tokens within the electronic document that are associated with the portion of the keyword;

receiving, at the end-user device from the machine learning system based on the request, the at least one string within the electronic document that is associated with the keyword includes receiving at least one string within the electronic document that is associated with the portion of the keyword, and further comprising:

detecting input corresponding to the entire keyword;

sending, from the end-user device to the machine learning system, an additional request to determine one or more tokens within the electronic document that are associated with the entire keyword; and

receiving, at the client device from the machine learning system based on the additional request, at least one additional string within the electronic document that is associated with the entire keyword.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2020
From: BUI, TRUNG; GONG, YU; DUBLISH, TUSHAR; SPALA, SASHA; SONI, SACHIN; MILLER, NICHOLAS; KIM, JOON; DERNONCOURT, FRANCK; DOCKHORN, CARL; KALE, AJINKYA
To: ADOBE INC.
Reel/Frame 052410/0659 →
Continuity (1)
Related Publication 20210326371A1 · Oct 21, 2021