IP Library Granted Patent US 11,403,457
Granted Patent B2
US 11,403,457 · App. 16/549,846 · Granted Aug 2, 2022

Processing referral objects to add to annotated corpora of a machine learning engine

Inventor: Joy Mustafi (Hyderabad, IN)
Assignee: salesforce.com, inc.
G06F40/169G06F16/955G06N20/00G06V10/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,403,457
App. No.
16/549,846
Granted
Aug 2, 2022
Kind
B2
Abstract

A system is provided for referral object processing for textual annotations. The system comprises a memory storing machine executable code and one or more processors coupled to the memory and configurable to execute the machine executable code to cause the one or more processors to parse a document to identify a reference identifier to an external object, the external object associated with information not contained in the document, retrieve the external object using the reference identifier, extract the information associated with the external object based on at least one data pattern detected in the external object, convert the extracted information into textual annotations associated with the reference identifier in the document, and enter the textual annotations to a corpus of content for the document so that the extracted information is associated with the reference in the document for the system.

Claims (65)

1. A system for referral object processing for textual annotations, the system comprising:

a memory storing machine executable code; and

one or more processors coupled to the memory and configurable to execute the machine executable code to cause the one or more processors to:

parse a document for a reference identifier to an external object, the external object associated with information not contained in the document, wherein parsing the document comprises:

performing optical character recognition on the document,

identifying, using a neural network model, the reference identifier based on a calculated similarity value determined from comparing data from the optical character recognition to one or more reference identifiers used to train the neural network model, and

determining a portion of the document that corresponds to a location within the document where the reference identifier is identified, wherein the portion references the external object using the reference identifier,

retrieve the external object using the reference identifier from parsing the document;

extract the information associated with the external object based on at least one data pattern detected in the external object;

convert the extracted information into the textual annotations associated with the reference identifier in the document;

combine the textual annotations with the portion of the document in a corpus of content so that the extracted information is associated with the reference identifier in the portion of the document by the system, wherein combining comprises integrating, using a natural language processing framework, the textual annotation with text from the portion of the document;

convert the textual annotations to first word embeddings for a machine learning model of a machine learning engine used to search the corpus of content including the document; and

combine the first word embeddings with second word embeddings of the portion of the document having the reference identifier, wherein the machine learning model previously comprises the second word embeddings.

2. The system of claim 1 , wherein the reference identifier comprises one of a hyperlink, a page identifier, a heading, a location identifier, an image identifier, a callout banner, or a table number.

3. The system of claim 1 , wherein the machine executable code further causes the one or more processors to:

train the machine learning model of the machine learning engine using the first word embeddings and the second word embeddings from at least the document and the textual annotations.

4. The system of claim 1 , wherein the machine executable code further causes the one or more processors to:

execute, using the machine learning engine, a search of the corpus of content based on a received search query, wherein the search is performed using at least the document and the textual annotations.

5. The system of claim 4 , wherein the machine executable code further causes the one or more processors to:

in response to the search, determine a portion of the document identified by the search comprises one of the textual annotations; and

provide the information associated with the external object based on the portion comprising the one of the textual annotations.

6. The system of claim 1 , wherein the textual annotations comprise searchable text generated using the information not contained within the document, and wherein the searchable text is associated with a portion of the document having the reference identifier in the corpus of content.

7. The system of claim 1 , wherein the information is extracted using at least one of natural language processing, image processing, further optical character recognition, or website data extraction.

8. A method for referral object processing for textual annotations, the method comprising:

parsing a document for a reference identifier to an external object, the external object associated with information not contained in the document, wherein parsing the document comprises:

performing optical character recognition on the document,

identifying, using a neural network model, the reference identifier based on a calculated similarity value determined from comparing data from the optical character recognition to one or more reference identifiers used to train the neural network model, and

determining a portion of the document that corresponds to a location within the document where the reference identifier is identified, wherein the portion references the external object using the reference identifier,

retrieving the external object using the reference identifier from parsing the document;

extracting the information associated with the external object based on at least one data pattern detected in the external object;

converting the extracted information into the textual annotations associated with the reference identifier in the document;

combining the textual annotations with the portion of the document in a corpus of content so that the extracted information is associated with the reference identifier in the portion of the document, wherein combining comprises integrating, using a natural language processing framework, the textual annotation with text from the portion of the document;

converting the textual annotations to first word embeddings for a machine learning model of a machine learning engine used to search the corpus of content including the document; and

combining the first word embeddings with second word embeddings of the portion of the document having the reference identifier, wherein for the machine learning model previously comprises the second word embeddings.

9. The method of claim 8 , wherein the reference identifier comprises one of a hyperlink, a page identifier, a heading, a location identifier, an image identifier, a callout banner, or a table number.

10. The method of claim 8 , further comprising:

training the machine learning model of the machine learning engine using the first word embeddings and the second word embeddings from at least the document and the textual annotations.

11. The method of claim 8 , further comprising:

executing, using the machine learning engine, a search of the corpus of content based on a received search query, wherein the search is performed using at least the document and the textual annotations.

12. The method of claim 11 , further comprising:

in response to the search, determining a portion of the document identified by the search comprises one of the textual annotations; and

providing the information associated with the external object based on the portion comprising the one of the textual annotations.

13. The method of claim 8 , wherein the textual annotations comprise searchable text generated using the information not contained within the document, and wherein the searchable text is associated with a portion of the document having the reference identifier in the corpus of content.

14. The method of claim 8 , wherein the information is extracted using at least one of natural language processing, image processing, further optical character recognition, or website data extraction.

15. A non-transitory machine-readable medium having stored thereon instructions for performing a method comprising machine executable code which when executed by at least one machine, causes the machine to:

parsing a document for a reference identifier to an external object, the external object associated with information not contained in the document, wherein parsing the document comprises:

performing optical character recognition on the document, and

identifying, using a neural network model, the reference identifier based on a calculated similarity value determined from comparing data from the optical character recognition to one or more reference identifiers used to train the neural network model, and

determining a portion of the document that corresponds to a location within the document where the reference identifier is identified, wherein the portion references the external object using the reference identifier,

retrieving the external object using the reference identifier from parsing the document;

extracting the information associated with the external object based on at least one data pattern detected in the external object;

converting the extracted information into textual annotations associated with the reference identifier in the document;

combining the textual annotations with the portion of the document in a corpus of content so that the extracted information is associated with the reference identifier in the portion of the document, wherein combining comprises integrating, using a natural language processing framework, the textual annotation with text from the portion of the document;

converting the textual annotations to first word embeddings for a machine learning model of a machine learning engine used to search the corpus of content including the document; and

combining the first word embeddings with second word embeddings of the portion of the document having the reference identifier, wherein for the machine learning model previously comprises the second word embeddings.

16. The non-transitory machine-readable medium of claim 15 , wherein the reference identifier comprises one of a hyperlink, a page identifier, a heading, a location identifier, an image identifier, a callout banner, or a table number.

17. The non-transitory machine-readable medium of claim 15 , storing the instructions which when executed by the at least one machine, further causes the machine to:

training the machine learning model of the machine learning engine using the first word embeddings and the second word embeddings from at least the document and the textual annotations.

18. The non-transitory machine-readable medium of claim 15 , storing the instructions which when executed by the at least one machine, further causes the machine to:

executing, using the machine learning engine, a search of the corpus of content based on a received search query, wherein the search is performed using at least the document and the textual annotations.

19. The non-transitory machine-readable medium of claim 18 , storing the instructions which when executed by the at least one machine, further causes the machine to:

in response to the search, determining a portion of the document identified by the search comprises one of the textual annotations; and

providing the information associated with the external object based on the portion comprising the one of the textual annotations.

20. The non-transitory machine-readable medium of claim 15 , wherein the textual annotations comprise searchable text generated using the information not contained within the document, and wherein the searchable text is associated with a portion of the document having the reference identifier in the corpus of content.

21. The non-transitory machine-readable medium of claim 15 , wherein the information is extracted using at least one of natural language processing, image processing, further optical character recognition, or website data extraction.

Assignments (2)
CHANGE OF NAME Recorded Dec 18, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 069717/0444 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2019
From: MUSTAFI, JOY
To: SALESFORCE.COM, INC.
Reel/Frame 050153/0693 →
Continuity (1)
Related Publication 20210056164A1 · Feb 25, 2021