IP Library › Granted Patent US 12,361,217
Granted Patent B2
US 12,361,217 · App. 17/005,028 · Granted Jul 15, 2025

System and method to extract customized information in natural language text

Inventors: Yashu Seth (Patna, IN); Badri Nath (Edison, NJ); Amrit Seshadri Diggavi (Bangalore, IN); Vijayendra Mysore Shamanna (Bangalore, IN); Henry Thomas Peter (Mountain House, CA); Simha Sadasiva (San Jose, CA)
Assignee: Ushur, Inc.
G06F40/289G06F40/284G06F40/47
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,217
App. No.
17/005,028
Granted
Jul 15, 2025
Kind
B2
Abstract

The present disclosure relates to systems and methods to extract customized keywords and their corresponding values occurring in a given natural language text. The desired keyword or keywords may occur in different forms, synonyms, abbreviations, and spellings. The disclosed automatic extraction method captures the meaning and context of the desired keywords by transforming the extraction problem into a question answering problem together with capturing the context to narrow down the answer to a unique value for a given keyword. A trained model on an existing corpus of text is used to get a value as an answer to the question phrased using the keyword. When the answer is ambiguous, a context model that uses conditional random field (CRF) is used to provide a most likely value.

Claims (43)

1. A computer-implemented method for automatically recognizing keywords and corresponding values in a word sequence appearing in a natural language text, the method comprising:

training a model for natural language processing to identify, using a deep neural network, desired keywords contextually, irrespective of variation in form of a particular keyword in a plurality of training word sequences;

identifying, by a processing device executing the trained model for natural language processing in the computer, a plurality of keywords, wherein each keyword of the plurality of keywords is included in the word sequence appearing in the natural language text that is directly received from a customer without a speech to text conversion;

generating custom tags from the plurality of keywords;

dynamically framing, by the processing device, using the trained model, one or more questions from the plurality of keywords and their contexts in the natural language text directly received from the customer, wherein the one or more questions are not framed by a human user;

obtaining, without building a question repository, one or more answers to the one or more questions using the custom tags in the trained model, wherein the one or more answers are already present in their entirety in the word sequence appearing in the natural language text that is directly received from the customer;

extracting, by an extractor module in the processing device, the one or more answers as corresponding values to be associated with the plurality of keywords, wherein the extractor module filters out contextually irrelevant custom tags; and

providing the plurality of keywords and the corresponding values as output in a form of key-value pair.

2. The method of claim 1 , wherein the output in the form of key-value pair is provided as a tuple.

3. The method of claim 1 , wherein a desired text that is input to the extractor module is devoid of any of the plurality of keywords in the natural language text, and an associated text to be extracted is understood from form and context of the desired text.

4. The method of claim 1 , where an answer returned from the trained model for natural language processing is determined to be unique.

5. The method of claim 1 , where an answer returned from trained model for natural language processing is determined to have multiple values.

6. The method of claim 1 , further comprising:

responsive to determining that the output contains ambiguous answers, filtering the output to obtain a unique value using a trained conditional random field (CRF) model.

7. The method of claim 6 , further comprising:

providing the answers as a string, wherein the string is encoded as a sentence embedding.

8. The method of claim 7 , further comprising:

using the trained CRF model to tag a word as a keyword;

extracting a prefix or a suffix of the keyword as a corresponding value; and

providing the keyword and the corresponding value as a tuple as a final output.

9. The method of claim 1 , wherein the deep neural network is based on a Bidirectional Encoder Representations from Transformers (BERT) model.

10. The method of claim 1 , wherein the deep neural network model is a sequence-to-sequence model.

11. The method of claim 6 , wherein the CRF model is trained on sentence embedding.

12. The method of claim 6 , wherein the CRF model is trained on word embedding.

13. The method of claim 6 , where the CRF model is trained on possible word sequences that include a desired keyword.

14. A computer system comprising:

a memory; and

a processing device executing a model for natural language processing in the computer, the processing device being operatively coupled to the memory, performing operations for automatically recognizing keywords and corresponding values in a word sequence appearing in a natural language text, the operations comprising:

training the model for natural language processing to identify, using a deep neural network, desired keywords contextually, irrespective of variation in form of a particular keyword in a plurality of training word sequences;

identifying a plurality of keywords, wherein each keyword of the plurality of keywords is included in the word sequence appearing in the natural language text that is directly received from a customer without a speech to text conversion;

generating custom tags from the plurality of keywords;

dynamically framing, using the trained model, one or more questions from the plurality of keywords and their contexts in the natural language text directly received from the customer, wherein the one or more questions are not framed by a human user;

obtaining, without building a question repository, one or more answers to the one or more questions using the custom tags in the trained model, wherein the one or more answers are already present in their entirety in the word sequence appearing in the natural language text that is directly received from a customer;

extracting, by an extractor module in the processing device in the computer, the one or more answers as corresponding values to be associated with the plurality of keywords, wherein the extractor module filters out contextually irrelevant custom tags; and

providing the plurality of keywords and the corresponding values as output in a form of key-value pair.

15. A non-transitory computer readable medium comprising instructions for a model for natural language processing, which when executed by a processing device in the computer, cause the processing device to perform operations for automatically recognizing keywords and corresponding values in a word sequence appearing in a natural language text, the operations comprising:

training the model for natural language processing to identify, using a deep neural network, desired keywords contextually, irrespective of variation in form of a particular keyword in a plurality of training word sequences;

identifying a plurality of keywords, wherein each keyword of the plurality of keywords is included in the word sequence appearing in the natural language text that is directly received from a customer without a speech to text conversion;

generating custom tags from the plurality of keywords;

dynamically framing, using the trained model, one or more questions from the plurality of keywords and their contexts in the natural language text directly received from the customer, wherein the one or more questions are not framed by a human user;

obtaining, without building a question repository, one or more answers to the one or more questions using the custom tags in the trained model, wherein the one or more answers are already present in their entirety in the word sequence appearing in the natural language text that is directly received from the customer;

extracting, by an extractor module in the processing device, the one or more answers as corresponding values to be associated with the plurality of keywords, wherein the extractor module filters out contextually irrelevant custom tags; and

providing the plurality of keywords and the corresponding values as output in a form of key-value pair.

Assignments (2)
SECURITY INTEREST Recorded Jun 9, 2025
From: USHUR, INC
To: HERCULES CAPITAL, INC.
Reel/Frame 071362/0734 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2020
From: SETH, YASHU; NATH, BADRI; DIGGAVI, AMRIT SESHADRI; SHAMANNA, VIJAYENDRA MYSORE; PETER, HENRY THOMAS; SADASIVA, SIMHA
To: USHUR, INC.
Reel/Frame 054450/0426 →
Continuity (2)
Provisional Application 62892412 · Aug 27, 2019
Related Publication 20210064821A1 · Mar 4, 2021
References Cited (38)
US 6633846B1 · Bennett · 2003 [cited by examiner]
US 8000956B2 · Brun et al. · 2011 [cited by applicant]
US 9009134B2 · Xu et al. · 2015 [cited by applicant]
US 9009153B2 · Khan et al. · 2015 [cited by applicant]
US 9190055B1 · Kiss et al. · 2015 [cited by applicant]
US 9767094B1 · Beller · 2017 [cited by examiner]
US 11055355B1 · Monti · 2021 [cited by examiner]
US 20050165613A1 · Kim · 2005 [cited by examiner]
US 20080052262A1 · Kosinov et al. · 2008 [cited by applicant]
US 20090222395A1 · Light et al. · 2009 [cited by applicant]
US 20090326923A1 · Yan et al. · 2009 [cited by applicant]
US 20160275148A1 · Jiang · 2016 [cited by examiner]
US 20160358094A1 · Fan · 2016 [cited by examiner]
US 20180053107A1 · Wang et al. · 2018 [cited by applicant]
US 20180300317A1 · Bradbury · 2018 [cited by applicant]
US 20180336183A1 · Lee · 2018 [cited by examiner]
US 20190065576A1 · Peng et al. · 2019 [cited by applicant]
US 20200034357A1 · Panuganty · 2020 [cited by examiner]
US 20200034764A1 · Panuganty · 2020 [cited by examiner]
US 20200065342A1 · Panuganty · 2020 [cited by examiner]
EP 1245023A1 · 2002 [cited by applicant]
JP 2014085947A · 2014 [cited by examiner]
WO 0046701A1 · 2000 [cited by applicant]
WO 0135391A1 · 2001 [cited by applicant]
George He “Multi-Task Deep Neural Networks for Generalized Text Understanding”, last modified date: Mar. 28, 2019, URLs: https://web.stanford.edu/class/archive/cs/cs224n/cs224n. 1194/reports/default/15734641.pdf and htt… [cited by examiner]
Chakraborty, Nilesh, et al. “Introduction to neural network based approaches for question answering over knowledge graphs.” arXiv preprint arXiv:1907.09361 (Jul. 2019). (Year: 2019). [cited by examiner]
M. Wakchaure et al., “A Scheme of Answer Selection in Community Question Answering Using Machine Learning Techniques,” 2019 International Conference on Intelligent Computing and Control Systems (ICCS), Conference: May 2… [cited by examiner]
Kim et al. (M. -K. Kim and H. -J. Kim, “Design of Question Answering System with Automated Question Generation,” 2008 Fourth International Conference on Networked Computing and Advanced Information Management, Gyeongju,… [cited by examiner]
Wang et al. (“QG-net: a data-driven question generation model for educational content.” Proc. of the fifth annual ACM conference on learning at scale. 2018. (https://dl.acm.org/doi/pdf/10.1145/3231644.3231654)) (Year: 2… [cited by examiner]
International Search Report and Written Opinion for International Application No. PCT/US2020/48266, mailed on Nov. 17, 2020, 85 pages. [cited by applicant]
Hirotaka Funayama et al., “Bottom up named entity recognition using a two-stage machine learning method”, https://www.aclweb.org/anthology/W09-2908.pdf, ACL-IJCNLP, 2009, pp. 55-62 , Singapore. [cited by applicant]
Jiang Guo et al., “Revisiting Embedding Features for Simple Semi-supervised Learning”, http://people.csail.mit.edu/liang_guo/papers/emnlp2014-semiemb.pdf, In proceedings of EMNLP, Oct. 2014, pp. 111-120, Beijing, China. [cited by applicant]
Jacob Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, https://www.aclweb.org/anthology/N19-1423/, In Proceedings of the Conference on NAACL, 2019, pp. 4171-4186. [cited by applicant]
Jenny Rose Finkel et al., “Incorporating Non-local Information into Information Extraction Systems by Gibbs Sampling”, https://www.aclweb.org/anthology/P05-1045.pdf, Proceedings of the 43nd Annual Meeting of the Associa… [cited by applicant]
Pranav Rajpurkar et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text”, https://arxiv.org/abs/1606.05250, In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, 2016. … [cited by applicant]
Soumay Wadhwa et al., “WadhwaComparative Analysis of Neural QA models on SQUAD”, https://arxiv.org/abs/1806.06972, Workshop on Machine Reading for Question Answering (MRQA), ACL 2018, 9 pages. [cited by applicant]
Motoki Sato et al., “Segment-Level Neural Conditional Random Fields for Named Entity Recognition”, https://www.aclweb.org/anthology/117-2017/, Proceedings of the Eighth International Joint Conference on Natural Language… [cited by applicant]
Extended European Search Report Serial No. EP 20859060.4, dated Jul. 26, 2023, 7 pages. [cited by applicant]