IP Library Granted Patent US 12,373,641
Granted Patent B2
US 12,373,641 · App. 17/187,764 · Granted Jul 29, 2025

Methods and apparatus for natural language understanding in conversational systems using machine learning processes

Inventor: Deepa Mohan (San Jose, CA)
Assignee: Walmart Apollo, LLC
G06F40/284G06F16/3329G06F16/3347G06F40/295G06F40/40G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,641
App. No.
17/187,764
Granted
Jul 29, 2025
Kind
B2
Abstract

This application relates to apparatus and methods for natural language understanding in conversational systems using machine learning processes. In some examples, a computing device receives a request that identifies textual data. The computing device applies a natural language model to the textual data to generate first embeddings. In some examples, the natural language model is trained on retail data, such as item descriptions and chat session data. The computing device also applies a dependency based model to the textual data to generate second embeddings. Further, the computing device concatenates the first and second embeddings, and applies an intent and entity classifier to the concatenated embeddings to determine entities, and an intent, for the request. The computing device may generate a response to the request based on the determined intent and entities.

Claims (47)

1. A system, comprising:

a database;

a processor communicatively coupled to the database; and non-transitory memory storing instructions that, when executed, cause the processor to:

receive input data comprising a plurality of characters;

generate word embeddings based on the plurality of characters;

train a Bidirectional Encoder Representation from Transformers (BERT) model based on catalog data identifying one or more attributes of a plurality of items and chat session data;

apply the BERT model to the word embeddings to generate first output embeddings and a pooled output;

train a dependency embedding generation model to generate dependency based embeddings, wherein the dependency embedding generation model comprises a syntactic dependency embedding model to generate syntactic dependency-based word embeddings, a transformer encoder comprising a single layer transformer block to generate encoded data from the syntactic dependency-based word embeddings, and a linear network to generate the dependency based embeddings from the encoded data;

generate concatenated embeddings by stepwise concatenation of the first output embeddings and the dependency based embeddings;

apply a first linear layer to the concatenated embeddings to generate second output embeddings;

apply a second linear layer to the concatenated embeddings and the pooled output to predict an intent class; and

store the second output embeddings and the intent class in the database.

2. The system of claim 1 , wherein the plurality of characters comprise a previous context and a current context.

3. The system of claim 2 , wherein the previous context is based on a previous chat session, and the current context is based on an inquiry.

4. The system of claim 1 , wherein the processor reads the instructions to apply a neural network to previously generated intent data to generate third output embeddings, where the linear layer is applied to the first output embeddings and the third output embeddings to generate the second output embeddings.

5. The system of claim 1 , wherein the processor reads the instructions to:

normalize the concatenated embeddings, and

generate the second output embeddings based on the normalized concatenated embeddings.

6. The system of claim 1 , wherein the processor reads the instructions to tokenize the input data into a plurality of tokens, wherein the BERT model is applied to the tokenized input data.

7. The system of claim 1 , wherein the second output embeddings characterize a named entity of the plurality of characters.

8. The system of claim 1 , wherein the plurality of characters identify a first item title, and the second output embeddings characterize a second item title that is shorter than the first item title.

9. The system of claim 1 , wherein the input data is received in a request from a second computing device, and wherein the computing device is further configured to generate a response to the request based on the output values.

10. A method, comprising:

receiving input data comprising a plurality of characters;

generating word embeddings based on the plurality of characters;

training a Bidirectional Encoder Representation from Transformers (BERT) model based on catalog data identifying one or more attributes of a plurality of items and chat session data;

applying the BERT model to the word embeddings to generate first output embeddings and a pooled output;

training a dependency embedding generation model to generate dependency based embeddings, wherein the dependency embedding generation model comprises a syntactic dependency embedding model to generate syntactic dependency-based word embeddings, a transformer encoder comprising a single layer transformer block to generate encoded data from the syntactic dependency-based word embeddings, and a linear network to generate the dependency based embeddings from the encoded data;

generating concatenated embeddings by stepwise concatenation of the first output embeddings and the dependency based embeddings;

applying a first linear layer to the concatenated embeddings to generate second output embeddings;

applying a second linear layer to the concatenated embeddings and the pooled output to predict an intent class; and

storing the second output embeddings and the intent class in a database.

11. The method of claim 10 , further comprising applying a neural network to previously generated intent data to generate third output embeddings, where the linear layer is applied to the first output embeddings and the third output embeddings to generate the second output embeddings.

12. The method of claim 11 , further comprising concatenating the concatenated embeddings and the third output embeddings to generate second concatenated embeddings, wherein the linear layer is applied to the second concatenated embeddings.

13. The method of claim 10 , further comprising tokenizing the input data into a plurality of tokens, wherein the BERT model is applied to the tokenized input data.

14. The method of claim 10 , wherein the second output embeddings characterize a named entity of the plurality of characters.

15. The method of claim 10 , wherein the plurality of characters identify a first item title, and the second output embeddings characterize a second item title that is shorter than the first item title.

16. A non-transitory computer readable medium having instructions stored thereon, wherein the instructions, when executed by at least one processor, cause a device to perform operations comprising:

receiving input data comprising a plurality of characters;

generating word embeddings based on the plurality of characters;

training a Bidirectional Encoder Representation from Transformers (BERT) model based on catalog data identifying one or more attributes of a plurality of items and chat session data;

applying the BERT model to the word embeddings to generate first output embeddings and a pooled output;

training a dependency embedding generation model to generate dependency based embeddings, wherein the dependency embedding generation model comprises a syntactic dependency embedding model to generate syntactic dependency-based word embeddings, a transformer encoder comprising a single layer transformer block to generate encoded data from the syntactic dependency-based word embeddings, and a linear network to generate the dependency based embeddings from the encoded data;

generating concatenated embeddings by stepwise concatenation of the first output embeddings and the dependency based embeddings;

applying a first linear layer to the concatenated embeddings to generate second output embeddings;

applying a second linear layer to the concatenated embeddings and the pooled output to predict an intent class; and

storing the second output embeddings and the intent class in a database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2021
From: MOHAN, DEEPA
To: WALMART APOLLO, LLC
Reel/Frame 058075/0487 →
Continuity (1)
Related Publication 20220277142A1 · Sep 1, 2022
References Cited (38)
US 7092871B2 · Pentheroudakis et al. · 2006 [cited by applicant]
US 8620836B2 · Ghani et al. · 2013 [cited by applicant]
US 20180121788A1 · Hashimoto et al. · 2018 [cited by applicant]
US 20180232438A1 · Gu · 2018 [cited by examiner]
US 20200050949A1 · Sundararaman et al. · 2020 [cited by applicant]
US 20200143265A1 · Jonnalagadda et al. · 2020 [cited by applicant]
US 20200265116A1 · Chatterjee et al. · 2020 [cited by applicant]
US 20200311198A1 · Poon · 2020 [cited by examiner]
US 20200388396A1 · Lindvall · 2020 [cited by applicant]
US 20200410337A1 · Huang et al. · 2020 [cited by applicant]
US 20210117623A1 · Aly · 2021 [cited by examiner]
US 20210133535A1 · Zhao · 2021 [cited by examiner]
US 20210142181A1 · Liu et al. · 2021 [cited by applicant]
US 20210255862A1 · Volkovs et al. · 2021 [cited by applicant]
US 20220092267A1 · Hou et al. · 2022 [cited by applicant]
US 20220129633A1 · Shah · 2022 [cited by examiner]
US 20220164626A1 · Bird · 2022 [cited by examiner]
US 20220198327A1 · Wang · 2022 [cited by examiner]
US 20220229993A1 · Vu · 2022 [cited by examiner]
US 20220237377A1 · Zhang · 2022 [cited by examiner]
CN 111640425A1 · 2020 [cited by applicant]
CN 112613326B · 2022 [cited by examiner]
Xingran, Zhu, Cross-lingual Word Sense Disambiguation using mBERT Embeddings with Syntactic Dependencies, Dec. 9, 2020 (Year: 2020). [cited by examiner]
“SpaCy 101: Everything you need to know”, retrieved Feb. 29, 2020 via Wayback Machine, Explosion (Year: 2020). [cited by examiner]
Wu, Shuanzhi; Zhang, Dongdong; Zhang, Zhirui; Yang, Nan; Li, Mu; Zhou, Ming, “Dependency-to-Dependency Neural Machine Translation”, Jul. 13, 2018, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26… [cited by examiner]
Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; Uszkoreit, Jakob; Jones, Llion; Gomez, Aidan N.; Kaiser, Lukasz, “Attention Is All You Need”, 2017, 31st Conference on Neural Information Processing Systems (Year: 2017). [cited by examiner]
Bunk, Tanja et al., “Diet: Lightweight Language Understanding for Dialogue Systems”, arXiv:2004.09936v3, May 11, 2020, 9 pages. [cited by applicant]
Google Research “Use BERT fine-tuned model for Tensorflow serving”, Nov. 20, 2018, 45 pages. [cited by applicant]
Mukerjee, Snehasish, “Discriminative Pre-training for Low Resource Title Compression in Conversational Grocery”, SIGIR eCom'20, Jul. 30, 2020, Virtual Event, China, 7 pages. [cited by applicant]
Zivkovic, Nikola M. “Deploying Machine Learning Models—pt. 3: gRPC and TensorFlow Serving”, Rubikscode.net, found at <https://rubikscode.net/2020/02/24/deploying-machine-learning-models-pt-3-grpc-and-tensorflow-serving/… [cited by applicant]
R. Horey, “BERT: State of the Art NLP Model, Explained,” https://www.kdnuggets.com/2018/12/bert-sota-nlp-model-explained.html, (2018), 8 pages. [cited by applicant]
Q. Liu, et al., “A Survey of Contextual Embeddings,” arXiv preprint ar Xiv:2003.07278, (2020), 13 pages. [cited by applicant]
D. Sundararaman, et al., “Syntax-Infused Transformer and BERT models for Machine Translation and Natural Language Understanding,” ArXiv:1911.06156, Department of Electrical and Computer Engineering, Duke University, (20… [cited by applicant]
S. MacAvaney, et al. “A deeper Look into Dependency-Based Word Embeddings,” Proceedings of NAACL-HLT 2018: Student Research Workshop, New Orleans, Louisiana, Jun. 2-4, 2008, pp. 4-45. [cited by applicant]
Google Research, “Use BERT fine-tuned model for Tensorflow serving,” Nov. 19, 2018, 47 pages. [cited by applicant]
Z. Jie et al., “Dependency-Guided LSTM-CRF for Named Entity Recognition,” Singapore University of Technology and Design, arXiv preprint arXiv:1909.10148, (2019), 13 pages. [cited by applicant]
Renhao Cui et al., “Tweets can tell: activity recognition using hybrid gated recurrent neural networks,” Social Network Analysis and Mining, Year: 2020, pp. 1-15. [cited by applicant]
H. Tang et al., “Dependency Graph Enhanced Dual-transformer Structure for Aspect-based Sentiment Classification,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5-10, 2020,… [cited by applicant]