IP Library Granted Patent US 12,198,685
Granted Patent B2
US 12,198,685 · App. 18/184,432 · Granted Jan 14, 2025

Systems and methods for formatting informal utterances

Inventors: Sandro Cavallari (Tanjong Pagar, SG); Yuzhen Zhuo (Tiong Bahru, SG); Van Hoang Nguyen (Clementi New Town, SG); Quan Jin Ferdinand Tang (Tanglin, SG); Gautam Vasappanavara (Fremont, CA)
Assignee: PAYPAL, INC.
G10L15/19G06F40/205G06F40/253G06F40/284G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,685
App. No.
18/184,432
Granted
Jan 14, 2025
Kind
B2
Abstract

Methods and systems are presented for translating informal utterances into formal texts. Informal utterances may include words in abbreviation forms or typographical errors. The informal utterances may be processed by mapping each word in an utterance into a well-defined token. The mapping from the words to the tokens may be based on a context associated with the utterance derived by analyzing the utterance in a character-by-character basis. The token that is mapped for each word can be one of a vocabulary token that corresponds to a formal word in a pre-defined word corpus, an unknown token that corresponds to an unknown word, or a masked token. Formal text may then be generated based on the mapped tokens. Through the processing of informal utterances using the techniques disclosed herein, the informal utterances are both normalized and sanitized.

Claims (59)

1. A system, comprising:

a non-transitory memory; and

one or more hardware processors coupled with the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform operations comprising:

receiving, from a user device and via an interactive voice response system, a chat utterance comprising a plurality of words;

dividing the chat utterance into a plurality of character components, wherein each character component in the plurality of character components consists of a single character;

analyzing each character component in the plurality of character components with respect to remaining character components in the plurality of character components;

determining, from a plurality of contexts, a particular context associated with the chat utterance based on the analyzing;

mapping the plurality of words of the chat utterance to a plurality of respective tokens from a group of tokens based on the particular context determined for the chat utterance and one or more characters within each word in the plurality of words, wherein a first word in the plurality of words is mapped to a first token based on a determination that the first word comprises data of a particular type;

generating formatted texts corresponding to the plurality of words based on the mapping, wherein the generating the formatted texts comprises masking the first word in the plurality of words based on the first token to which the first word is mapped;

generating data based on the formatted texts; and

transmitting the data to the user device via the interactive voice response system.

2. The system of claim 1 , wherein the data represents a response to the chat utterance.

3. The system of claim 1 , wherein the chat utterance comprises voice data, and wherein the operations further comprise:

translating, using a voice recognition module, the voice data to the plurality of words.

4. The system of claim 1 , wherein the operations further comprise:

converting the data to audio data; and

providing the audio data to the user device via the interactive voice response system.

5. The system of claim 1 , wherein the masking the first word comprises:

determining a length of the first word; and

replacing the first word with a series of symbols having the length of the first word.

6. The system of claim 1 , wherein the operations further comprise:

presenting the formatted texts on a second device.

7. The system of claim 1 , wherein the operations further comprise:

determining that the first word is mapped to a second token corresponding to a first vocabulary from a dictionary based on the mapping; and

replacing the second token with the first token corresponding to a mask token based on the particular context.

8. A method, comprising:

receiving, from a device and via an interactive voice response system, a chat utterance comprising a plurality of words;

analyzing the plurality of words in a character-by-character basis, wherein the analyzing the plurality of words comprises (i) dividing the chat utterance into a plurality of characters, and (ii) analyzing each character in the plurality of characters with respect to one or more remaining characters in the plurality of characters;

determining, from a plurality of contexts, a particular context associated with the chat utterance based on the analyzing;

determining a mapping between a first word from the plurality of words and a first token from a group of tokens based on one or more characters within the first word;

adjusting, for the first word, the mapping from the first token to a second token of the group of tokens based on the particular context;

generating a formatted chat utterance based at least in part on the adjusted mapping;

generating data based on the formatted chat utterance; and

transmitting the data to the device via the interactive voice response system.

9. The method of claim 8 , wherein the data comprises a response to the chat utterance.

10. The method of claim 8 , wherein the group of tokens comprises a plurality of vocabulary tokens that corresponds to respective vocabularies in a dictionary, wherein the second token corresponds to a particular vocabulary from the dictionary, and wherein the generating the formatted chat utterance comprises replacing the first word with the particular vocabulary.

11. The method of claim 8 , wherein the chat utterance comprises voice data, and wherein the method further comprises generating an audio response to the voice data based on the formatted chat utterance.

12. The method of claim 11 , further comprising:

translating, using a voice recognition module, the voice data into the plurality of words.

13. The method of claim 8 , wherein the second token corresponds to a mask token, and wherein the generating the formatted chat utterance comprises:

determining a length of the first word; and

replacing the first word with a series of symbols having the length of the first word.

14. The method of claim 8 , further comprising presenting the formatted chat utterance on a second device.

15. A non-transitory machine-readable medium having stored thereon machine-readable instructions executable to cause a machine to perform operations comprising:

receiving, from a user device and via an interactive voice response system, a chat utterance comprising a plurality of words;

dividing the chat utterance into a plurality of character components, wherein each character component in the plurality of character components consists of a character;

analyzing each character component in the plurality of character components with respect to one or more other character components in the plurality of character components;

determining, from a plurality of contexts, a particular context associated with the chat utterance based on the analyzing;

mapping the plurality of words of the chat utterance to a plurality of respective tokens from a group of tokens based on the particular context determined for the chat utterance and one or more characters within each word in the plurality of words;

generating formatted texts corresponding to the plurality of words based on the mapping;

generating data based on the formatted texts; and

transmitting the data to the user device via the interactive voice response system.

16. The non-transitory machine-readable medium of claim 15 , wherein the chat utterance comprises voice data, and wherein the operations further comprise generating an audio response based on the data.

17. The non-transitory machine-readable medium of claim 16 , wherein the operations further comprise translating the voice data to the plurality of words.

18. The non-transitory machine-readable medium of claim 16 , wherein the audio response is generated further based on the formatted texts.

19. The non-transitory machine-readable medium of claim 18 , wherein a first word from the plurality of words is mapped to a particular token from the group of tokens that corresponds to a mask token, and wherein the generating the formatted texts comprises:

determining a length of the first word; and

replacing the first word with a series of symbols having the length of the first word.

20. The non-transitory machine-readable medium of claim 15 , wherein a first word from the plurality of words is mapped to a first token from the group of tokens that corresponds to a vocabulary, and wherein the generating the formatted texts comprises replacing the first word with the vocabulary.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2023
From: TANG, QUAN JIN FERDINAND; VASAPPANAVARA, GAUTAM; CAVALLARI, SANDRO; ZHUO, YUZHEN; NGUYEN, VAN HOANG
To: PAYPAL, INC.
Reel/Frame 062994/0314 →
Continuity (2)
Continuation 16831058 · Mar 26, 2020
Related Publication 20230290344A1 · Sep 14, 2023
References Cited (45)
US 20120072204A1 · Nasri et al. · 2012 [cited by applicant]
US 20140365222A1 · Weider · 2014 [cited by examiner]
US 20160335244A1 · Weisman · 2016 [cited by examiner]
US 20190251165A1 · Bachrach et al. · 2019 [cited by applicant]
GB 2327133A · 1999 [cited by examiner]
WO 9924955A1 · 1999 [cited by applicant]
WO WO2019052811A1 · 2019 [cited by examiner]
O. Ali and A. Ouda, “A classification module in data masking framework for Business Intelligence platform in healthcare,” 2016 IEEE 7th Annual Information Technology, Electronics and Mobile Communication Conference (IEM… [cited by examiner]
Ribeiro, Eugénio, Ricardo Ribeiro, and David Martins de Matos. 2019. “A Multilingual and Multidomain Study on Dialog Act Recognition Using Character-Level Tokenization” Information 10, No. 3: 94. https://doi.org/10.3390… [cited by examiner]
Bahdanau D., et al., “Neural Machine Translation by Jointly Learning to Align and Translate,” International Conference on Learning Representations Conference Paper, 2015, pp. 1-15. [cited by applicant]
Brill, Eric et al., “An improved error model for noisy channel spelling correction”, Proceedings of the 38th Annual Meeting on Association for Computational Linguistics, pp. 286-293. Association for Computational Lingui… [cited by applicant]
Cavallari S., et al., “Neural Text Normalisation and Sanitisation for Customer Services,” 8 pages. [cited by applicant]
Cho, Kyunghyun et al., “Learning phrase representations using rnn encoder-decoder for statistical machine translation”, In Conference on Empirical Methods in Natural Language Processing (EMNLP 2014), 2014. [cited by applicant]
Church, Kenneth W. et al., “Probability scoring for spelling correction”, Statistics and Computing, 1 (2): pp. 93-103, 1991, 11 pages. [cited by applicant]
Cook P., et al., “An Unsupervised Model for Text Message Normalization,” Proceedings of the NAACL HLT Workshop on Computational Approaches to Linguistic Creativity, Jun. 2009, 8 pages. [cited by applicant]
Dai, Zihang et al., “Transformer-xl: Attentive language models beyond a fixed-length context”, Conference: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019. [cited by applicant]
Devlin J., et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: H… [cited by applicant]
Dong, Li, et al., “Unified language model pre-training for natural language understanding and generation”, 33rd Conference on Neural Information Processing Systems—NeuriPS 2019,2019. [cited by applicant]
Edizel, Bora et al., “Misspelling oblivious word embeddings”, arXiv preprintarXiv: 1905.09755, 2019, 9 pages. [cited by applicant]
Goldberg, Yoav et al., “word2vec explained: deriving mikolov et al.'s negative-samplingword-embedding method”, arXiv preprint arXiv:1402.3722, 2014, 5 pages. [cited by applicant]
Graves, Alex et al., “Neural turing machines”, arXiv:141 0.5401, 2014, 26 pages. [cited by applicant]
Gu, Jiatao et al., “Levenshtein transformer”, 33rd Conference on NeuralInformatonProcessing System (NeuriPS 2019), Vancouver, Canada, 11 pages. [cited by applicant]
Gulcehre, Cagier et al., “Memory augmented neural networks with wormhole connections”, arXiv preprint arXiv:1701.08718, 2017, 27 pages. [cited by applicant]
Gulcehre, Cagier et al., “Pointing the unknown words”, In 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016—Long Papers (pp. 140-149). (54th Annual Meeting of the Association for Computation… [cited by applicant]
Guyon, Isabelle et al., “Discovering informative patterns and data cleaning”, Advances inKnowledge Discovery and Data Mining, pp. 181-203, 1996. [cited by applicant]
Han B., et al., “Lexical Normalisation of Short Text Messages: Makn Sens a #twitter,” Department of Computer Science and Software Engineering, The University of Melbourne, Proceedings of the 49th Annual Meeting of the A… [cited by applicant]
Hinton, Geoffrey E., “Learning distributed representations of concepts”, Proceedings of the eighth annual conference of the cognitive science society, vol. 1, Amherst, MA, 1986, 12 pages. [cited by applicant]
Hutto, Clayton J., et al., “Vader: A parsimonious rule-based model for sentiment analysis ofsocial media text”, Eighth international AAAI conference on weblogs and social media, 2014, 10 pages. [cited by applicant]
Luong, Thang et al., “Effective approaches to attention-based neural machine translation”, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing,pp. 1412-1421, 2015. [cited by applicant]
Merity, Stephen et al., “Pointer sentinel mixture models”, arXiv preprint arXiv:1609.07843,2016, 13 pages. [cited by applicant]
Mikolov, et al., “Distributed representations of words and phrases and theircompositionality”, Advances in neural information processing systems, pp. 3111-3119,2013, 9 pages. [cited by applicant]
Min, Wookhee et al., “Ncsu sas wookhee: a deep contextual long-short term memorymodel for text normalization”, Proceedings of the Workshop on Noisy User-generated Text,pp. 111-119,2015, 9 pages. [cited by applicant]
Rahm, Erhard et al., “Data cleaning: Problems and current approaches”, IEEE Data Eng.Bull., 23(4):3-13, 2000. [cited by applicant]
Sanchez, David, “Detecting sensitive information from textual documents: an informationtheoreticapproach”, International Conference on Modeling Decisions for Artificial Intelligence,pp. 173-184. Springer, 2012. [cited by applicant]
See, Abigail et al., “Get to the point: Summarization with pointer generator networks”, Proceedings of the 55th Annual Meeting of the Association for ComputationalLinguistics (vol. 1: Long Papers), pp. 1073-1083, 2017. [cited by applicant]
Sennrich, Rico et al., “Neural machine translation of rare words with subword units”, arXivpreprint arXiv:1508.07909, 2015, 12 pages. [cited by applicant]
Sproat, Richard et al., “Rnn approaches to text normalization: A challenge” arXivpreprint arXiv:1611.00068, 2016. [cited by applicant]
Sukhbaatar, Sainbayar et al., “End-to-end memory networks”, Advancesin neural information processing systems, pp. 2440-2448, 2015. [cited by applicant]
Sutskever, Ilya et al. “Sequence to sequence learning with neural networks” Advances inneural information processing systems, pp. 3104-3112, 2014. [cited by applicant]
Sweeney, Latanya, “Replacing personally-identifying information in medical records, thescrub system”, Proceedings of the AMIA annual fall symposium, pp. 333. American MedicalInformatics Association, 1996, 5 pages. [cited by applicant]
Vaswani A., et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 6, 2017, 15 pages. [cited by applicant]
Vinyals, Oriol et al., “Pointer networks”, Advances in Neural Information Processing Systems, pp. 2692-2700, 2015. [cited by applicant]
Weston, Jason et al., “Memory networks”, arXiv preprint arXiv:1410.3916, 2014, 9 pages. [cited by applicant]
Yang, Zhilin et al., “XLNet: Generalized Autoregressive Pretraining for LanguageUnderstanding”, 33rd Conference on Neural Information Processing System (NeuriPS)Canada, 2019, 18 pages. [cited by applicant]
Zhang, Hao et al., “Neural models of text normalization for speech applications”, Computational Linguistics, vol. 45, No. 2, pp. 293-337, 2019. [cited by applicant]