IP Library › Granted Patent US 12,387,729
Granted Patent B2
US 12,387,729 · App. 18/086,742 · Granted Aug 12, 2025

Intelligent voice assistant

Inventors: Vaishnavi Mysore Sridhar (Karnataka, IN); Anupriya Gupta (Uttar Pradesh, IN); Gunjan Narulkar (Karnataka, IN); Ankur Jaiswal (Maharashtra, IN); Suwetha Selvakumar (Tamil Nadu, IN); Mari Pavithra (Tamil Nadu, IN)
Assignee: FMR LLC
G10L15/26G10L15/02H04M3/5175G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,729
App. No.
18/086,742
Granted
Aug 12, 2025
Kind
B2
Abstract

A computer-implemented method is provided for recommending at least one pertinent electronic document for supporting a call between a customer and an agent. The method includes converting in real time content of the call between the customer and the agent from speech to digitized text, isolating a predefined number of words in the digitized text of the converted call content as the call is in progress and converting the predefined number of words in text to a phoneme sequence. The method also includes identifying at least one probable business category associated with the phoneme sequence and detecting sections of one or more documents associated with the probable business category that are similar to the content of the call.

Claims (35)

1. A computer-implemented method for recommending at least one pertinent electronic document for supporting a call between a customer and an agent, the method comprising:

converting in real time, by a computing device, content of the call between the customer and the agent from speech to digitized text;

isolating, by the computing device, a predefined number of words in the digitized text of the converted call content as the call is in progress;

converting, by the computing device, the predefined number of words in text to a phoneme sequence;

identifying, by the computing device, at least one probable business category associated with the phoneme sequence, wherein the probable business category is associated with one or more documents;

detecting, by the computing device, sections of the one or more documents that are similar to the content of the call;

ranking, by the computing device, the one or more documents based on corresponding degrees of relevancy between the similar sections of respective ones of the documents and the content of the call; and

presenting, by the computing device, the ranking to the agent during the call via a user interface.

2. The computer-implemented method of claim 1 , wherein the similar document sections are determined in batches in real time as the call progresses, each batch relating to the predetermined number of words isolated from the call as the call progresses.

3. The computer-implemented method of claim 1 , wherein converting the predefined number of words from text to a phoneme sequence comprises applying a trained transformer model based on neural networks that is configured to convert (i) the predefined number of words to a grapheme representation and (ii) the grapheme representation to the phoneme sequence.

4. The computer-implemented method of claim 1 , wherein identifying the at least one probable business category comprises applying a multi-stage convolutional neural network trained to predict relationships between phoneme sequences and business domains.

5. The computer-implemented method of claim 1 , wherein detecting the sections of the one or more documents that are similar to the content of the call comprises applying a Siamese bidirectional long short term memory (LSTM) network model to capture phrase similarity using phoneme embedding.

6. The computer-implemented method of claim 5 , wherein the Siamese bidirectional LSTM model is trained to detect similarities in phoneme representations of text.

7. The computer-implemented method of claim 1 , wherein the user interface includes: (i) a chat transcription section displaying in real time client-side conversation to the agent and (ii) a domain section identifying the at least one business category pertinent to the conversation displayed in the chat transcriptions section.

8. The computer-implemented method of claim 7 , wherein the user interface further comprises a Top Results section configured to display links to the documents in the at least one pertinent business category that include the similar content, wherein the links are ranked in accordance with the corresponding degrees of relevancy between the documents and the content of the call.

9. The computer-implemented method of claim 8 , wherein the similar content of each linked document is highlighted within each document.

10. The computer-implemented method of claim 7 , further comprising updating the user interface as the call progresses with updated client-side conversation as well as business category identification and similar content identification pertinent to the updated conversation.

11. A computer-implemented system for recommending at least one pertinent electronic document for supporting a call between a customer and an agent, the computer-implemented system comprising a computing device having a memory for storing instructions, wherein the instructions, when executed, configure the computer-implemented system to provide:

a speech transcription module configured to convert in real time content of the call between the customer and the agent from speech to digitized text;

a batching module configured to isolate a predefined number of words in the digitized text of the converted call content as the call is in progress;

a phoneme conversion module configured to convert the predefined number of words in text to a phoneme sequence;

a phoneme based domain detection module configured to identify at least one probable business category associated with the phoneme sequence, wherein the probable business category is associated with one or more documents;

a phoneme based similarity detection module configured to detect sections of the one or more documents that are similar to the content of the call; and

a user interface configured to present to the agent during the call a ranking of the one or more documents based on corresponding degrees of relevancy between the similar sections of respective ones of the documents and the content of the call.

12. The computer-implemented system of claim 11 , wherein the similarity detection module is configured to determine similar document sections in batches in real time as the call progresses, each batch relating to the predetermined number of words isolated by the batching module.

13. The computer-implemented system of claim 11 , wherein the phoneme conversion module converts the predefined number of words by applying a trained transformer model based on neural networks configured to convert (i) the predefined number of words to a grapheme representation and (ii) the grapheme representation to the phoneme sequence.

14. The computer-implemented system of claim 11 , wherein the domain detection module is configured to identify the at least one probable business category by applying a multi-stage convolutional neural network trained to predict relationships between phoneme sequences and business domains.

15. The computer-implemented system of claim 11 , wherein the similarity detection module is configured to detect the sections of the one or more documents that are similar to the content of the call by applying a Siamese bidirectional long short term memory (LSTM) network model to capture phrase similarity using phoneme embedding.

16. The computer-implemented system of claim 15 , wherein the similarity detection module trains the Siamese LSTM network model to detect similarities in phoneme representations of text.

17. The computer-implemented system of claim 11 , wherein the user interface comprises:

a chat transcription section displaying in real time client-side conversation to the agent; and

a domain section identifying the at least one business category pertinent to the conversation displayed in the chat transcription section.

18. The computer-implemented system of claim 17 , wherein the user interface further comprises a Top Results section configured to display links to the documents in the at least one pertinent business category that include the similar content, wherein the links are ranked in accordance with the corresponding degrees of relevancy between the documents and the content of the call.

19. The computer-implemented system of claim 18 , wherein the similar content of each linked document is highlighted within each document.

20. The computer-implemented system of claim 18 , wherein the user interface is configured to update one or more of the chat transcription section, the domain section and the Top Results section as the call progresses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2023
From: SRIDHAR, VAISHNAVI MYSORE; GUPTA, ANUPRIYA; NARULKAR, GUNJAN; JAISWAL, ANKUR; SELVAKUMAR, SUWETHA; PAVITHRA, MARI
To: FMR LLC
Reel/Frame 062480/0174 →
Continuity (1)
Related Publication 20240212684A1 · Jun 27, 2024
References Cited (9)
US 20090216740A1 · Ramakrishnan · 2009 [cited by examiner]
US 20160275945A1 · Elisha · 2016 [cited by examiner]
US 20210312900A1 · Lalithsena · 2021 [cited by examiner]
US 20210350795A1 · Kenter · 2021 [cited by examiner]
US 20220156298A1 · Mahmoud · 2022 [cited by examiner]
“Cmusphinx / G2pse2seq”, GitHub, Inc., Retrieved from the Internet: <https://github.com/cmusphinx/g2p-seq2seq>, Retrieved from the Internet on: Mar. 23, 2023, pp. 1-6. [cited by applicant]
“NVIDIA / NeMo”, GitHub, Inc., Retrieved from the Internet: <https://github.com/NVIDIA/NeMo>, Retrieved from the Internet on: Mar. 23, 2023, pp. 1-6. [cited by applicant]
“Phonemes”, Wikipedia, Retrieved from the Internet: <https://en.wikipedia.org/wiki/Phoneme>, Retrieved from the Internet on: Mar. 23, 2023, pp. 1-13. [cited by applicant]
“Siamese Neural Network”, Wikipedia, Retrieved from the Internet: <https://en.wikipedia.org/wiki/Siamese_neural_network>, Retrieved from the Internet on: Mar. 23, 2023, pp. 1-4. [cited by applicant]