IP Library Granted Patent US 12,462,812
Granted Patent B2
US 12,462,812 · App. 18/061,499 · Granted Nov 4, 2025

Relationship-driven virtual assistant

Inventors: Ashok Kumar Iyengar (Encinitas, CA); Trudy L. Hewitt (Cary, NC); Venkata Vishwanath Gadepalli (Apex, NC); Jeremy R. Fox (Georgetown, TX)
Assignee: International Business Machines Corporation
G10L17/22G10L13/027G10L25/54
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,812
App. No.
18/061,499
Granted
Nov 4, 2025
Kind
B2
Abstract

A computer-implemented method, a computer system and a computer program product generate a query response in an environment based on predicted relationships between users. The method includes capturing a question with a device in the environment, where the question includes a speaking voice and is selected from a group consisting of: video data, audio data and text data. The method also includes identifying the speaker of the question based on the speaking voice. The method further includes determining a relationship between the question and each user interaction in a database of user interactions, where each user interaction is associated with a user. In addition, the method includes selecting a response from the database of user interactions based on the relationship. Lastly, the method includes transmitting the response to the speaker of the question in the environment, where the transmission of the response uses a voice of the user.

Claims (56)

1 . A computer-implemented method for generating a query response in an environment based on predicted relationships between users, the computer-implemented method comprising:

capturing a question with a device in the environment, wherein the question includes a speaking voice and is selected from a group consisting of: video data, audio data and text data;

identifying a speaker of the question based on the speaking voice;

using a supervised machine learning model, generating a relationship map based on historic interactions associated with a user and other related information stored in a database;

analyzing, using the supervised machine model having a speech recognition algorithm, a tone of speech, voice print or voice quality for comparison to a previous user interaction stored in said database, wherein said stored information in said database has been recorded over time and contains audio, video and/or text questions and answers previously obtained;

determining a relationship between the question and each user interaction based on the relationship map generated;

selecting a response from the database of user interactions based on the relationship; and

transmitting the response to the environment, wherein transmission of the response uses a voice of the user.

2 . The computer-implemented method of claim 1 , wherein the selecting the response from the database of user interactions further comprises:

calculating a confidence score for each user interaction in the database of user interactions based on the relationship and the user associated with a user interaction; and

selecting the user interaction as the response when the confidence score for the user interaction is above a threshold.

3 . The computer-implemented method of claim 1 , further comprising:

monitoring an interaction between the speaker of the question and the response; and

updating the database of user interactions based on the interaction between the speaker of the question and the response.

4 . The computer-implemented method of claim 1 , further comprising

adding the question as the user interaction to the database of user interactions, wherein the user interaction is associated with the speaker of the question, when the relationship between the question and the user interaction in the database is not determined.

5 . The computer-implemented method of claim 1 , wherein the determining the relationship between the question and each user interaction in the database of user interactions includes determining the relationship between the speaker of the question and the user associated with the user interaction.

6 . The computer-implemented method of claim 1 , wherein the transmission of the response uses a hologram of the user speaking in the voice of the user.

7 . The computer-implemented method of claim 1 , wherein the transmission of the response uses an audible response in the voice of the user.

8 . A computer system for generating a query response in an environment based on predicted relationships between users, the computer system comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored on at least one of the one or more tangible storage media for execution by at least one of the one or more processors via at least one of the one or more memories, wherein the computer system is capable of performing a method comprising:

capturing a question with a device in the environment, wherein the question includes a speaking voice and is selected from a group consisting of: video data, audio data and text data;

identifying a speaker of the question based on the speaking voice;

using a supervised machine learning model, generating a relationship map based on historic interactions associated with a user and other related information stored in a database;

analyzing, using the supervised machine model having a speech recognition algorithm, a tone of speech, voice print or voice quality for comparison to a previous user interaction stored in said database, wherein said stored information in said database has been recorded over time and contains audio, video and/or text questions and answers previously obtained;

determining a relationship between the question and each user interaction based on the relationship map generated;

selecting a response from the database of user interactions based on the relationship; and transmitting the response to the environment, wherein transmission of the response uses a voice of the user.

9 . The computer system of claim 8 , wherein the selecting the response from the database of user interactions further comprises:

calculating a confidence score for each user interaction in the database of user interactions based on the relationship and the user associated with the user interaction; and

selecting the user interaction as the response when the confidence score for the user interaction is above a threshold.

10 . The computer system of claim 8 , further comprising:

monitoring an interaction between the speaker of the question and the response; and

updating the database of user interactions based on the interaction between the speaker of the question and the response.

11 . The computer system of claim 8 , further comprising

adding the question as the user interaction to the database of user interactions, wherein the user interaction is associated with the speaker of the question, when the relationship between the question and the user interaction in the database is not determined.

12 . The computer system of claim 8 , wherein the determining the relationship between the question and each user interaction in the database of user interactions includes determining the relationship between the speaker of the question and the user associated with the user interaction.

13 . The computer system of claim 8 , wherein the transmission of the response uses a hologram of the user speaking in the voice of the user.

14 . The computer system of claim 8 , wherein the transmission of the response uses an audible response in the voice of the user.

15 . A computer program product for generating a query response in an environment based on predicted relationships between users, the computer program product comprising:

a computer-readable storage device having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform a method comprising:

capturing a question with a device in the environment, wherein the question includes a speaking voice and is selected from a group consisting of: video data, audio data and text data;

identifying a speaker of the question based on the speaking voice;

using a supervised machine learning model, generating a relationship map based on historic interactions associated with a user and other related information stored in a database;

analyzing, using the supervised machine model having a speech recognition algorithm, a tone of speech, voice print or voice quality for comparison to a previous user interaction stored in said database, wherein said stored information in said database has been recorded over time and contains audio, video and/or text questions and answers previously obtained;

determining a relationship between the question and each user interaction based on the relationship map generated;

selecting a response from the database of user interactions based on the relationship; and transmitting the response to the environment, wherein transmission of the response uses a voice of the user.

16 . The computer program product of claim 15 , wherein the selecting the response from the database of user interactions further comprises:

calculating a confidence score for each user interaction in the database of user interactions based on the relationship and the user associated with a user interaction; and

selecting the user interaction as the response when the confidence score for the user interaction is above a threshold.

17 . The computer program product of claim 15 , further comprising:

monitoring an interaction between the speaker of the question and the response; and

updating the database of user interactions based on the interaction between the speaker of the question and the response.

18 . The computer program product of claim 15 , further comprising

adding the question as the user interaction to the database of user interactions, wherein the user interaction is associated with the speaker of the question, when the relationship between the question and the user interaction in the database is not determined.

19 . The computer program product of claim 15 , wherein the determining the relationship between the question and each user interaction in the database of user interactions includes determining the relationship between the speaker of the question and the user associated with the user interaction.

20 . The computer program product of claim 15 , wherein the transmission of the response uses a hologram of the user speaking in the voice of the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2022
From: IYENGAR, ASHOK KUMAR; HEWITT, TRUDY L.; GADEPALLI, VENKATA VISHWANATH; FOX, JEREMY R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061967/0769 →
Continuity (1)
Related Publication 20240185862A1 · Jun 6, 2024
References Cited (25)
US 8156060B2 · Borzestowski · 2012 [cited by applicant]
US 9313646B2 · Baldwin · 2016 [cited by applicant]
US 9634855B2 · Poltorak · 2017 [cited by applicant]
US 9864431B2 · Keskin · 2018 [cited by applicant]
US 10152719B2 · Navaratnam · 2018 [cited by examiner]
US 10628635B1 · Carpenter, II · 2020 [cited by examiner]
US 10878174B1 · Vontobel · 2020 [cited by examiner]
US 20150111607A1 · Baldwin · 2015 [cited by applicant]
US 20160196491A1 · Chandrasekaran · 2016 [cited by examiner]
US 20170308905A1 · Navaratnam · 2017 [cited by applicant]
US 20170329404A1 · Keskin et al. · 2017 [cited by applicant]
US 20190005024A1 · Somech · 2019 [cited by examiner]
US 20190156222A1 · Emma · 2019 [cited by examiner]
US 20190220727A1 · Dohrmann · 2019 [cited by applicant]
US 20220036013A1 · Liu · 2022 [cited by examiner]
Author Unknown, “Create an Alexa skill using Watson Assistant and OpenWhisk”, https://github.com/IBM/alexa-skill-watson-assistant, Accessed Aug. 17, 2022, pp. 1-24. [cited by applicant]
Author Unknown, “Gatebox”, Gatebox Inc.—Gatebox, https://www.gatebox.ai/en, Accessed Aug. 17, 2022, pp. 1-6. [cited by applicant]
Author Unknown, “Gaze and commit”, Gaze and commit—Mixed Reality | Microsoft Docs, https://docs.microsoft.com/en-us/windows/mixed-reality/design/gaze-and-commit#composite-gestures, Aug. 4, 2022, pp. 1-15. [cited by applicant]
Author Unknown, “speech recognition”, Speech recognition—Windows apps | Microsoft Docs, https://docs.microsoft.com/en-us/windows/apps/design/input/speech-recognition, Jun. 24, 2021, pp. 1-13. [cited by applicant]
Author Unknown, “Voice input”, Voice input—Mixed Reality | Microsoft Docs, https://docs.microsoft.com/en-us/windows/mixed-reality/design/voice-input, Mar. 7, 2022, pp. 1-16. [cited by applicant]
Author Unknown, “Watson Assistant: Intelligent virtual agent”, Virtual Agent—IBM Watson Assistant IBM, https://www.ibm.com/products/watson-assistant, Accessed Aug. 17, 2022, pp. 1-13. [cited by applicant]
Johnson, “Humanizing digital communications”, Humanizing digital communications | IBM, https://www.ibm.com/case-studies/2mee/, Accessed Aug. 17, 2022, pp. 1-14. [cited by applicant]
KH, “3 Efficient Ways to Supply Your App with a Virtual Assistant”, How to Create Virtual Assistant Apps like Siri and Google Assistant, https://www.cleveroad.com/blog/how-to-create-virtual-assistant-apps-like-siri-and-… [cited by applicant]
Koetsier, “This No-Headset Holographic Display Enables 3D Presence, Remotely”, Consumer Tech, https://www.forbes.com/sites/johnkoetsier/2021/09/11/this-no-headset-holographic-display-enables-3d-presence-remotely/?sh=7c8… [cited by applicant]
Kramar et al. “Augmented Reality-assisted Cyber-Physical Systems of Smart University Campus.”, https://ieeexplore.ieee.org/document/9321951, 2020 IEEE 15th International Conference on Computer Sciences and Information T… [cited by applicant]