IP Library Granted Patent US 12,374,340
Granted Patent B2
US 12,374,340 · App. 17/806,286 · Granted Jul 29, 2025

Individual recognition using voice detection

Inventors: Ethan S. Headings (Columbus, OH); Feng-wei Chen (Cary, NC); Neha S Deshpande (Cary, NC); Madhavi Kolachala (Cary, NC)
Assignee: International Business Machines Corporation
G10L17/22G10L17/02G10L17/06G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,374,340
App. No.
17/806,286
Granted
Jul 29, 2025
Kind
B2
Abstract

A method, a computer program product, and a computer system determine a name of an individual based on voice detection. The method includes determining voice characteristics of a voice of the individual based on an audio input received and recorded via an audio input device. The method includes comparing the voice characteristics of the voice to further voice characteristics of voice profiles. The method includes as a result of the voice and one of the voice profiles meeting a similarity threshold, determining a name associated with the one of the voice profile. The method includes providing the name to a user who is having a conversation with the individual.

Claims (59)

1. A computer-implemented method for determining a name of an individual based on voice detection, the computer-implemented method comprising:

receiving, by an audio input device, an audio input associated with a conversation between a user and the individual;

recording, by the audio input device, the received audio input;

determining voice characteristics of a voice of the individual based on the recorded audio input, wherein the audio input includes a decibel level of the conversation;

instructing, based on the decibel level of the conversation, the audio input device to change from a passive state to an active state;

determining that an amity level of the conversation is greater than a threshold value, wherein

the amity level is one of an acquaintance or a friend, and

the amity level is determined based on:

the instructing of the audio input device to change from the passive state to the active state; and

contextual information associated with specific words or phrases used in the conversation, the voice characteristics of the individual, and at least one of frequency of meetings between the user and the individual, a time since previous meeting between the user and the individual, or a time since first meeting between the user and the individual;

comparing, based on the determination that the amity level is greater than the threshold value, the voice characteristics of the voice of the individual to voice characteristics of voice profiles;

as a result of the voice and one of the voice profiles meeting a similarity threshold, determining the name associated with one of the voice profiles; and

providing the name to the user who is having the conversation with the individual.

2. The computer-implemented method of claim 1 , further comprising:

monitoring an audio space in proximity to the user; and

detecting a presence of speech in the audio input captured in the audio space, wherein the speech is associated with the conversation.

3. The computer-implemented method of claim 2 , further comprising

determining whether the conversation includes a named greeting, the named greeting being an utterance from one of the users or the individual including the name of the individual, wherein the voice characteristics are determined as a result of the named greeting being absent in the conversation.

4. The computer-implemented method of claim 1 , wherein the voice characteristics include a tone, a pitch, and a dialect.

5. A computer-readable storage media that configures a computer to perform program instructions stored on the computer-readable storage media for determining a name of an individual based on voice detection, the program instructions comprising:

receiving, by an audio input device, an audio input associated with a conversation between a user and the individual;

recording, by the audio input device, the received audio input;

determining voice characteristics of a voice of the individual based on the recorded audio input, wherein the audio input includes a decibel level of the conversation;

instructing, based on the decibel level of the conversation, the audio input device to change from a passive state to an active state;

determining that an amity level of the conversation is greater than a threshold value, wherein

the amity level is one of an acquaintance or a friend, and

the amity level is determined based on:

the instructing of the audio input device to change from the passive state to the active state; and

contextual information associated with specific words or phrases used in the conversation, the voice characteristics of the individual, and at least one of frequency of meetings between the user and the individual, a time since previous meeting between the user and the individual, or a time since first meeting between the user and the individual;

comparing, based on the determination that the amity level is greater than the threshold value, the voice characteristics of the voice of the individual to voice characteristics of voice profiles;

as a result of the voice and one of the voice profiles meeting a similarity threshold, determining the name associated with one of the voice profiles; and

providing the name to the user who is having the conversation with the individual.

6. The computer-readable storage media of claim 5 , wherein the program instructions further comprise:

monitoring an audio space in proximity to the user; and

detecting a presence of speech in the audio input captured in the audio space, wherein the speech is associated with the conversation.

7. The computer-readable storage media of claim 6 , wherein the program instructions further comprise:

determining whether the conversation includes a named greeting, the named greeting being an utterance from one of the users or the individual including the name of the individual, wherein the voice characteristics are determined as a result of the named greeting being absent in the conversation.

8. The computer-readable storage media of claim 5 , wherein the voice characteristics include a tone, a pitch, and a dialect.

9. A computer system to determine a name of an individual based on voice detection, the computer system comprising:

one or more computer processors;

one or more computer-readable storage media; and

program instructions stored on the one or more of the computer-readable storage media, the program instructions executable by at least one processor of the one or more computer processors to cause the at least one processor;

receive, by an audio input device, an audio input associated with a conversation between a user and the individual;

record, by the audio input device, the received audio input;

determine voice characteristics of a voice of the individual based on the recorded audio input, wherein the audio input includes a decibel level of the conversation;

instruct, based on the decibel level of the conversation, the audio input device to change from a passive state to an active state;

determine that an amity level of the conversation is greater than a threshold value, wherein

the amity level is one of an acquaintance or a friend, and

the amity level is determined based on:

the instruction to the audio input device to change from the passive state to the active state; and

contextual information associated with specific words or phrases used in the conversation, the voice characteristics of the individual, and at least one of frequency of meetings between the user and the individual, a time since previous meeting between the user and the individual, or a time since first meeting between the user and the individual;

compare, based on the determination that the amity level is greater than the threshold value, the voice characteristics of the voice of the individual to voice characteristics of voice profiles;

as a result of the voice and one of the voice profiles meeting a similarity threshold, determine the name associated with one of the voice profiles; and

providing the name to the user who is having the conversation with the individual.

10. The computer system of claim 9 , wherein the program instructions further cause the at least one processor to:

monitor an audio space in proximity to the user; and

detect a presence of speech in the audio input captured in the audio space, wherein the speech is associated with the conversation.

11. The computer system of claim 10 , wherein the program instructions further cause the at least one processor to:

determine whether the conversation includes a named greeting, the named greeting being an utterance from one of the users or the individual including the name of the individual, wherein the voice characteristics are determined as a result of the named greeting being absent in the conversation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2022
From: HEADINGS, ETHAN S.; CHEN, FENG-WEI; DESHPANDE, NEHA S.; KOLACHALA, MADHAVI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 060158/0373 →
Continuity (1)
Related Publication 20230402041A1 · Dec 14, 2023
References Cited (24)
US 8571865B1 · Hewinson · 2013 [cited by examiner]
US 10217465B2 · Grahm · 2019 [cited by examiner]
US 10460728B2 · Anbazhagan · 2019 [cited by applicant]
US 11094316B2 · Visser · 2021 [cited by applicant]
US 20080043996A1 · Dolph · 2008 [cited by examiner]
US 20140188846A1 · Kurabayashi · 2014 [cited by examiner]
US 20160098992A1 · Renard · 2016 [cited by examiner]
US 20160329053A1 · Grahm · 2016 [cited by examiner]
US 20210043216A1 · Wang · 2021 [cited by examiner]
EP 2899609A1 · 2015 [cited by applicant]
WO 2009094415A1 · 2009 [cited by applicant]
Amazon, “BytNotes”, https://www.amazon.com/MSGA-bytNotes/dp/B004Q7F7NS, https://www.amazon.com/MSGA-bytNotes/dp/B004Q7F7NS, accessed Mar. 29, 2022, pp. 1-4. [cited by applicant]
Apple, “Remember App”, https://itunes.apple.com/us/app/remember-app-remember-everyone-you-met/id1096350872?mt=8, accessed Mar. 29, 2022, pp. 1-4. [cited by applicant]
Contacts Journal, “Contacts Journal CRM”, http://www.contactsjournal.com/, accessed Mar. 29, 2022, pp. 1-9. [cited by applicant]
Das et al., “Voice Recognition System, Speech-To-Text”, https://www.researchgate.net/publication/304651244, Jul. 1, 2018, pp. 1-6. [cited by applicant]
Disclosed Anonymously, “Method for Identifying IoT Devices to Enable Seamless Transfer of Chat Sessions”, IPCOM000263406D; Aug. 27, 2020, pp. 1-5. [cited by applicant]
Disclosed Anonymously, “Dynamic Adjustment of Hotword Detection Threshold”, IPCOM000257597D; Feb. 22, 2019, pp. 1-7. [cited by applicant]
Disclosed Anonymously, “Intelligent Voice Assistant Extended Through Voice Relay System”, IPCOM000255132D; Sep. 4, 2018, pp. 1-23. [cited by applicant]
http://namerick.com/, “Remember Names”, accessed Mar. 29, 2022, pp. 1-3. [cited by applicant]
http://namesharkapp.com/, “Name Shark”, Mar. 29, 2022, pp. 1-6. [cited by applicant]
http://www.nameorize.com/, “Namorize”, accessed Mar. 29, 2022, p. 1. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, pp. 1-7. [cited by applicant]
Mphego, “Designing a Smart Home Automation With Voice Recognition Using a Rasberry PI and Arduino”, https://dev.to/mmphego/smart-home-automation-using-raspberry-pi-and-arduino-4k71, Jun. 13, 2018, pp. 1-116. [cited by applicant]
Sudharsan, et al., “Smart Speaker Design and Implementation With Biometric Authentication and Advanced Voice Interaction Capability”, https://www.researchgate.net/publication/338984616, Dec. 2019, pp. 1-14. [cited by applicant]