IP Library › Granted Patent US 12,609,124
Granted Patent B2
US 12,609,124 · App. 18/446,635 · Granted Apr 21, 2026

Voice agent system

Inventors: Aaron J. Neustedter (Milwaukee, WI); Thong T. Nguyen (New Berlin, WI); Paul D. Schmirler (Glendale, WI)
Assignee: Rockwell Automation Technologies, Inc.
G10L17/22G10L17/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,609,124
App. No.
18/446,635
Granted
Apr 21, 2026
Kind
B2
Abstract

An illustrative method includes a voice agent system establishing a plurality of user-agent conversations, wherein each user-agent conversation between a user and the voice agent system is established in response to the user speaking a trigger word and the plurality of user-agent conversations continue simultaneously in a same physical area, detecting an utterance in an audio stream associated with the physical area, determining, based on the utterance, that the utterance potentially belongs to a particular user-agent conversation among the plurality of user-agent conversations and determining a confidence score that the utterance belongs to the particular user-agent conversation, identifying a candidate action to be performed by the voice agent system based on the utterance, determining an overall confidence score of the candidate action based on the confidence score that the utterance belongs to the particular user-agent conversation, and performing an operation based on the overall confidence score of the candidate action.

Claims (91)

1 . A method comprising:

establishing, by a voice agent system, a plurality of user-agent conversations, wherein:

each user-agent conversation between a user and the voice agent system is established in response to the user speaking a trigger word and the plurality of user-agent conversations continue simultaneously in a same physical area, and

the voice agent system is communicatively coupled to one or more industrial devices in one or more industrial automation systems;

detecting, by the voice agent system, an utterance in an audio stream associated with the same physical area;

determining, by the voice agent system and based on the utterance, that the utterance potentially belongs to a particular user-agent conversation among the plurality of user-agent conversations and determining a confidence score that the utterance belongs to the particular user-agent conversation;

identifying, by the voice agent system, a candidate action to be performed by the voice agent system based on the utterance, wherein the candidate action is associated with an industrial device of the one or more industrial devices;

retrieving device information associated with the industrial device;

determining, by the voice agent system, an overall confidence score of the candidate action based at least on:

the confidence score that the utterance belongs to the particular user-agent conversation, and

the retrieved device information; and

performing, by the voice agent system, an operation on the industrial device based on the overall confidence score of the candidate action.

2 . The method of claim 1 , wherein:

the user-agent conversation between the user and the voice agent system is associated with a voice signature profile of the user.

3 . The method of claim 1 , wherein determining that the utterance potentially belongs to the particular user-agent conversation includes:

determining a voice signature of the utterance; and

determining that the voice signature of the utterance matches a voice signature profile associated with the particular user-agent conversation.

4 . The method of claim 1 , wherein determining the confidence score that the utterance belongs to the particular user-agent conversation includes:

determining a voice signature of the utterance; and

determining, with a first confidence score, that the voice signature of the utterance matches a voice signature profile associated with the particular user-agent conversation.

5 . The method of claim 4 , wherein determining the confidence score that the utterance belongs to the particular user-agent conversation includes:

determining an utterance content of the utterance; and

determining, with a second confidence score, that the utterance content of the utterance is relevant to a topic of the particular user-agent conversation.

6 . The method of claim 5 , wherein determining the confidence score that the utterance belongs to the particular user-agent conversation includes:

determining a user orientation of a particular user associated with the particular user-agent conversation at an utterance timestamp of the utterance; and

determining, with a third confidence score, that the particular user orientates towards a different person at the utterance timestamp of the utterance based on the user orientation of the particular user.

7 . The method of claim 6 , wherein determining the confidence score that the utterance belongs to the particular user-agent conversation includes:

determining a weighted average value of the first confidence score that the voice signature of the utterance matches the voice signature profile associated with the particular user-agent conversation, the second confidence score that the utterance content of the utterance is relevant to the topic of the particular user-agent conversation, and the third confidence score that the particular user orientates towards the different person at the utterance timestamp of the utterance; and

determining the confidence score that the utterance belongs to the particular user-agent conversation to be the weighted average value.

8 . The method of claim 1 , further comprising:

determining, by the voice agent system, that the confidence score that the utterance belongs to the particular user-agent conversation satisfies a confidence score threshold; and

including, by the voice agent system in response to determining that the confidence score that the utterance belongs to the particular user-agent conversation satisfies the confidence score threshold, the utterance in the particular user-agent conversation.

9 . The method of claim 1 , wherein identifying the candidate action includes:

determining an utterance content of the utterance; and

identifying the candidate action to be performed by the voice agent system based on the utterance content of the utterance.

10 . The method of claim 1 , wherein determining the overall confidence score of the candidate action includes:

determining that the confidence score that the utterance belongs to the particular user-agent conversation satisfies a confidence score threshold; and

adjusting, in response to determining that the confidence score that the utterance belongs to the particular user-agent conversation satisfies the confidence score threshold, the overall confidence score of the candidate action by a predefined amount.

11 . The method of claim 1 , wherein determining the overall confidence score of the candidate action includes:

determining a confidence score of an utterance content of the utterance;

determining an emotional distress level of the utterance;

determining a user location of a particular user associated with the particular user-agent conversation relative to the industrial device; and

wherein determining the overall confidence score of the candidate action is further based on one or more of the confidence score of the utterance content of the utterance, the emotional distress level of the utterance, the device information associated with the industrial device, and the user location of the particular user relative to the industrial device.

12 . The method of claim 1 , wherein performing the operation based on the overall confidence score of the candidate action includes one of:

performing, in response to determining that the overall confidence score of the candidate action satisfies a first overall confidence score threshold, the candidate action;

ignoring, in response to determining that the overall confidence score of the candidate action does not satisfy a second overall confidence score threshold, the utterance without performing the candidate action; or

requesting, in response to determining that the overall confidence score of the candidate action satisfies the second overall confidence score threshold and does not satisfy the first overall confidence score threshold, a user confirmation of the candidate action from a particular user associated with the particular user-agent conversation.

13 . The method of claim 1 , further comprising:

detecting, by the voice agent system, a different utterance in the audio stream associated with the physical area;

determining, by the voice agent system, a voice signature of the different utterance;

determining, by the voice agent system, that the different utterance is not spoken by a plurality of users associated with the plurality of user-agent conversations based on the voice signature of the different utterance and a plurality of voice signature profiles associated with the plurality of user-agent conversations; and

ignoring, by the voice agent system and in response to determining that the different utterance is not spoken by the plurality of users associated with the plurality of user-agent conversations, the different utterance.

14 . The method of claim 13 , further comprising:

determining, by the voice agent system, that an utterance content of the different utterance is relevant to a topic of the particular user-agent conversation associated with a particular user among the plurality of user-agent conversations; and

using, by the voice agent system and in response to determining that the utterance content of the different utterance is relevant to the topic of the particular user-agent conversation, the utterance content of the different utterance in processing a subsequent utterance in the particular user-agent conversation.

15 . The method of claim 14 , wherein:

the voice agent system detects the different utterance in a conversation between the particular user located at the physical area and a different person located remotely from the physical area via a user device of the particular user.

16 . A voice agent system comprising:

a memory storing instructions; and

a processor communicatively coupled to the memory and configured to execute the instructions to:

establish a plurality of user-agent conversations, wherein:

each user-agent conversation between a user and the voice agent system is established in response to the user speaking a trigger word and the plurality of user-agent conversations continue simultaneously in a same physical area, and

the voice agent system is communicatively coupled to one or more industrial devices in one or more industrial automation systems;

detect an utterance in an audio stream associated with the same physical area;

determine, based on the utterance, that the utterance potentially belongs to a particular user-agent conversation among the plurality of user-agent conversations and determine a confidence score that the utterance belongs to the particular user-agent conversation;

identify a candidate action to be performed by the voice agent system based on the utterance, wherein the candidate action is associated with an industrial device of the one or more industrial devices;

retrieve device information associated with the industrial device;

determine an overall confidence score of the candidate action based at least on:

the confidence score that the utterance belongs to the particular user-agent conversation, and

the retrieved device information; and

perform an operation on the industrial device based on the overall confidence score of the candidate action.

17 . The voice agent system of claim 16 , wherein:

the user-agent conversation between the user and the voice agent system is associated with a voice signature profile of the user.

18 . The voice agent system of claim 16 , wherein determining that the utterance potentially belongs to the particular user-agent conversation includes:

determining a voice signature of the utterance; and

determining that the voice signature of the utterance matches a voice signature profile associated with the particular user-agent conversation.

19 . The voice agent system of claim 16 , wherein the processor is further configured to execute the instructions to:

determine that the confidence score that the utterance belongs to the particular user-agent conversation satisfies a confidence score threshold; and

include, in response to determining that the confidence score that the utterance belongs to the particular user-agent conversation satisfies the confidence score threshold, the utterance in the particular user-agent conversation.

20 . A non-transitory computer-readable medium storing instructions that, when executed, direct a processor of a voice agent system to:

establish a plurality of user-agent conversations, wherein:

each user-agent conversation between a user and the voice agent system is established in response to the user speaking a trigger word and the plurality of user-agent conversations continue simultaneously in a same physical area, and

the voice agent system is communicatively coupled to one or more industrial devices in one or more industrial automation systems;

detect an utterance in an audio stream associated with the same physical area;

determine, based on the utterance, that the utterance potentially belongs to a particular user-agent conversation among the plurality of user-agent conversations and determine a confidence score that the utterance belongs to the particular user-agent conversation;

identify a candidate action to be performed by the voice agent system based on the utterance, wherein the candidate action is associated with an industrial device of the one or more industrial devices;

retrieve device information associated with the industrial device;

determine an overall confidence score of the candidate action based at least on:

the confidence score that the utterance belongs to the particular user-agent conversation, and

the retrieved device information; and

perform an operation on the industrial device based on the overall confidence score of the candidate action.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2023
From: NEUSTEDTER, AARON J.; NGUYEN, THONG T.; SCHMIRLER, PAUL D.
To: ROCKWELL AUTOMATION TECHNOLOGIES, INC.
Reel/Frame 064534/0393 →
Continuity (1)
Related Publication 20250054501A1 · Feb 13, 2025
References Cited (20)
US 20200073367A1 · Nguyen · 2020 [cited by examiner]
US 20200251107A1 · Wang · 2020 [cited by examiner]
US 20200272690A1 · Howard · 2020 [cited by examiner]
US 20200348651A1 · Nguyen · 2020 [cited by applicant]
US 20210065020A1 · Schmirler · 2021 [cited by applicant]
US 20220093094A1 · Krishnan · 2022 [cited by examiner]
US 20220253044A1 · Tremblay · 2022 [cited by applicant]
US 20240212689A1 · Mohammad · 2024 [cited by examiner]
“AI chatbot that's easy to use”, https://www.ibm.com/products/watson-assistant/artificial-intelligence?utm_content=SRCWW&p1=Search&p4=43700074369651641&p5=p&&msclkid=81d683c118a1141023f0e7c739e0c0f0&gclid=81d683c118a114… [cited by applicant]
“Bot in the Bunch: Facilitating Group Chat Discussion by Improving Efficiency and Participation with a Chatbot”, https://www.researchgate.net/publication/339438103_Bot_in_the_Bunch_Facilitating_Group_Chat_Discussion_by_… [cited by applicant]
“Channel and Group chat conversations with a Microsoft Teams bot”, https://learn.microsoft.com/en-us/microsoftteams/platform/resources/bot-v3/bot-conversations/bots-conv-channel, Microsoft.com, Nov. 25, 2022, 7 pages. [cited by applicant]
“Conversation Mode helps interactions with Alexa feel more natural”, https://www.aboutamazon.com/news/devices/conversation-mode-helps-interactions-with-alexa-feel-more-natural, Amazon.com, Written by Amazon Staff, Sep. … [cited by applicant]
“Have a conversation with your speaker or display”, https://support.google.com/googlenest/answer/7685981?hl=en&co=genie.platform%3dandroid, google.com, accessed Mar. 16, 2023, 2 pages. [cited by applicant]
“Have more natural conversations with Google Assistant”, https://blog.google/products/assistant/assistant-io-2022, google.com, Sissie Hsiao, May 11, 2022, 3 pages. [cited by applicant]
“New Alexa feature enables natural, multiparty interactions”, Amazon Science, https://www.amazon.science/blog/new-alexa-feature-enables-natural-multiparty-interactions, Alexa AI Team, Nov. 18, 2021, 8 pages. [cited by applicant]
“Set up voice recognition on HomePod or HomePod Mini”, https://support.apple.com/en-us/ht204753, apple.com, accessed Mar. 16, 2023, 3 pages. [cited by applicant]
“Start a conversation with your personal productivity assistant in Outlook with Cortana”, https://techcommunity.microsoft.com/t5/outlook-blog/start-a-conversation-with-your-personal-productivity-assistant/ba-p/2071416, … [cited by applicant]
“Voice Match and media on shared Google Nest or Home devices”, https://support.google.com/googlenest/answer/7342711?hl=en&co=genie.platform%3dandroid, google.com, accessed Mar. 16, 2023, 3 pages. [cited by applicant]
“What Is Alexa Voice ID?”, https://www.amazon.com/gp/help/customer/display.html?nodeld=GYCXKY2AB2QWZT2X, Amazon.com, accessed Mar. 16, 2023, 2 pages. [cited by applicant]
“What's Microsoft's vision for conversational AI? Computers that understand you”, https://news.microsoft.com/source/features/ai/microsoft-build-future-of-natural-language, microsoft.com, John Roach, May 6, 2019, 9 pages. [cited by applicant]