IP Library Granted Patent US 12,322,397
Granted Patent B2
US 12,322,397 · App. 18/147,902 · Granted Jun 3, 2025

Processing audio information captured by interactive virtual assistant

Inventors: Joseph Namm (Plantation, FL); Melanie King (Wesley Chapel, FL); Jeet K. Pawani (Sunrise, FL)
Assignee: MOTOROLA SOLUTIONS, INC.
G10L17/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,397
App. No.
18/147,902
Granted
Jun 3, 2025
Kind
B2
Abstract

Methods and systems for processing audio information captured by an interactive virtual assistant (IVA). An example method includes: converting sound received by a microphone into electrical signals; generating voiceprints corresponding to voices represented in the electrical signals; distinguishing signals representing background speech and signals representing voice commands directed at the IVA; and based on various signals representing background speech and voice commands, labeling each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag. The method also includes transmitting, in response to a trigger event, a message with at least one of an estimated number of occupants in a geofenced area corresponding to the IVA and an alert reporting an indication of unlawful activity thereat.

Claims (108)

1. An apparatus implementing an interactive virtual assistant (IVA), the apparatus comprising:

a microphone to convert received sound into electrical signals;

a communication interface; and

an electronic processor connected to the microphone and the communication interface and configured to:

generate a plurality of voiceprints corresponding to a plurality of voices represented in the electrical signals;

distinguish, in the electrical signals, signals representing background speech and signals representing voice commands directed at the IVA;

based on the signals representing background speech and the signals representing voice commands, label each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag; and

in response to a trigger event, transmit, through the communication interface, a message with at least one of:

an estimated number of occupants in a geofenced area corresponding to the IVA; and

an alert reporting an indication of unlawful activity in the background speech,

wherein the indication of unlawful activity is attributed to a first subset or a second subset of the plurality of voiceprints, the first subset being labeled with the IVA interactor occupant tag, the second subset being labeled with the non-occupant tag.

2. The apparatus of claim 1 ,

wherein the electronic processor is further configured to convert a signal representing a voice command into a corresponding device command; and

wherein the apparatus is configured to execute the device command.

3. The apparatus of claim 1 , wherein, for a voiceprint labeled with the IVA interactor occupant tag or the IVA non-interactor occupant tag, the electronic processor is configured to determine a characteristic selected from the group consisting of a corresponding person's age, gender, accent, dialect, and speech abnormality.

4. The apparatus of claim 1 , wherein the electronic processor is configured to compute the estimated number of occupants by summing a number of voiceprints labeled with the IVA interactor occupant tag and a number of voiceprints labeled with the IVA non-interactor occupant tag.

5. The apparatus of claim 1 , wherein the trigger event is an event selected from the group consisting of:

a citizen's complaint regarding the geofenced area;

an emergency call regarding the geofenced area;

a warrant to monitor the geofenced area;

the estimated number of occupants being greater than a threshold value; and

a number of voiceprints labeled with the non-occupant tag being greater than another threshold value.

6. The apparatus of claim 1 , wherein the indication of unlawful activity includes at least one of:

an aggressive or threatening posture in the background speech attributed to a voiceprint of the first subset; and

higher than a threshold usage of a keyword in the background speech attributed to the first and second subsets.

7. The apparatus of claim 6 , wherein the keyword indicates coercion, manipulation, exploitation, or explicitly unlawful activity.

8. The apparatus of claim 1 , wherein the message is directed to a law-enforcement agency or a child welfare agency.

9. A method of processing information captured by an interactive virtual assistant (IVA), the method comprising:

converting sound received by a microphone into electrical signals;

generating, with an electronic processor, a plurality of voiceprints corresponding to a plurality of voices represented in the electrical signals;

distinguishing, with the electronic processor, in the electrical signals, signals representing background speech and signals representing voice commands directed at the IVA;

based on the signals representing background speech and the signals representing voice commands, labeling, with the electronic processor, each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag; and

in response to a trigger event, transmitting, through a communication interface connected to the electronic processor, a message with at least one of:

an estimated number of occupants in a geofenced area corresponding to the IVA; and

an alert reporting an indication of unlawful activity in the background speech,

wherein the indication of unlawful activity is attributed to a first subset or a second subset of the plurality of voiceprints, the first subset being labeled with the IVA interactor occupant tag, the second subset being labeled with the non-occupant tag.

10. The method of claim 9 , further comprising:

converting, with the electronic processor, a signal representing a voice command into a corresponding device command; and

causing the IVA to execute the device command.

11. The method of claim 9 , further comprising, for a voiceprint labeled with the IVA interactor occupant tag or the IVA non-interactor occupant tag, determining, with the electronic processor, a characteristic selected from the group consisting of a corresponding person's age, gender, accent, dialect, and speech abnormality.

12. The method of claim 9 , further comprising computing, with the electronic processor, the estimated number of occupants by summing a number of voiceprints labeled with the IVA interactor occupant tag and a number of voiceprints labeled with the IVA non-interactor occupant tag.

13. The method of claim 9 , wherein the trigger event is an event selected from the group consisting of:

a citizen's complaint regarding the geofenced area;

an emergency call regarding the geofenced area;

a warrant to monitor the geofenced area;

the estimated number of occupants being greater than a threshold value; and

a number of voiceprints labeled with the non-occupant tag being greater than another threshold value.

14. The method of claim 9 , wherein the trigger event is generated in response to detecting one or more predetermined elements associated with human trafficking within the plurality of voiceprints.

15. A non-transitory computer-readable medium storing instructions that, when executed by the electronic processor, cause the electronic processor to perform operations comprising the method of claim 9 .

16. An apparatus implementing an interactive virtual assistant (IVA), the apparatus comprising:

a microphone to convert received sound into electrical signals;

a communication interface; and

an electronic processor connected to the microphone and the communication interface and configured to:

generate a plurality of voiceprints corresponding to a plurality of voices represented in the electrical signals;

distinguish, in the electrical signals, signals representing background speech and signals representing voice commands directed at the IVA;

based on the signals representing background speech and the signals representing voice commands, label each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag;

in response to a trigger event, transmit, through the communication interface, a message with at least one of:

an estimated number of occupants in a geofenced area corresponding to the IVA; and

an alert reporting an indication of unlawful activity in the background speech; and

label a voiceprint with the IVA interactor occupant tag in response to:

counting a first number of instances in which said voiceprint is present in the background speech, the first number being greater than a threshold value; and

counting a second number of voice commands corresponding to said voiceprint, the second number being greater than another threshold value.

17. An apparatus implementing an interactive virtual assistant (IVA), the apparatus comprising:

a microphone to convert received sound into electrical signals;

a communication interface; and

an electronic processor connected to the microphone and the communication interface and configured to:

generate a plurality of voiceprints corresponding to a plurality of voices represented in the electrical signals;

distinguish, in the electrical signals, signals representing background speech and signals representing voice commands directed at the IVA;

based on the signals representing background speech and the signals representing voice commands, label each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag;

in response to a trigger event, transmit, through the communication interface, a message with at least one of:

an estimated number of occupants in a geofenced area corresponding to the IVA; and

an alert reporting an indication of unlawful activity in the background speech; and

label a voiceprint with the IVA non-interactor occupant tag in response to:

counting a first number of instances in which said voiceprint is present in the background speech, the first number being greater than a threshold value; and

counting a second number of voice commands corresponding to said voiceprint, the second number being smaller than another threshold value.

18. An apparatus implementing an interactive virtual assistant (IVA), the apparatus comprising:

a microphone to convert received sound into electrical signals;

a communication interface; and

an electronic processor connected to the microphone and the communication interface and configured to:

generate a plurality of voiceprints corresponding to a plurality of voices represented in the electrical signals;

distinguish, in the electrical signals, signals representing background speech and signals representing voice commands directed at the IVA;

based on the signals representing background speech and the signals representing voice commands, label each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag; and

in response to a trigger event, transmit, through the communication interface, a message with at least one of:

an estimated number of occupants in a geofenced area corresponding to the IVA; and

an alert reporting an indication of unlawful activity in the background speech,

wherein the trigger event is generated in response to detecting one or more predetermined elements associated with human trafficking within the plurality of voiceprints.

19. A method of processing information captured by an interactive virtual assistant (IVA), the method comprising:

converting sound received by a microphone into electrical signals;

generating, with an electronic processor, a plurality of voiceprints corresponding to a plurality of voices represented in the electrical signals;

distinguishing, with the electronic processor, in the electrical signals, signals representing background speech and signals representing voice commands directed at the IVA;

based on the signals representing background speech and the signals representing voice commands, labeling, with the electronic processor, each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag;

in response to a trigger event, transmitting, through a communication interface connected to the electronic processor, a message with at least one of:

an estimated number of occupants in a geofenced area corresponding to the IVA; and

an alert reporting an indication of unlawful activity in the background speech; and

causing the electronic processor to label a voiceprint with the IVA interactor occupant tag in response to:

counting a first number of instances in which said voiceprint is present in the background speech, the first number being greater than a threshold value; and

counting a second number of voice commands corresponding to said voiceprint, the second number being greater than another threshold value.

20. A method of processing information captured by an interactive virtual assistant (IVA), the method comprising:

converting sound received by a microphone into electrical signals;

generating, with an electronic processor, a plurality of voiceprints corresponding to a plurality of voices represented in the electrical signals;

distinguishing, with the electronic processor, in the electrical signals, signals representing background speech and signals representing voice commands directed at the IVA;

based on the signals representing background speech and the signals representing voice commands, labeling, with the electronic processor, each of the voiceprints with a tag selected from the group consisting of an IVA interactor occupant tag, an IVA non-interactor occupant tag, and a non-occupant tag;

in response to a trigger event, transmitting, through a communication interface connected to the electronic processor, a message with at least one of:

an estimated number of occupants in a geofenced area corresponding to the IVA; and

an alert reporting an indication of unlawful activity in the background speech; and

causing the electronic processor to label a voiceprint with the IVA non-interactor occupant tag in response to:

counting a first number of instances in which said voiceprint is present in the background speech, the first number being greater than a threshold value; and

counting a second number of voice commands corresponding to said voiceprint, the second number being smaller than another threshold value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2023
From: NAMM, JOSEPH; KING, MELANIE; PAWANI, JEET K.
To: MOTOROLA SOLUTIONS, INC.
Reel/Frame 062574/0499 →
Continuity (1)
Related Publication 20240221759A1 · Jul 4, 2024
References Cited (13)
US 8055534B2 · Ashby et al. · 2011 [cited by applicant]
US 10380852B2 · Horling · 2019 [cited by examiner]
US 11170089B2 · Goldstein et al. · 2021 [cited by applicant]
US 20180330589A1 · Horling · 2018 [cited by examiner]
US 20200035239A1 · Ni · 2020 [cited by examiner]
US 20210249033A1 · Hsu · 2021 [cited by examiner]
US 20210326421A1 · Khoury · 2021 [cited by examiner]
US 20220115032A1 · Iwagaki · 2022 [cited by examiner]
US 20240112681A1 · Rohatgi · 2024 [cited by examiner]
US 20240195905A1 · Guan · 2024 [cited by examiner]
JP 7204283B2 · 2023 [cited by examiner]
Dasgupta, et al., “Audio Analytics-based Human Trafficking Detection Framework for Autonomous Vehicles,” Alabama Transportation Institute, <https://arxiv.org/ftp/arxiv/papers/2209/2209.04071.pdf> webpage accessed Dec. 1… [cited by applicant]
Podder, et al., “Designing Intelligent Automation based Solutions for Complex Social Problems,” ICML Workshop on #Data4Good: Machine Learning in Social Good Applications, <https://arxiv.org/pdf/1606.05275.pdf> 2016 (5 p… [cited by applicant]