IP Library Granted Patent US 12,579,986
Granted Patent B2
US 12,579,986 · App. 18/602,835 · Granted Mar 17, 2026

Systems and methods for distinguishing between human speech and machine generated speech

Inventors: Ronen Reouveni (San Diego, CA); Robert Piro (Ellensburg, WA); Jonathan Wiggs (Seattle, WA); Christina Quinn (Portland, OR); Mohammed Soliman (Newton, MA)
Assignee: Outbound AI Inc.
G10L17/14G06F40/30G10L15/26G10L17/00G10L17/02G10L17/22H04M3/5166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,986
App. No.
18/602,835
Granted
Mar 17, 2026
Kind
B2
Abstract

Systems, devices, and methods for determining whether a segment of speech was generated by a human or by a machine, such as a robotic voice that is synthesized and used as part of an IVR system. The disclosed approach can be used to assist in implementing a process to automate the detection of the start and end of a hold time during a call to a call center and in response execute a desired action.

Claims (35)

1 . A method, comprising;

processing one or more speech segments generated by a first entity or by a second entity during a call placed by the first entity and received by the second entity by;

converting each segment of speech into text using an automatic speech recognition (ASR) system;

creating a processing window of a configurable size;

processing each word in the text in the configurable window using a text normalization or standardization process;

accessing a set of keywords determining if one or more of the keywords or are in the text in the configurable window, wherein the set of keywords represent a speech pattern or characteristics of a human speaker; and

automatically performing an indicated action if a segment of speech is determined to be speech generated by a human or automatically performing a different indicated action if the segment of speech is determined to be speech generated by a machine.

2 . The method of claim 1 , wherein the second entity is a call center which connects the call to an IVR system associated with the call center, wherein the IVR system generates one or more prompts in the form of speech segments that are navigated through to be connected to a human call center representative, and wherein after navigation through one or more prompts, the call is placed into an on-hold state by the call center.

3 . The method of claim 2 , wherein the speech segments are navigated through using a trained model.

4 . The method of 2 , wherein the speech segments are processed by a service provided to the first entity, and if a segment of speech is determined to be speech generated by a human, then the indicated action is to alert the first entity that the human call center representative is available.

5 . The method of claim 4 , further comprising:

in response to the generated alert, executing a trained classifier to classify a speech segment generated by the human call center representative, and in response to the classification, generate a speech segment for presentation to the human call center representative.

6 . The method of claim 4 , wherein the speech segments are processed by a service provided to the second entity.

7 . The method of claim 6 , wherein if the segment of speech is determined to be speech generated by a machine, then the indicated action is to prevent the call being routed to a call center representative.

8 . The method of claim 1 , wherein the first entity is an automated process, the second entity is a human, the speech segments are processed by a service provided to the human, and the indicated action is to alert the human if the speech segments are machine generated.

9 . The method of claim 1 , wherein the text normalization or standardization process is one of removing punctuation, removing stop words, or removing hesitation words.

10 . The method of claim 1 , wherein multiple speech segments in multiple calls are processed and used to determine a distribution of the duration of a set of calls or of a set of sections of a call.

11 . The method of claim 1 , wherein the size of the configurable processing window is no larger than a maximum keyword size for a speech segment or speech segments.

12 . The method of claim 11 , wherein the maximum keyword size is determined by a process that includes forming a set of n-grams based on the text generated from a speech segment or speech segments.

13 . A system, comprising:

one or more electronic processors configured to execute a set of computer-executable instructions; and

one or more non-transitory computer-readable media containing the set of computer-executable instructions, wherein when executed, the instructions cause the one or more electronic processors or a device or apparatus in which they are contained to

process one or more speech segments generated by a first entity or by a second entity during a call placed by the first entity and received by the second entity by;

converting each segment of speech into text using an automatic speech recognition (ASR) system;

creating a processing window of a configurable size;

processing each word in the text in the configurable window using a text normalization or standardization process;

accessing a set of keywords determining if one or more of the keywords or are in the text in the configurable window, wherein the set of keywords represent a speech pattern or characteristics of a human speaker; and

automatically perform an indicated action if a segment of speech is determined to be speech generated by a human or automatically performing a different indicated action if the segment of speech is determined to be speech generated by a machine.

14 . The system of claim 13 , wherein the second entity is a call center which connects the call to an IVR system associated with the call center, wherein the IVR system generates one or more prompts in the form of speech segments that are navigated through to be connected to a human call center representative, and wherein after navigation through one or more prompts, the call is placed into an on-hold state by the call center.

15 . The system of claim 14 , wherein the speech segments are navigated through using a trained model.

16 . The system of claim 14 , wherein the speech segments are processed by a service provided to the first entity, and if a segment of speech is determined to be speech generated by a human, then the indicated action is to generate an alert when the human call center representative is available.

17 . The system of claim 16 , wherein in response to the generated alert, executing a trained classifier to classify a speech segment generated by the human call center representative, and in response to the classification, generate a speech segment for presentation to the human call center representative.

18 . The system of claim 14 , wherein the first entity is an automated process, the second entity is a human, the speech segments are processed by a service provided to the human, and the indicated action is to alert the human if the speech segments are machine generated.

19 . The system of claim 14 , wherein the text normalization or standardization process is one of removing punctuation, removing stop words, or removing hesitation words.

20 . The system of claim 14 , wherein multiple speech segments in multiple calls are processed and used to determine a distribution of the duration of a set of calls or of a set of sections of a call.

Assignments (2)
SECURITY INTEREST Recorded Dec 31, 2025
From: OUTBOUND AI, INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY
Reel/Frame 073346/0933 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2024
From: REOUVENI, RONEN; PIRO, ROBERT; WIGGS, JONATHAN; QUINN, CHRISTINA; SOLIMAN, MOHAMMED
To: OUTBOUND AI INC.
Reel/Frame 067111/0778 →
Continuity (2)
Provisional Application 63451985 · Mar 14, 2023
Related Publication 20240312466A1 · Sep 19, 2024
References Cited (25)
US 11074407B2 · Kannan · 2021 [cited by examiner]
US 11128636B1 · Jorasch · 2021 [cited by examiner]
US 11158327B2 · Choi · 2021 [cited by examiner]
US 11410657B2 · Kim · 2022 [cited by examiner]
US 11449668B2 · Mann · 2022 [cited by examiner]
US 11672479B2 · Jorasch · 2023 [cited by examiner]
US 12061861B2 · Liu · 2024 [cited by examiner]
US 12348674B2 · Reouveni · 2025 [cited by examiner]
US 20200023856A1 · Kim · 2020 [cited by examiner]
US 20200035244A1 · Kim · 2020 [cited by examiner]
US 20200035249A1 · Choi · 2020 [cited by examiner]
US 20200302015A1 · Kannan · 2020 [cited by examiner]
US 20210272040A1 · Johnson · 2021 [cited by examiner]
US 20220006813A1 · Jorasch · 2022 [cited by examiner]
US 20230315247A1 · Pastrana · 2023 [cited by examiner]
US 20230351098A1 · Liu · 2023 [cited by examiner]
US 20240312466A1 · Reouveni · 2024 [cited by examiner]
AU 2023221266A1 · 2024 [cited by examiner]
CN 107122154A · 2017 [cited by examiner]
CN 112262431A · 2021 [cited by examiner]
JP 2018132754A · 2018 [cited by examiner]
JP 6448723B2 · 2019 [cited by examiner]
JP 2025507385A · 2025 [cited by examiner]
KR 20240132373A · 2024 [cited by examiner]
WO WO2023158745A1 · 2023 [cited by examiner]