IP Library Granted Patent US 12,230,256
Granted Patent B2
US 12,230,256 · App. 17/886,859 · Granted Feb 18, 2025

Input-aware and input-unaware iterative speech recognition

Inventors: Michael Levy (Alpharetta, GA); Jay Miller (Alpharetta, GA)
Assignee: Verint Americas Inc.
G10L15/1815G10L15/05G10L15/1822G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,230,256
App. No.
17/886,859
Granted
Feb 18, 2025
Kind
B2
Abstract

An interactive voice response (IVR) system including iterative speech recognition with semantic interpretation is deployed to process an audio input in a manner that optimizes and conserves computing resources and facilitates low-latency discovery of start-of-speech events that can be used to support external processes such as barge-in operations. The IVR system can repeatedly receive an audio input at a speech processing component and apply an input-aware recognition process to the audio input. In response to generating a start-of-speech event, the IVR system can apply an input-unaware recognition process to the remaining audio input and determine a semantic meaning in relation to the relevant portion of the audio input.

Claims (45)

1. A method for processing an audio input using an iterative speech recognition system, the method comprising:

receiving the audio input at a speech processing component;

applying an input-aware recognition process to the audio input to generate a start-of-speech event;

responsive to generating the start-of-speech event, applying an input-unaware recognition process to a remaining audio input to determine silence, unexpected speech, or expected speech; and

initiating a response in accordance with a result of the input-unaware recognition process.

2. The method of claim 1 , wherein the input-aware recognition process comprises iteratively analyzing the audio input using an overlapping, high-frequency constant size audio window.

3. The method of claim 1 , wherein the input-unaware recognition process comprises iteratively analyzing the audio input using an increasing audio window size.

4. The method of claim 1 , wherein applying the input-unaware recognition process further comprises:

generating a semantic interpretation result describing a semantic meaning of the audio input.

5. The method of claim 4 , further comprising:

using the semantic interpretation result to control speech complete detection.

6. The method of claim 1 , wherein determining a semantic meaning of the audio input further comprises:

generating a confidence value associated with the semantic meaning.

7. The method of claim 6 , further comprising:

determining whether the confidence value meets or exceeds a confidence threshold.

8. A computer system for processing an audio input, the system comprising:

a processor; and

a memory operably coupled to the processor, the memory having computer-executable instructions stored thereon that, when executed by the processor, causes the system to:

receive the audio input at a speech processing component;

apply an input-aware recognition process to the audio input to generate a start-of-speech event;

responsive to generating the start-of-speech event, apply an input-unaware recognition process to a remaining audio input to determine silence, unexpected speech, or expected speech; and

initiate a response in accordance with a result of the input-unaware recognition process.

9. The computer system of claim 8 , wherein the input-aware recognition process comprises iteratively analyzing the audio input using an overlapping, high-frequency constant size audio window.

10. The computer system of claim 8 , wherein the input-unaware recognition process comprises iteratively analyzing the audio input using an increasing audio window size.

11. The computer system of claim 8 , wherein applying the input-unaware recognition process further comprises:

generating a semantic interpretation result describing a semantic meaning of the audio input.

12. The computer system of claim 11 , wherein the computer-executable instructions further include instructions to cause the processor to:

use the semantic interpretation result to control speech complete detection.

13. The computer system of claim 11 , wherein the computer-executable instructions further include instructions to cause the processor to:

generate a confidence value associated with the semantic meaning.

14. The computer system of claim 13 , wherein the computer-executable instructions further include instructions to cause the processor to:

determine whether the confidence value meets or exceeds a confidence threshold.

15. A non-transitory computer readable medium comprising instructions that, when executed by a processor of a processing system, cause the processing system to perform a method for processing an audio input, comprising instructions to:

receive the audio input at a speech processing component;

apply an input-aware recognition process to the audio input to generate a start-of-speech event;

responsive to generating the start-of-speech event, apply an input-unaware recognition process to a remaining audio input to determine silence, unexpected speech, or expected speech; and

initiate a response in accordance with a result of the input-unaware recognition process.

16. The non-transitory computer readable medium of claim 15 , wherein the input-aware recognition process comprises iteratively analyzing the audio input using an overlapping, high-frequency constant size audio window.

17. The non-transitory computer readable medium of claim 15 , wherein the input-unaware recognition process comprises iteratively analyzing the audio input using an increasing audio window size.

18. The non-transitory computer readable medium of claim 15 , wherein applying the input-unaware recognition process further comprises:

generating a semantic interpretation result describing a semantic meaning of the audio input.

19. The non-transitory computer readable medium of claim 18 , wherein the instructions further include instructions to cause the processor to:

generate a confidence value associated with the semantic meaning.

20. The non-transitory computer readable medium of claim 19 , wherein the instructions further include instructions to cause the processor to:

determine whether the confidence value meets or exceeds a confidence threshold.

Assignments (2)
SECURITY INTEREST Recorded Dec 23, 2025
From: VERINT AMERICAS INC.
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 074034/0292 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2022
From: LEVY, MICHAEL; MILLER, JAY
To: VERINT AMERICAS INC.
Reel/Frame 061492/0198 →
Continuity (1)
Related Publication 20240054995A1 · Feb 15, 2024
References Cited (8)
US 9953638B2 · Willett · 2018 [cited by examiner]
US 20160027440A1 · Gelfenbeyn · 2016 [cited by examiner]
US 20160260434A1 · Gelfenbeyn · 2016 [cited by examiner]
US 20170316780A1 · Lovitt · 2017 [cited by examiner]
US 20190325870A1 · Mitic · 2019 [cited by examiner]
US 20230176550A1 · Cella · 2023 [cited by examiner]
US 20230176557A1 · Cella · 2023 [cited by examiner]
US 20240055018A1 · Levy · 2024 [cited by examiner]