IP Library › Granted Patent US 12,322,410
Granted Patent B2
US 12,322,410 · App. 17/806,565 · Granted Jun 3, 2025

System and method for handling unsplit segments in transcription of air traffic communication (ATC)

Inventor: Jitender Kumar Agarwal (Bangalore, IN)
Assignee: HONEYWELL INTERNATIONAL, INC.
G10L21/10G10L15/26G10L25/78G10L25/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,322,410
App. No.
17/806,565
Granted
Jun 3, 2025
Kind
B2
Abstract

Systems and methods are provided for a transcription system with voice activity detection (VAD). The system includes a VAD module to receive incoming audio and generate an audio segment; and a speech decoder with a split predictor to perform, in a first pass, a decode operation to transcribe text from an audio segment into a message; wherein in the first pass, if the message is determined not to contain a split point based on a content-based analysis performed by the split predictor, the speech decoder forwards the message for display and if the message is determined based on the content-based analysis to contain the split point, the speech decoder performs in a second pass, a re-decode operation to transcribe text from the audio segment based on the split point wherein the split point is configured within an audio domain of the audio segment by the split predictor and forward the message for display.

Claims (31)

1. A transcription system with voice activity detection (VAD) comprising:

a VAD module that is configured with an input channel to receive incoming audio and to generate at least one audio segment based on the incoming audio; and

a speech decoder in operable communication with a split predictor and the VAD module, the speech decoder is configured to perform, in a first pass, a decode operation to transcribe text from the at least one audio segment into a message;

wherein in the first pass, if the message is determined not to contain at least one split point based on a content-based analysis performed by the split predictor, the speech decoder forwards the message for display;

wherein in the first pass, if the message is determined, based on the content-based analysis performed by the split predictor, to contain the at least one split point, the speech decoder performs, in a second pass, a re-decode operation to transcribe text from the at least one audio segment based on the at least one split point and forwards the message for display configured in multiple text segments based on at least one audio split point, wherein the at least one split point is configured within an audio domain of the at least one audio segment by the split predictor;

wherein the system further comprises a natural language processor (NLP) in operable communication with the speech decoder configured to communicate with the split predictor to provide content information about the message for the content-based analysis;

wherein the content-based analysis performed by the split predictor comprises an intelligent application that determines at least repetitive usage of call signs and critical information in content of the message, or multiple speaker dialogues for defining the at least one split point in the message; and

wherein the message containing the at least one split point is an unsplit message.

2. The transcription system of claim 1 , wherein the unsplit message is caused by at least a pause in the incoming audio that is not detected by the VAD module.

3. The transcription system of claim 2 , wherein the pause is less than a threshold value configured in a set of ranges of approximately 10, 20, and 30 milliseconds or less.

4. A method of implementing a transcription system, the method comprising:

receiving, by a voice activity detection (VAD) module, incoming audio to generate at least one audio segment;

in a first pass, decoding, by a speech decoder coupled to the VAD module, text from the at least one audio segment to generate a message; and

determining, by a split predictor coupled to the speech decoder, and based on a content-based analysis, whether the message contains a split point;

wherein if the message does not contain the split point then enabling display of the message;

wherein if the message does contain the split point, then enabling re-decode by the speech decoder of the at least one audio segment based on the split point configured by the split point predictor from the content-based analysis, and enabling the display of the message configured in multiple text segments based on at least one audio split point, wherein the split point in the message is defined in an audio domain of the at least one audio segment;

wherein the method further comprises configuring a natural language processor (NLP) in operable communication with the speech decoder configured to communicate with the split predictor to provide content information about the message for the content-based analysis;

wherein the content-based analysis performed by the split predictor comprises an intelligent application that determines at least repetitive usage of call signs and critical information in content of the message, or multiple speaker dialogues for defining the at least one split point in the message; and

wherein the message containing the at least one split point is an unsplit message.

5. The method of claim 4 , wherein the unsplit message is caused by at least a pause in the incoming audio that is not detected by the VAD module.

6. The method of claim 5 , wherein the pause is less than a threshold value configured in a set of ranges of approximately 10, 20, and 30 milliseconds or less.

7. At least one non-transient computer-readable medium having instructions stored thereon that are configurable to cause at least one processor to perform a method to segment a transcribed textual message by a transcription system, the method comprising:

receiving, by a voice activity detection (VAD) module incoming audio to generate at least one audio segment;

in a first pass, decoding, by a speech decoder coupled to the VAD module, text from the at least one audio segment to generate a message; and

determining, by a split predictor coupled to the speech decoder, and based on a content-based analysis whether the message contains a split point;

wherein if the message does not contain the split point then enabling display of the message;

wherein if the message does contain the split point then enabling re-decode by the speech decoder of the at least one audio segment based on the split point configured by the split point predictor from the content-based analysis, and enabling the display of the message configured in multiple text segments based on at least one audio split point wherein the split point in the message is defined in an audio domain of the at least one audio segment by the split predictor;

wherein the method further comprises configuring a natural language processor (NLP) in operable communication with the speech decoder configured to communicate with the split predictor to provide content information about the message for the content-based analysis;

wherein the content-based analysis performed by the split predictor comprises an intelligent application that determines at least repetitive usage of call signs and critical information in content of the message, or multiple speaker dialogues for defining the at least one split point in the message; and

wherein the message containing the at least one split point is an unsplit message and wherein the unsplit message is caused by at least a pause in the incoming audio that is not detected by the VAD module.

8. The method of claim 7 , wherein the pause is less than a threshold value configured in a set of ranges of approximately 10, 20, and 30 milliseconds or less.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2022
From: AGARWAL, JITENDER KUMAR
To: HONEYWELL INTERNATIONAL INC.
Reel/Frame 060181/0112 →
Priority Claims (1)
IN 202211025236 · Apr 29, 2022 · national
Continuity (1)
Related Publication 20230352042A1 · Nov 2, 2023
References Cited (51)
US 5333275A · Wheatley et al. · 1994 [cited by applicant]
US 6230131B1 · Kuhn · 2001 [cited by examiner]
US 6463413B1 · Applebaum et al. · 2002 [cited by applicant]
US 8249870B2 · Roy et al. · 2012 [cited by applicant]
US 8306675B2 · Prus et al. · 2012 [cited by applicant]
US 8577913B1 · Hansson · 2013 [cited by examiner]
US 8626498B2 · Lee · 2014 [cited by applicant]
US 9355094B2 · Cuthbert et al. · 2016 [cited by applicant]
US 9368108B2 · Liu et al. · 2016 [cited by applicant]
US 9786283B2 · Baker · 2017 [cited by applicant]
US 10152968B1 · Agrusa et al. · 2018 [cited by applicant]
US 10403274B2 · Girod et al. · 2019 [cited by applicant]
US 10515625B1 · Metallinou · 2019 [cited by examiner]
US 10573304B2 · Gemmeke et al. · 2020 [cited by applicant]
US 10629186B1 · Slifka · 2020 [cited by applicant]
US 10878807B2 · Tomar et al. · 2020 [cited by applicant]
US 11423887B2 · Pust · 2022 [cited by examiner]
US 12087276B1 · Nokob · 2024 [cited by examiner]
US 20050165602A1 · Cote et al. · 2005 [cited by applicant]
US 20130144414A1 · Kajarekar · 2013 [cited by examiner]
US 20130197917A1 · Dong et al. · 2013 [cited by applicant]
US 20150100311A1 · Kar · 2015 [cited by examiner]
US 20150217870A1 · McCullough et al. · 2015 [cited by applicant]
US 20160093302A1 · Bilek · 2016 [cited by examiner]
US 20160379640A1 · Joshi et al. · 2016 [cited by applicant]
US 20180047387A1 · Nir · 2018 [cited by applicant]
US 20180129635A1 · Saptharishi et al. · 2018 [cited by applicant]
US 20190147858A1 · Letsu-Dake et al. · 2019 [cited by applicant]
US 20190310981A1 · Sevenster · 2019 [cited by examiner]
US 20200027457A1 · Gelinske et al. · 2020 [cited by applicant]
US 20200075044A1 · Jankowski, Jr. et al. · 2020 [cited by applicant]
US 20200104362A1 · Yang · 2020 [cited by examiner]
US 20200135204A1 · Robichaud · 2020 [cited by examiner]
US 20200171671A1 · Huang et al. · 2020 [cited by applicant]
US 20200183983A1 · Abe · 2020 [cited by examiner]
US 20210020168A1 · Pabla et al. · 2021 [cited by applicant]
US 20210074277A1 · Duncan · 2021 [cited by applicant]
US 20210225371A1 · Takacs et al. · 2021 [cited by applicant]
US 20210233411A1 · Saptharishi et al. · 2021 [cited by applicant]
US 20210342634A1 · Chen · 2021 [cited by examiner]
US 20220115019A1 · Bradley et al. · 2022 [cited by applicant]
US 20220115020A1 · Bradley et al. · 2022 [cited by applicant]
US 20220238118A1 · Mazzoccoli · 2022 [cited by applicant]
CN 111785257A · 2020 [cited by applicant]
CN 112954122A · 2021 [cited by applicant]
EP 2669889A2 · 2013 [cited by applicant]
EP 4095853A1 · 2022 [cited by applicant]
WO 2009104332A1 · 2009 [cited by applicant]
Furui Sadaoki; “Recent advances in robust speech recognition” Assistant-based speech recognition from ATM applications. Apr. 17, 1997 pp. 11-20 Section 4.1 XP093050804 Retrieved from the Internet: URL:https//www.isca-sp… [cited by applicant]
Park, Tae Jin, et al.: “A review of speaker diarization: Recent advances with deep learning”, arXiv article, Jan. 24, 2021 (Jan. 24, 2021), XP055935769, DOI: 10.1016/j.csl.2021.101317 Retrieved from the Internet:URL:htt… [cited by applicant]
Nikolaos Flemotomos, et al.: “Linguistically Aided Speaker Diarization Using Speaker Role Information”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 18, 2019 (Nov. 18… [cited by applicant]