IP Library Granted Patent US 12,190,861
Granted Patent B2
US 12,190,861 · App. 17/659,596 · Granted Jan 7, 2025

Adaptive speech recognition methods and systems

Inventors: Jitender Kumar Agarwal (Bangalore, IN); Chaya Garg (Plymouth, MN); Vasantha Paulraj (Madurai, IN); Ramakrishnan Raman (Bangalore, IN); Mahesh Kumar Sampath (Madurai, IN); Mohan M Thippeswamy (Bangalore, IN)
Assignee: HONEYWELL INTERNATIONAL INC.
G10L15/063B64D11/0015G10L15/01G10L15/19G10L15/22G10L15/30G10L2015/0635G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,861
App. No.
17/659,596
Granted
Jan 7, 2025
Kind
B2
Abstract

Methods and systems are provided for assisting operation of a vehicle using speech recognition. One method involves analyzing a transcription of an audio communication with respect to the vehicle to characterize a nonstandard pattern within the transcription of the audio communication, obtaining a ground truth for the transcription of the audio communication, determining one or more performance metrics associated with the nonstandard pattern within the transcription based on a relationship between the transcription of the audio communication and the ground truth for the transcription, updating a speech recognition vocabulary for the vehicle to include the nonstandard pattern based at least in part on the one or more performance metrics and determining an updated speech recognition model for the vehicle using the updated speech recognition vocabulary and the audio communication.

Claims (55)

1. A method of assisting operation of a vehicle, the method comprising:

analyzing a transcription of an audio communication with respect to the vehicle to characterize a nonstandard pattern within the transcription of the audio communication;

obtaining a ground truth for the transcription of the audio communication;

determining one or more performance metrics associated with the nonstandard pattern within the transcription based on a relationship between the transcription of the audio communication and the ground truth for the transcription;

updating a speech recognition vocabulary for the vehicle to include the nonstandard pattern based at least in part on the one or more performance metrics, resulting in an updated speech recognition vocabulary; and

determining an updated speech recognition model for the vehicle using the updated speech recognition vocabulary and the audio communication.

2. The method of claim 1 , further comprising pushing the updated speech recognition model to the vehicle over a communications network.

3. The method of claim 2 , further comprising analyzing, at the vehicle, a second transcription of a subsequent audio communication with respect to the vehicle to detect the nonstandard pattern within the second transcription of the subsequent audio communication using at least one of the updated speech recognition vocabulary and the updated speech recognition model.

4. The method of claim 1 , further comprising analyzing a second transcription of a subsequent audio communication with respect to the vehicle to detect the nonstandard pattern within the second transcription of the subsequent audio communication using at least one of the updated speech recognition vocabulary and the updated speech recognition model.

5. The method of claim 1 , wherein analyzing the transcription of the audio communication with respect to the vehicle to characterize the nonstandard pattern within the transcription of the audio communication comprises:

determining at least one of a phraseology pattern subject category and a phraseology pattern structure type associated with the transcription of the audio communication based at least in part on content of the transcription of the audio communication;

determining whether the content of the transcription of the audio communication corresponds to one of a plurality of standard phraseology patterns based at least in part on the at least one of the phraseology pattern subject category and the phraseology pattern structure type; and

determining the transcription of the audio communication comprises the nonstandard pattern when the transcription of the audio communication does not correspond to any of the plurality of standard phraseology patterns.

6. The method of claim 5 , further comprising assigning a pattern identifier associated with an existing nonstandard pattern to the transcription of the audio communication when the content of the transcription of the audio communication corresponds to the existing nonstandard pattern based at least in part on the at least one of the phraseology pattern subject category and the phraseology pattern structure type.

7. The method of claim 1 , wherein:

analyzing the transcription of the audio communication with respect to the vehicle to characterize the nonstandard pattern comprises identifying a phraseology pattern portion associated with the nonstandard pattern within the transcription of the audio communication; and

determining the one or more performance metrics associated with the nonstandard pattern within the transcription comprises determining a phraseology pattern performance metric based on a relationship between the phraseology pattern portion of the transcription of the audio communication and a second phraseology pattern portion of the ground truth.

8. The method of claim 1 , wherein:

determining the one or more performance metrics comprises determining a pattern-based performance metric associated with the nonstandard pattern within the transcription based on a relationship between a phraseology pattern portion of the transcription of the audio communication and the phraseology pattern portion of the ground truth for the transcription; and

updating the speech recognition vocabulary comprises updating the speech recognition vocabulary based on the pattern-based performance metric.

9. A non-transitory computer-readable medium having computer-executable instructions stored thereon that, when executed by a processing system, cause the processing system to:

analyze a transcription of an audio communication with respect to a vehicle to characterize a nonstandard pattern within the transcription of the audio communication;

obtain a ground truth for the transcription of the audio communication;

determine one or more performance metrics associated with the nonstandard pattern within the transcription based on a relationship between the transcription of the audio communication and the ground truth for the transcription;

update a speech recognition vocabulary to include the nonstandard pattern based at least in part on the one or more performance metrics, resulting in an updated speech recognition vocabulary; and

determine an updated speech recognition model for the vehicle using the updated speech recognition vocabulary and the audio communication.

10. The non-transitory computer-readable medium of claim 9 , wherein the computer-executable instructions cause the processing system to push the updated speech recognition model to the vehicle over a communications network.

11. The non-transitory computer-readable medium of claim 9 , wherein the computer-executable instructions cause the processing system to analyze a second transcription of a subsequent audio communication with respect to the vehicle to detect the nonstandard pattern within the second transcription of the subsequent audio communication using at least one of the updated speech recognition vocabulary and the updated speech recognition model.

12. The non-transitory computer-readable medium of claim 9 , wherein the computer-executable instructions cause the processing system to analyze the transcription of the audio communication with respect to the vehicle to characterize the nonstandard pattern within the transcription of the audio communication by:

determining at least one of a phraseology pattern subject category and a phraseology pattern structure type associated with the transcription of the audio communication based at least in part on content of the transcription of the audio communication;

determining whether the content of the transcription of the audio communication corresponds to one of a plurality of standard phraseology patterns based at least in part on the at least one of the phraseology pattern subject category and the phraseology pattern structure type; and

determining the transcription of the audio communication comprises the nonstandard pattern when the transcription of the audio communication does not correspond to any of the plurality of standard phraseology patterns.

13. The non-transitory computer-readable medium of claim 12 , wherein the computer-executable instructions cause the processing system to assign a pattern identifier associated with an existing nonstandard pattern to the transcription of the audio communication when the content of the transcription of the audio communication corresponds to the existing nonstandard pattern based at least in part on the at least one of the phraseology pattern subject category and the phraseology pattern structure type.

14. The non-transitory computer-readable medium of claim 9 , wherein:

analyzing the transcription of the audio communication with respect to the vehicle to characterize the nonstandard pattern comprises identifying a phraseology pattern portion associated with the nonstandard pattern within the transcription of the audio communication; and

determining the one or more performance metrics associated with the nonstandard pattern within the transcription comprises determining a phraseology pattern performance metric based on a relationship between the phraseology pattern portion of the transcription of the audio communication and a second phraseology pattern portion of the ground truth.

15. The non-transitory computer-readable medium of claim 9 , wherein the computer-executable instructions cause the processing system to:

determine a pattern-based performance metric associated with the nonstandard pattern within the transcription based on a relationship between a phraseology pattern portion of the transcription of the audio communication and the phraseology pattern portion of the ground truth for the transcription; and

update the speech recognition vocabulary to include the nonstandard pattern based on the pattern-based performance metric.

16. A computing device comprising:

at least one computer-readable storage medium to store computer-executable instructions; and

at least one processor, coupled to the at least one computer-readable storage medium, to execute the computer-executable instructions to:

analyze a transcription of an audio communication with respect to a vehicle to characterize a nonstandard pattern within the transcription of the audio communication;

obtain a ground truth for the transcription of the audio communication;

determine one or more performance metrics associated with the nonstandard pattern within the transcription based on a relationship between the transcription of the audio communication and the ground truth for the transcription;

update a speech recognition vocabulary to include the nonstandard pattern based at least in part on the one or more performance metrics, resulting in an updated speech recognition vocabulary; and

determine an updated speech recognition model for the vehicle using the updated speech recognition vocabulary and the audio communication.

17. The computing device of claim 16 , wherein the computer-executable instructions cause the at least one processor to push the updated speech recognition model to the vehicle over a communications network.

18. The computing device of claim 16 , wherein the computer-executable instructions cause the at least one processor to analyze a second transcription of a subsequent audio communication with respect to the vehicle to detect the nonstandard pattern within the second transcription of the subsequent audio communication using at least one of the updated speech recognition vocabulary and the updated speech recognition model.

19. The computing device of claim 16 , wherein the computer-executable instructions cause the at least one processor to:

identify a phraseology pattern portion associated with the nonstandard pattern within the transcription of the audio communication; and

determine a phraseology pattern performance metric based on a relationship between the phraseology pattern portion of the transcription of the audio communication and a second phraseology pattern portion of the ground truth.

20. The computing device of claim 16 , wherein the computer-executable instructions cause the at least one processor to:

determine a pattern-based performance metric associated with the nonstandard pattern within the transcription based on a relationship between a phraseology pattern portion of the transcription of the audio communication and the phraseology pattern portion of the ground truth for the transcription; and

update the speech recognition vocabulary based on the pattern-based performance metric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2022
From: AGARWAL, JITENDER KUMAR; GARG, CHAYA; PAULRAJ, VASANTHA; RAMAN, RAMAKRISHNAN; SAMPATH, MAHESH KUMAR; THIPPESWAMY, MOHAN M
To: HONEYWELL INTERNATIONAL INC.
Reel/Frame 059626/0677 →
Priority Claims (1)
IN 202111018599 · Apr 22, 2021 · national
Continuity (1)
Related Publication 20220343897A1 · Oct 27, 2022
References Cited (84)
US 6992626B2 · Smith · 2006 [cited by applicant]
US 7184863B2 · Weineck · 2007 [cited by applicant]
US 7415326B2 · Komer et al. · 2008 [cited by applicant]
US 7668719B2 · Nakagawa et al. · 2010 [cited by applicant]
US 7733903B2 · Bhogal et al. · 2010 [cited by applicant]
US 7809405B1 · Rand et al. · 2010 [cited by applicant]
US 7881832B2 · Komer et al. · 2011 [cited by applicant]
US 8149141B2 · Coulmeau et al. · 2012 [cited by applicant]
US 8180503B2 · Estabrook et al. · 2012 [cited by applicant]
US 8280741B2 · Colin et al. · 2012 [cited by applicant]
US 8340839B2 · Yogesha et al. · 2012 [cited by applicant]
US 8681040B1 · Rathinam et al. · 2014 [cited by applicant]
US 8704701B2 · Pschierer et al. · 2014 [cited by applicant]
US 8768698B2 · Mengibar et al. · 2014 [cited by applicant]
US 8793139B1 · Serban et al. · 2014 [cited by applicant]
US 8812316B1 · Chen · 2014 [cited by applicant]
US 8909392B1 · Carrico · 2014 [cited by applicant]
US 8957790B2 · Cornell et al. · 2015 [cited by applicant]
US 9047870B2 · Ballinger et al. · 2015 [cited by applicant]
US 9190073B2 · Dong et al. · 2015 [cited by applicant]
US 9443433B1 · Conway et al. · 2016 [cited by applicant]
US 9487167B2 · Graumann et al. · 2016 [cited by applicant]
US 9620119B2 · Bilek et al. · 2017 [cited by applicant]
US 9642184B2 · Plocher et al. · 2017 [cited by applicant]
US 9665645B2 · Hawley · 2017 [cited by applicant]
US 9666178B2 · Loubiere et al. · 2017 [cited by applicant]
US 9704405B2 · Kashi et al. · 2017 [cited by applicant]
US 9830829B1 · Doyen et al. · 2017 [cited by applicant]
US 9881608B2 · Lebeau et al. · 2018 [cited by applicant]
US 10056085B2 · Klose et al. · 2018 [cited by applicant]
US 10204430B2 · Gowda · 2019 [cited by applicant]
US 10490085B2 · Cotdeloup et al. · 2019 [cited by applicant]
US 10535351B2 · Gaston et al. · 2020 [cited by applicant]
US 10818192B2 · Chen et al. · 2020 [cited by applicant]
US 20040124998A1 · Dame · 2004 [cited by applicant]
US 20040263381A1 · Mitchell et al. · 2004 [cited by applicant]
US 20050144187A1 · Che et al. · 2005 [cited by applicant]
US 20050203700A1 · Merritt · 2005 [cited by applicant]
US 20060229873A1 · Eide et al. · 2006 [cited by applicant]
US 20070189328A1 · Judd · 2007 [cited by applicant]
US 20070288128A1 · Komer et al. · 2007 [cited by applicant]
US 20080201148A1 · Desrochers · 2008 [cited by applicant]
US 20110028147A1 · Calderhead, Jr. et al. · 2011 [cited by applicant]
US 20110125503A1 · Dong et al. · 2011 [cited by applicant]
US 20110137653A1 · Ljolje et al. · 2011 [cited by applicant]
US 20110202351A1 · Plocher et al. · 2011 [cited by applicant]
US 20110231036A1 · Yogesha et al. · 2011 [cited by applicant]
US 20120078448A1 · Dorneich et al. · 2012 [cited by applicant]
US 20130093612A1 · Pschierer et al. · 2013 [cited by applicant]
US 20130103297A1 · Bilek et al. · 2013 [cited by applicant]
US 20150081138A1 · Lacko et al. · 2015 [cited by applicant]
US 20150162001A1 · Kar et al. · 2015 [cited by applicant]
US 20150212671A1 · Judy et al. · 2015 [cited by applicant]
US 20150212701A1 · Rodney et al. · 2015 [cited by applicant]
US 20160063999A1 · Gaston · 2016 [cited by examiner]
US 20160125744A1 · Shamasundar et al. · 2016 [cited by applicant]
US 20160155435A1 · Mohideen · 2016 [cited by applicant]
US 20160379640A1 · Joshi et al. · 2016 [cited by applicant]
US 20170039858A1 · Wang et al. · 2017 [cited by applicant]
US 20180061243A1 · Shloosh · 2018 [cited by applicant]
US 20190147858A1 · Letsu-Dake et al. · 2019 [cited by applicant]
US 20190244528A1 · Srinivasan et al. · 2019 [cited by applicant]
US 20200322040A1 · Middlestead et al. · 2020 [cited by applicant]
US 20200372916A1 · Delpech · 2020 [cited by applicant]
US 20210020168A1 · Dame · 2021 [cited by examiner]
US 20210233411A1 · Saptharishi · 2021 [cited by examiner]
US 20210295840A1 · John et al. · 2021 [cited by applicant]
CN 110335609A · 2019 [cited by applicant]
DE 102009025530A1 · 2010 [cited by applicant]
EP 0613110A1 · 1994 [cited by applicant]
EP 0618565A2 · 1994 [cited by applicant]
EP 1318492A2 · 2003 [cited by applicant]
EP 2026328A1 · 2009 [cited by applicant]
EP 3664065A1 · 2020 [cited by applicant]
EP 3889947A1 · 2021 [cited by applicant]
FR 3032574A1 · 2016 [cited by applicant]
FR 3032575A1 · 2016 [cited by applicant]
FR 3009759B1 · 2017 [cited by applicant]
IN 111785257A · 2020 [cited by applicant]
WO 2016076939A1 · 2016 [cited by applicant]
Cardosi, Kim and Tracy Lennertz “Loss of Controller-Pilot Voice Communications in Domestic En Route Airspace.”DOT-VNTSC-FAA-17-04, dated Feb. 2016. [cited by applicant]
“Air-Ground Voice Communications,” SKYbrary, downloaded from Internet Mar. 28, 2018. [cited by applicant]
“Loss of Communication,” SKYbrary, downloaded from Internet Mar. 28, 2018. [cited by applicant]
Saptharishi, et al. Contextual Speech Recognition Methods and Systems; Filed with the USPTO on Jun. 22, 2021 and assigned U.S. Appl. No. 17/354,580. [cited by applicant]