IP Library › Granted Patent US 12,351,119
Granted Patent B2
US 12,351,119 · App. 18/062,163 · Granted Jul 8, 2025

Systems and methods for performing commands in a vehicle using speech and image recognition

Inventors: Sumit Bhattacharya (Maharashtra, IN); Jason Conrad Roche (Santa Clara, CA); Niranjan Avadhanam (Saratoga, CA)
Assignee: NVIDIA Corporation
B60R16/0373B60R25/01B60R25/257B60R25/305G05B13/027G06F21/32G06N3/08G06V10/25G06V20/59G10L17/00G10L17/06G10L17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,351,119
App. No.
18/062,163
Granted
Jul 8, 2025
Kind
B2
Abstract

Systems and methods are disclosed herein for implementation of a vehicle command operation system that may use multi-modal technology to authenticate an occupant of the vehicle to authorize a command and receive natural language commands for vehicular operations. The system may utilize sensors to receive data indicative of a voice command from an occupant of the vehicle. The system may receive second sensor data to aid in the determination of the corresponding vehicular operation in response to the received command. The system may retrieve authentication data for the occupants of the vehicle. The system authenticates the occupant to authorize a vehicular operation command using a neural network based on at least one of the first sensor data, the second sensor data, and the authentication data. Responsive to the authentication, the system may authorize the operation to be performed in the vehicle based on the vehicular operation command.

Claims (84)

1. A method comprising:

receiving audio data generated using one or more microphones of a machine, the audio data representative of user speech from a first occupant of the machine;

processing the audio data to determine that the user speech from the first occupant indicates an identifier of a second occupant of the machine and an operation associated with a component of the machine, the second occupant being different than the first occupant;

receiving image data generated using one or more image sensors of the machine, the image data representative of an image depicting at least a portion of an interior of the machine;

based at least on the user speech from the first occupant indicating the identifier of the second occupant, determining, using at least the image data, that the image depicts the second occupant within a region of the machine that is associated with the component; and

causing the operation to be performed with respect to the component.

2. The method of claim 1 , further comprising:

storing data that associates the component with the region of the machine,

wherein the causing the operation to be performed with respect to the component is based at least on the image data depicting the second occupant within the region and the component being associated with the region.

3. The method of claim 1 , wherein the first occupant is associated with a second region within the machine that is different than the region associated with the second occupant within the machine.

4. The method of claim 1 , wherein the identifier of the second occupant comprises one or more of:

a synonym of the second occupant,

a colloquial phase of the second occupant,

a shorthand name of the second occupant,

a name of the second occupant, or

a related descriptor of the second occupant in a different language than a voice command interface.

5. The method of claim 1 , further comprising storing, based at least on a position of the second occupant with the machine, data that associates the second occupant with at least one of the component or the region of the machine that is associated with the component.

6. The method of claim 1 , further comprising:

storing first data representative of an image depicting the second occupant;

storing second data that associates the first data with the identifier of the second occupant; and

identifying, based at least on the first data and the second data, that a portion of the image data represents the second occupant,

wherein the determining that the image depicts the second occupant within the region of the machine that is associated with the component is based at least on the identifying the portion of the image data represents the second occupant.

7. The method of claim 1 , further comprising:

receiving second audio data generated using the one or more microphones of the machine, the second audio data representative of second user speech indicating at least a second identifier of a third occupant and a second operation associated with a second component;

determining, based at least on the second audio data and second image data, a position of the third occupant within the machine;

determining, based at least on the position of the third occupant within the machine, the second component is associated with the second operation; and

causing the second operation to be performed with respect to the second component.

8. A system comprising:

one or more processing units to:

receive audio data generated using one or more microphones of a machine, the audio data representative of user speech from a first occupant of the machine;

process the audio data to determine that the user speech indicates an identifier of a second occupant of the machine and an operation associated with a component of the machine, the second occupant being different than the first occupant;

receive image data generated using one or more image sensors of the machine, the image data representative of an image depicting at least a portion of an interior of the machine;

based at least on the user speech from the first occupant indicating the identifier of the second occupant, determine, using the image data, that the image represents the second occupant within a region of the machine that is associated with the component; and

cause the operation to be performed using the component.

9. The system of claim 8 , wherein the first occupant is within a second region of the machine that is different than the region of the machine.

10. The system of claim 8 , wherein the identifier of the second occupant comprises one or more of:

a synonym of the second occupant,

a colloquial phase of the second occupant,

a shorthand name of the second occupant,

a name of the second occupant, or

a related descriptor of the first occupant in a different language than a voice command interface.

11. The system of claim 8 , wherein the one or more processing units are further to:

store data that associates the component with the region of the machine,

wherein the determination that the image represents the second occupant located within the region of the machine that is associated with the component is further based at least on the data that associates the component with the region.

12. The system of claim 8 , wherein the one or more processing units are further to:

store data that associates the second occupant with at least one of a position within the machine or the component.

13. The system of claim 8 , wherein the one or more processing units are further determine, based at least on the second occupant being located within the region of the machine that is associated with the component and the user speech identifying the operation, to use the component to perform the operation.

14. The system of claim 8 , wherein the one or more processing units are further to:

store second image data representative of a second image depicting the second occupant;

wherein the determination that the image represents the second occupant located within the region of the machine that is associated with the component is further based at least on the second image data.

15. The system of claim 8 , wherein the one or more processing units are further to:

determine a position of a third occupant within the machine;

receive second audio data generated using the one or more microphones of the machine, the second audio data representative of second user speech from the first occupant, the second user speech indicating at least a second operation associated with a second component; and

based at least on the position of the third occupant, cause the second operation to be performed with respect to the second component.

16. The system of claim 8 , wherein a position of the first occupant is determined based at least on the first occupant entering the machine.

17. The system of claim 8 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for generating or presenting at least one of virtual reality content or augmented reality content;

a system for performing deep learning operations;

a system implemented using a robot;

a system for performing conversational AI operations;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

18. A processor comprising:

one or more processing units to:

determine, based at least on audio data representative of user speech from a first occupant of a machine, that the user speech indicates an operation and an identifier of a second occupant that is different than the first occupant;

based at least on the user speech from the first occupant indicating the identifier of the second occupant, determine, using data representative of an image, that the image represents the second occupant located within a region of the machine that is associated with a component;

determine, based at least on the second occupant being located within the region, to use the component to perform the operation; and

cause the operation to be performed with respect to the component.

19. The processor of claim 18 , wherein the processor is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for generating or presenting at least one of virtual reality content or augmented reality content;

a system for performing deep learning operations;

a system implemented using a robot;

a system for performing conversational AI operations;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

20. The processor of claim 18 , wherein the one or more processing units are further to:

determine that the image also represents the component,

wherein the determination to use the component to perform the operation is further based at least on the image further representing the component.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2022
From: BHATTACHARYA, SUMIT; ROCHE, JASON CONRAD; AVADHANAM, NIRANJAN
To: NVIDIA CORPORATION
Reel/Frame 062011/0868 →
Continuity (2)
Continuation 16867395 · May 5, 2020
Related Publication 20230095988A1 · Mar 30, 2023
References Cited (92)
US 5704008A · Duvall, Jr. · 1997 [cited by examiner]
US 7372370B2 · Stults · 2008 [cited by examiner]
US 7542826B2 · Hanzawa · 2009 [cited by examiner]
US 10157614B1 · Devaraj · 2018 [cited by examiner]
US 10178301B1 · Welbourne · 2019 [cited by examiner]
US 10580405B1 · Wang · 2020 [cited by examiner]
US 10674003B1 · Kang · 2020 [cited by examiner]
US 10885698B2 · Muthler et al. · 2021 [cited by applicant]
US 11023786B2 · Hiroki · 2021 [cited by examiner]
US 11145294B2 · Vescovi · 2021 [cited by examiner]
US 11180098B2 · Wheeler · 2021 [cited by examiner]
US 11355136B1 · Amman · 2022 [cited by examiner]
US 11590929B2 · Bhattacharya et al. · 2023 [cited by applicant]
US 20030144844A1 · Colmenarez · 2003 [cited by examiner]
US 20030210159A1 · Arunkumar · 2003 [cited by examiner]
US 20040170284A1 · Janse · 2004 [cited by examiner]
US 20050071159A1 · Boman · 2005 [cited by examiner]
US 20050275505A1 · Himmelstein · 2005 [cited by examiner]
US 20050285943A1 · Cutler · 2005 [cited by examiner]
US 20120054028A1 · Tengler · 2012 [cited by examiner]
US 20130145360A1 · Ricci · 2013 [cited by examiner]
US 20130183944A1 · Mozer · 2013 [cited by examiner]
US 20130200995A1 · Muramatsu · 2013 [cited by examiner]
US 20140195477A1 · Graumann · 2014 [cited by examiner]
US 20140214424A1 · Wang · 2014 [cited by examiner]
US 20140244069A1 · Yang · 2014 [cited by examiner]
US 20140375543A1 · Ng-Thow-Hing · 2014 [cited by examiner]
US 20150025888A1 · Sharp · 2015 [cited by examiner]
US 20150110287A1 · Holdren · 2015 [cited by examiner]
US 20150223032A1 · Nespolo · 2015 [cited by examiner]
US 20150348554A1 · Orr · 2015 [cited by examiner]
US 20160028730A1 · Natarajan · 2016 [cited by examiner]
US 20160098088A1 · Park · 2016 [cited by examiner]
US 20160288796A1 · Yuan · 2016 [cited by examiner]
US 20170061110A1 · Wright · 2017 [cited by examiner]
US 20170185362A1 · Cansino · 2017 [cited by examiner]
US 20170213541A1 · MacNeille · 2017 [cited by examiner]
US 20170236512A1 · Williams · 2017 [cited by examiner]
US 20170286407A1 · Chochowski · 2017 [cited by examiner]
US 20170345420A1 · Barnett, Jr. · 2017 [cited by examiner]
US 20180012433A1 · Ricci · 2018 [cited by examiner]
US 20180103022A1 · Tokunaga · 2018 [cited by examiner]
US 20180154899A1 · Tiwari · 2018 [cited by examiner]
US 20180176112A1 · Mclaughlin · 2018 [cited by examiner]
US 20180233139A1 · Finkelstein · 2018 [cited by examiner]
US 20180251122A1 · Golston · 2018 [cited by examiner]
US 20180261247A1 · Bellotti · 2018 [cited by examiner]
US 20180285062A1 · Ulaganathan · 2018 [cited by examiner]
US 20180314689A1 · Wang · 2018 [cited by examiner]
US 20190019516A1 · Van Hoecke · 2019 [cited by examiner]
US 20190054874A1 · Breed · 2019 [cited by examiner]
US 20190092169A1 · Thurimella · 2019 [cited by examiner]
US 20190121956A1 · Turgeman · 2019 [cited by examiner]
US 20190171409A1 · Boulanger · 2019 [cited by examiner]
US 20190219976A1 · Giorgi · 2019 [cited by examiner]
US 20190251973A1 · Kume · 2019 [cited by examiner]
US 20190258253A1 · Tremblay · 2019 [cited by examiner]
US 20190362725A1 · Himmelstein · 2019 [cited by examiner]
US 20190372986A1 · Singh · 2019 [cited by examiner]
US 20200013396A1 · Park · 2020 [cited by examiner]
US 20200017122A1 · Chatten · 2020 [cited by examiner]
US 20200023856A1 · Kim et al. · 2020 [cited by applicant]
US 20200047687A1 · Camhi · 2020 [cited by examiner]
US 20200135190A1 · Kaja · 2020 [cited by examiner]
US 20200152203A1 · Sugihara · 2020 [cited by examiner]
US 20200207358A1 · Katz · 2020 [cited by examiner]
US 20200216088A1 · Shin · 2020 [cited by examiner]
US 20200218442A1 · Shin · 2020 [cited by examiner]
US 20200249822A1 · Penilla · 2020 [cited by examiner]
US 20200269848A1 · Kang · 2020 [cited by examiner]
US 20200321000A1 · Kurihara · 2020 [cited by examiner]
US 20200356754A1 · Chen · 2020 [cited by examiner]
US 20200388266A1 · You · 2020 [cited by examiner]
US 20200388285A1 · Spiewla · 2020 [cited by examiner]
US 20210030308A1 · Grace · 2021 [cited by examiner]
US 20210035571A1 · Yun · 2021 [cited by examiner]
US 20210064029A1 · Tarakhovsky · 2021 [cited by examiner]
US 20210105619A1 · Kashani · 2021 [cited by examiner]
US 20210151057A1 · Sinha · 2021 [cited by examiner]
US 20210180977A1 · Wiesenberg · 2021 [cited by examiner]
US 20210309183A1 · Bielby · 2021 [cited by examiner]
US 20210312915A1 · Shaked · 2021 [cited by examiner]
EP 1699042A1 · 2006 [cited by applicant]
EP 2028061A2 · 2009 [cited by applicant]
EP 3470276A1 · 2019 [cited by applicant]
IEC 61508, “Functional Safety of Electrical/Electronic/Programmable Electronic Safety-related Systems,” https://en.wikipedia.org/wiki/IEC_61508, accessed on Apr. 1, 2022, 7 pgs. [cited by applicant]
ISO 26262, “Road vehicle—Functional safety,” International standard for functional safety of electronic system, https://en.wikipedia.org/wiki/ISO_26262, accessed on Sep. 13, 2021, 8 pgs. [cited by applicant]
Bhattacharya, Sumit; Notice of Allowance for U.S. Appl. No. 16/867,395, filed May 5, 2020, mailed Sep. 30, 2022, 12 pgs. [cited by applicant]
Avadhanam, Niranjan; International Preliminary Report on Patentability for PCT Application No. PCT/US2021/030585, filed May 4, 2021; mailed Nov. 17, 2022, 9 pgs. [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Society of Automotive Engineers (SAE), Standard No. J3016-201609, pp. 30 (Sep. 30, 2016). [cited by applicant]
“Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Society of Automotive Engineers (SAE), Standard No. J3016-201806, pp. 35 (Jun. 15, 2018). [cited by applicant]
Bhattacharya, Sumit; International Search Report and Written Opinion for International Application No. PCT/US2021/030585, filed May 4, 2021, mailed Jul. 20, 2021, 15 pgs. [cited by applicant]