IP Library › Granted Patent US 12,603,101
Granted Patent B2
US 12,603,101 · App. 18/087,323 · Granted Apr 14, 2026

Machine learning methods for evaluating vehicle conditions

Inventors: Philip Schneider (Amherst, NY); Michael Pokora (Tonawanda, NY); Justas Birgiolas (Lockport, NY); Livio Forte, III (Lloyd Harbor, NY); Dennis Christopher Fedorishin (East Amherst, NY)
Assignee: ACV Auctions Inc.
G10L25/78G01M13/028G01M15/02G01M15/12G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,603,101
App. No.
18/087,323
Granted
Apr 14, 2026
Kind
B2
Abstract

Techniques for using a trained machine learning (ML) model to detect presence of vehicle defects from audio acquired at least in part during operation of an engine of a vehicle. The techniques include using at least one computer hardware processor to perform: obtaining, via at least one communication network, a first audio recording that was acquired, using at least one acoustic sensor, at least in part during operation of the engine; processing the first audio recording using the trained ML model to detect, from the first audio recording, presence of at least one vehicle defect, the processing comprising: generating an audio waveform from the first audio recording, generating a two-dimensional (2D) representation of the audio waveform, and processing the audio waveform and the 2D representation of the audio waveform using the trained ML model to obtain output indicative of presence or absence of the at least one vehicle defect.

Claims (71)

1 . A method for using a trained machine learning (ML) model to detect presence of vehicle defects from audio acquired at least in part during operation of an engine of a vehicle, the method comprising:

using at least one computer hardware processor to perform:

obtaining, via at least one communication network, a first audio recording that was acquired, using at least one acoustic sensor, at least in part during operation of the engine;

processing the first audio recording using the trained ML model to detect, from the first audio recording, presence of at least one vehicle defect, the processing comprising:

generating an audio waveform from the first audio recording,

generating a two-dimensional (2D) representation of the audio waveform, and

processing the audio waveform using the trained ML model and processing the 2D representation of the audio waveform using the trained ML model to obtain output indicative of presence or absence of the at least one vehicle defect, wherein the trained ML model comprises:

a first neural network portion comprising a plurality of one-dimensional (1D) convolutional layers configured to process the audio waveform;

a second neural network portion comprising a plurality of 2D convolutional layers configured to process the 2D representation of the audio waveform; and

a fusion neural network portion comprising one or more fully connected layers configured to combine outputs produced by the first neural network portion and the second neural network portion to obtain the output indicative of the presence or absence of the at least one vehicle defect.

2 . The method of claim 1 , wherein the audio recording comprises at least a first waveform for at least a first audio channel, and wherein generating the audio waveform from the first audio recording comprises:

resampling the first waveform to a target frequency to obtain a resampled waveform;

normalizing the resampled waveform by subtracting its mean and dividing by its standard deviation to obtain a normalized waveform; and

clipping the normalized waveform to a target maximum to obtain the audio waveform.

3 . The method of claim 1 , wherein the audio waveform is between 5 and 45 seconds long and wherein the frequency of the audio waveform is between 8 and 45 KHz.

4 . The method of claim 1 , wherein generating the two-dimensional (2D) representation of the audio waveform comprises generating a time-frequency representation of the audio waveform.

5 . The method of claim 4 , wherein generating the time-frequency representation of the audio waveform comprises using a short-time Fourier transform, a wavelet transform, a Gabor transform, or a chirplet transform to generate the time-frequency representation.

6 . The method of claim 5 , wherein generating the time-frequency representation of the audio waveform comprises generating a Mel-scale log spectrogram from the audio waveform.

7 . The method of claim 1 , further comprising:

obtaining, via the at least one communication network, metadata indicating one or more properties of the vehicle,

wherein using the trained ML model to detect the presence of the at least one vehicle defect further comprises generating metadata features from the metadata, and

wherein processing the audio waveform and the 2D representation of the audio waveform comprises processing the audio waveform, the 2D representation of the audio waveform, and the metadata features using the trained ML model to obtain the output indicative of the presence or absence of the at least one vehicle defect.

8 . The method of claim 7 , wherein the properties of the vehicle are selected from the group consisting of: a reading of the vehicle's odometer, a model of the vehicle, a make of the vehicle, an age of the vehicle, a type of drivetrain in the vehicle, a type of transmission in the vehicle, a measure of displacement of the engine, a fuel type for the vehicle, an indication of whether on-board diagnostics (OBD) codes could be obtained from the vehicle, a number of incomplete readiness monitors reported by the OBD scanner, one or more BlackBook-reported engine properties, a list of one or more OBD codes, location of the vehicle, information about weather at the location of the vehicle, and information about a seller of the vehicle.

9 . The method of claim 7 , wherein the metadata comprises text indicating at least one of the one or more properties, and generating the metadata features from the metadata comprises generating a numeric representation of the text.

10 . The method of claim 1 , wherein the output is indicative of the presence or absence of abnormal internal engine noise, timing chain noise, engine accessory noise, and/or exhaust noise.

11 . The method of claim 1 , further comprising:

obtaining, via the at least one communication network, metadata indicating one or more properties of the vehicle,

wherein using the trained ML model to detect the presence of the at least one vehicle defect further comprises generating metadata features from the metadata,

wherein processing the audio waveform and the two-dimensional representation of the audio waveform comprises processing the audio waveform, the two-dimensional representation of the audio waveform, and the metadata features, using the trained ML model to obtain output indicative of presence of the at least one vehicle defect,

wherein the trained ML model further comprises a third neural network portion comprising one or more fully connected layers configured to process the metadata features, and

wherein the one or more fully connected layers of the fusion neural network are configured to combine outputs produced by the first neural network portion, the second neural network portion, and the third neural network portion to obtain the output indicative of the presence or absence of the at least one vehicle defect.

12 . The method of claim 11 ,

wherein the trained ML model has at least one million parameters, and

wherein processing the first audio recording using the trained ML model to detect the presence of the at least one vehicle defect comprises computing the output using values of the at least one million parameters, the audio waveform, and the 2D representation of the audio waveform.

13 . The method of claim 1 , further comprising:

acquiring, using the at least one acoustic sensor, the first audio recording at least in part during operation of the engine.

14 . The method of claim 1 , further comprising:

determining, based on the output, that the at least one vehicle defect was detected using the first audio recording, and

generating an electronic vehicle condition report indicating that the at least one vehicle defect was detected using the first audio recording and a measure of confidence in that detection.

15 . The method of claim 14 , further comprising:

transmitting the electronic vehicle condition report, via the at least one communication network, to a remote device of an inspector of the vehicle.

16 . The method of claim 15 , further comprising:

receiving a second audio recording, via the at least one communication network, from the remote device of the inspector of the vehicle, the second audio recording being acquired after transmission of the electronic vehicle condition report and using the at least one acoustic sensor at least in part during operation of the engine; and

processing the second audio recording using the trained ML model to detect, from the second audio recording, presence of the at least one vehicle defect, the processing comprising:

generating a second audio waveform from the second audio recording,

generating a second two-dimensional (2D) representation of the second audio waveform, and

processing the second audio waveform and the second 2D representation of the audio waveform using the trained ML model to obtain second output indicative of presence or absence of the at least one vehicle defect.

17 . The method of claim 1 ,

wherein obtaining the first audio recording comprises receiving the first audio recording from a mobile device, via the at least one communication network, by at least one computing device at a location remote from a location of the mobile device,

wherein the processing is performed by the at least one computing device, and

wherein the mobile device comprises a smart phone or a mobile vehicle diagnostic device.

18 . A system, comprising:

at least one computer hardware processor; and

at least one non-transitory computer-readable storage medium storing processor executable instructions that when executed by the at least one computer hardware processor perform a method for using a trained machine learning (ML) model to detect presence of vehicle defects from audio acquired at least in part during operation of an engine of a vehicle, the method comprising:

obtaining, via at least one communication network, a first audio recording that was acquired, using at least one acoustic sensor, at least in part during operation of the engine;

processing the first audio recording using the trained ML model to detect, from the first audio recording, presence of at least one vehicle defect, the processing comprising:

generating an audio waveform from the first audio recording,

generating a two-dimensional (2D) representation of the audio waveform, and

processing the audio waveform using the trained ML model and processing the 2D representation of the audio waveform using the trained ML model to obtain output indicative of presence or absence of the at least one vehicle defect, wherein the trained ML model comprises:

a first neural network portion comprising a plurality of one-dimensional (1D) convolutional layers configured to process the audio waveform;

a second neural network portion comprising a plurality of 2D convolutional layers configured to process the 2D representation of the audio waveform; and

a fusion neural network portion comprising one or more fully connected layers configured to combine outputs produced by the first neural network portion and the second neural network portion to obtain the output indicative of the presence or absence of the at least one vehicle defect.

19 . At least one non-transitory computer-readable storage medium storing processor executable instructions that, when executed by a at least one computer hardware processor, cause the at least one processor to perform a method for using a trained machine learning (ML) model to detect presence of vehicle defects from audio acquired at least in part during operation of an engine of a vehicle, the method comprising:

obtaining, via at least one communication network, a first audio recording that was acquired, using at least one acoustic sensor, at least in part during operation of the engine;

processing the first audio recording using the trained ML model to detect, from the first audio recording, presence of at least one vehicle defect, the processing comprising:

generating an audio waveform from the first audio recording,

generating a two-dimensional (2D) representation of the audio waveform, and

processing the audio waveform using the trained ML model and processing the 2D representation of the audio waveform using the trained ML model to obtain output indicative of presence or absence of the at least one vehicle defect, wherein the trained ML model comprises:

a first neural network portion comprising a plurality of one-dimensional (1D) convolutional layers configured to process the audio waveform;

a second neural network portion comprising a plurality of 2D convolutional layers configured to process the 2D representation of the audio waveform; and

a fusion neural network portion comprising one or more fully connected layers configured to combine outputs produced by the first neural network portion and the second neural network portion to obtain the output indicative of the presence or absence of the at least one vehicle defect.

Assignments (2)
SUPPLEMENT TO PATENT SECURITY AGREEMENT Recorded Aug 28, 2025
From: ACV AUCTIONS INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 072716/0682 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 26, 2023
From: SCHNEIDER, PHILIP; POKORA, MICHAEL; FORTE, LIVIO, III; FEDORISHIN, DENNIS CHRISTOPHER; BIRGIOLAS, JUSTAS
To: ACV AUCTIONS INC.
Reel/Frame 063448/0074 →
Continuity (3)
Provisional Application 63293534 · Dec 23, 2021
Provisional Application 63293558 · Dec 23, 2021
Related Publication 20230206942A1 · Jun 29, 2023
References Cited (83)
US 4215412A · Bernier et al. · 1980 [cited by applicant]
US 4375672A · Kato et al. · 1983 [cited by applicant]
US 5854993A · Grichnik · 1998 [cited by applicant]
US 6175787B1 · Breed · 2001 [cited by applicant]
US 6275765B1 · Divljakovic et al. · 2001 [cited by applicant]
US 7054596B2 · Arntz · 2006 [cited by applicant]
US 8311973B1 · Zadeh · 2012 [cited by applicant]
US 8437904B2 · Mansouri et al. · 2013 [cited by applicant]
US 9200981B2 · Horlbeck et al. · 2015 [cited by applicant]
US 10127591B1 · Wollmer et al. · 2018 [cited by applicant]
US 10554802B2 · Moore et al. · 2020 [cited by applicant]
US 10740404B1 · Hjermstad et al. · 2020 [cited by applicant]
US 11157835B1 · Hjermstad et al. · 2021 [cited by applicant]
US 11327726B2 · Wang et al. · 2022 [cited by applicant]
US 11631289B2 · Campanella et al. · 2023 [cited by applicant]
US 11783851B2 · Schneider et al. · 2023 [cited by applicant]
US 12046254B2 · Pokora et al. · 2024 [cited by applicant]
US 12062257B2 · Campanella et al. · 2024 [cited by applicant]
US 20040086135A1 · Vaishya · 2004 [cited by applicant]
US 20040176879A1 · Menon et al. · 2004 [cited by applicant]
US 20050096873A1 · Klein · 2005 [cited by applicant]
US 20050149234A1 · Vian et al. · 2005 [cited by applicant]
US 20050169484A1 · Cascone et al. · 2005 [cited by applicant]
US 20050171833A1 · Jost et al. · 2005 [cited by applicant]
US 20050192722A1 · Noguchi · 2005 [cited by applicant]
US 20060064231A1 · Fekete et al. · 2006 [cited by applicant]
US 20080192954A1 · Honji et al. · 2008 [cited by applicant]
US 20120323531A1 · Pascu et al. · 2012 [cited by applicant]
US 20130185078A1 · Tzirkel-Hancock et al. · 2013 [cited by applicant]
US 20130277529A1 · Bolliger · 2013 [cited by applicant]
US 20140096608A1 · Themm et al. · 2014 [cited by applicant]
US 20140162219A1 · Stankoulov · 2014 [cited by applicant]
US 20140201126A1 · Zadeh et al. · 2014 [cited by applicant]
US 20150019533A1 · Moody et al. · 2015 [cited by applicant]
US 20150100448A1 · Binion et al. · 2015 [cited by applicant]
US 20150333789A1 · An · 2015 [cited by applicant]
US 20160025027A1 · Mentele · 2016 [cited by applicant]
US 20160034590A1 · Endras et al. · 2016 [cited by applicant]
US 20160036899A1 · Moody et al. · 2016 [cited by applicant]
US 20160055737A1 · Boken · 2016 [cited by applicant]
US 20160112216A1 · Sargent et al. · 2016 [cited by applicant]
US 20160161299A1 · Campbell et al. · 2016 [cited by applicant]
US 20160342945A1 · Doranth et al. · 2016 [cited by applicant]
US 20160377500A1 · Bizub · 2016 [cited by applicant]
US 20170169399A1 · Areshidze et al. · 2017 [cited by applicant]
US 20170201779A1 · Publicover et al. · 2017 [cited by applicant]
US 20170213541A1 · MacNeille et al. · 2017 [cited by applicant]
US 20170356936A1 · Ismail et al. · 2017 [cited by applicant]
US 20170364776A1 · Micks et al. · 2017 [cited by applicant]
US 20170374460A1 · Jung et al. · 2017 [cited by applicant]
US 20180005463A1 · Siegel et al. · 2018 [cited by applicant]
US 20180025392A1 · Helstab · 2018 [cited by applicant]
US 20180150805A1 · Shaver et al. · 2018 [cited by applicant]
US 20180204111A1 · Zadeh et al. · 2018 [cited by applicant]
US 20180350167A1 · Ekkizogloy et al. · 2018 [cited by applicant]
US 20190017487A1 · Rudnitzki et al. · 2019 [cited by applicant]
US 20190080528A1 · Bednar et al. · 2019 [cited by applicant]
US 20190228596A1 · Mondello et al. · 2019 [cited by applicant]
US 20190287079A1 · Shiraishi et al. · 2019 [cited by applicant]
US 20190294878A1 · Endras et al. · 2019 [cited by applicant]
US 20200057487A1 · Sicconi et al. · 2020 [cited by applicant]
US 20200064227A1 · Im et al. · 2020 [cited by applicant]
US 20200118367A1 · Dudar · 2020 [cited by applicant]
US 20200134933A1 · Covington et al. · 2020 [cited by applicant]
US 20200234517A1 · Campanella · 2020 [cited by examiner]
US 20210123832A1 · Johnson et al. · 2021 [cited by applicant]
US 20230091331A1 · Gibson et al. · 2023 [cited by applicant]
US 20230186690A1 · Usami · 2023 [cited by applicant]
US 20230204461A1 · Schneider et al. · 2023 [cited by applicant]
US 20230267780A1 · Campanella et al. · 2023 [cited by applicant]
US 20230360667A1 · Schneider et al. · 2023 [cited by applicant]
US 20240331722A1 · Schneider et al. · 2024 [cited by applicant]
US 20240355156A1 · Campanella et al. · 2024 [cited by applicant]
US 20250117313A1 · Tyomkin et al. · 2025 [cited by applicant]
WO WO2008022289A2 · 2008 [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2020/014645 mailed May 22, 2020. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2020/014645 mailed Aug. 5, 2021. [cited by applicant]
Extended European Search Report for European Application No. 20745586.6 dated Dec. 12, 2022. [cited by applicant]
Invitation to Pay Additional Fees for International Application No. PCT/US2022/053850 mailed Apr. 3, 2023. [cited by applicant]
Bilen et al., A framework for the robust evaluation of sound event detection. 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). May 4, 2020:61-5. [cited by applicant]
Fedorishin et al., Large-Scale Acoustic Automobile Fault Detection: Diagnosing Engines Through Sound. Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining Aug. 14, 2022. 11 pages. [cited by applicant]
Fedorishin et al., Waveforms and Spectrograms: Enhancing Acoustic Scene Classification Using Multimodal Feature Fusion. DCASE 2021:216-20. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/053850 mailed Jun. 2, 2023. [cited by applicant]