IP Library Granted Patent US 12,283,117
Granted Patent B2
US 12,283,117 · App. 18/418,670 · Granted Apr 22, 2025

Systems and methods for automated license plate recognition

Inventors: Morgan Kohler (Mill Valley, CA); Bo Shen (Fremont, CA)
Assignee: Hayden AI Technologies, Inc.
G06V20/625G06V10/95G06V20/58G06V30/153G06V30/18G06V30/19147G06V30/1916
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,117
App. No.
18/418,670
Granted
Apr 22, 2025
Kind
B2
Abstract

Disclosed herein are methods, systems, and apparatus for automated license plate recognition and methods for training a machine learning model to undertake automated license plate recognition. For example, a method can comprise dividing an image or video frame comprising a license plate into a plurality of image patches, determining a positional vector for each of the image patches, adding the positional vector to each of the image patches and inputting the image patches and their associated positional vectors to a text-adapted vision transformer. The text-adapted vision transformer can be configured to output a prediction concerning the license plate number of the license plate.

Claims (27)

1. One or more non-transitory computer-readable media comprising instructions stored thereon, that when executed by one or more processors, cause the one or more processors to perform operations comprising:

dividing an image or video frame comprising a license plate in the image or video frame into a plurality of image patches, wherein the image or video frame is divided horizontally and vertically to obtain the image patches, and wherein at least one of the image patches comprises a portion of a character of a license plate number of the license plate;

determining a positional vector for each of the image patches, wherein the positional vector represents a spatial position of each of the image patches in the image or video frame;

adding the positional vector to each of the image patches and inputting the image patches and their associated positional vectors to a transformer encoder of a text-adapted vision transformer run on one or more devices; and

obtaining a prediction, outputted by the text-adapted vision transformer, concerning the license plate number of the license plate.

2. The one or more non-transitory computer-readable media of claim 1 , wherein the image or video frame is divided horizontally and vertically into a N×N grid of image patches, wherein N is an integer between 4 and 256.

3. The one or more non-transitory computer-readable media of claim 2 , wherein the text-adapted vision transformer comprises a linear projection layer, wherein the linear projection layer is configured to flatten each image patch of the N×N grid of image patches to 1×((H/N)×(W/N)), wherein a resulting input to the transformer encoder is then M×1×((H/N)×(W/N)), wherein H is a height of the image or video frame in pixels, wherein W is a width of the image or video frame in pixels, and wherein M is equaled to N multiplied by N.

4. The one or more non-transitory computer-readable media of claim 3 , further comprising adding a two-dimensional vector representing a spatial position of each of the image patches to the image patches.

5. The one or more non-transitory computer-readable media of claim 1 , wherein the text-adapted vision transformer is run on an edge device, wherein the edge device is coupled to a carrier vehicle, and wherein the image or video frame is captured using one or more cameras of the edge device while the carrier vehicle is in motion.

6. The one or more non-transitory computer-readable media of claim 1 , wherein the text-adapted vision transformer is run on a server, wherein the image or video frame is captured using one or more cameras of an edge device communicatively coupled to the server, wherein the image or video frame is transmitted by the edge device to the server, wherein the edge device is coupled to a carrier vehicle, and wherein the image or video frame is captured by the one or more cameras of the edge device while the carrier vehicle is in motion.

7. The one or more non-transitory computer-readable media of claim 1 , wherein each character of the license plate number is separately predicted by the transformer encoder of the text-adapted vision transformer.

8. The one or more non-transitory computer-readable media of claim 1 , wherein the text-adapted vision transformer is trained on a plate dataset comprising a plurality of real plate-text pairs, wherein the real plate-text pairs comprise images or video frames of real-life license plates and an annotated license plate number associated with each of the real-life license plates.

9. The one or more non-transitory computer-readable media of claim 8 , wherein the real-life license plates in the plate dataset comprise license plates with differing U.S. state plate aesthetics or configurations, license plates with differing non-U.S. country or region plate aesthetics or configurations, license plates with differing plate character configurations or styles, license plates with differing levels of blur associated with the images or video frames, and license plate with differing levels of exposure associated with the images or video frames.

10. The one or more non-transitory computer-readable media of claim 8 , wherein the text-adapted vision transformer is pre-trained on a plurality of image-text pairs prior to being trained on the plate dataset, wherein the image-text pairs comprise images or video frames of real-life objects comprising text and an annotation of the text.

11. The one or more non-transitory computer-readable media of claim 8 , wherein the text-adapted vision transformer is further trained on artificially-generated plate-text pairs, wherein the artificially-generated plate-text pairs comprise images of non-real license plates artificially generated by a latent diffusion model and a non-real license plate number associated with each of the non-real license plates.

12. The one or more non-transitory computer-readable media of claim 11 , wherein the latent diffusion model is trained using the plate dataset used to train the text-adapted vision transformer.

13. The one or more non-transitory computer-readable media of claim 11 , wherein the non-real license plate number is generated by a random plate number generator, and wherein at least one of the images of the non-real license plates is generated based on the non-real license plate number and one or more plate features provided as inputs to the latent diffusion model.

14. The one or more non-transitory computer-readable media of claim 13 , wherein the one or more plate features comprise at least one of a U.S. state plate aesthetic or configuration, a non-U.S. country or region plate aesthetic or configuration, a plate configuration or style, a level of noise associated with a license plate image, a level of blur associated with the license plate image, and a level of exposure associated with the license plate image.

15. One or more non-transitory computer-readable media comprising instructions stored thereon, that when executed by one or more processors, cause the one or more processors to perform operations comprising:

pre-training a text-adapted vision transformer on a plurality of image-text pairs, wherein the image-text pairs comprise images or video frames of real-life objects comprising text and an annotation of the text;

training the text-adapted vision transformer on a plate dataset comprising a plurality of real plate-text pairs, wherein the real plate-text pairs comprise images or video frames of real-life license plates and an annotated license plate number associated with each of the real-life license plates; and

further training the text-adapted vision transformer on artificially-generated plate-text pairs, wherein the artificially-generated plate-text pairs comprise images of non-real license plates artificially generated by a latent diffusion model and a non-real license plate number associated with each of the non-real license plates.

16. The one or more non-transitory computer-readable media of claim 15 , wherein the non-real license plate number is generated by a random plate number generator, and wherein at least one of the images of the non-real license plates is generated based on the non-real license plate number and one or more plate features provided as inputs to the latent diffusion model.

17. The one or more non-transitory computer-readable media of claim 16 , wherein the one or more plate features comprise at least one of a U.S. state plate aesthetic or configuration, a non-U.S. country or region plate aesthetic or configuration, a plate character configuration or style, a level of blur associated with the artificially-generated image, and a level of exposure associated with the artificially-generated image.

18. The one or more non-transitory computer-readable media of claim 16 , wherein the one or more plate features are selected or changed based on an accuracy of predictions made by the text-adapted vision transformer.

19. The one or more non-transitory computer-readable media of claim 18 , wherein the instructions further cause the one or more processors to perform operations comprising providing a prompt to the latent diffusion model to generate additional images of non-real license plates based in part on common plate features resulting in low accuracy predictions made by the text-adapted vision transformer.

20. The one or more non-transitory computer-readable media of claim 15 , wherein the latent diffusion model is trained using the plate dataset used to train the text-adapted vision transformer.

Assignments (2)
SECURITY INTEREST Recorded Oct 27, 2025
From: HAYDEN AI TECHNOLOGIES INC.
To: BANK OF MONTREAL
Reel/Frame 072691/0585 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2024
From: KOHLER, MORGAN; SHEN, BO
To: HAYDEN AI TECHNOLOGIES, INC.
Reel/Frame 066198/0495 →
Continuity (2)
Continuation 18458631 · Aug 30, 2023
Related Publication 20250078538A1 · Mar 6, 2025
References Cited (135)
US 9070289B2 · Saund et al. · 2015 [cited by applicant]
US 10296794B2 · Ratti · 2019 [cited by applicant]
US 10438083B1 · Rivard · 2019 [cited by examiner]
US 10963719B1 · Hantehzadeh · 2021 [cited by examiner]
US 11003919B1 · Ghadiok et al. · 2021 [cited by applicant]
US 11164014B1 · Ghadiok et al. · 2021 [cited by applicant]
US 11322017B1 · Ghadiok et al. · 2022 [cited by applicant]
US 11361558B2 · Seo · 2022 [cited by applicant]
US 11688182B2 · Ghadiok et al. · 2023 [cited by applicant]
US 11689701B2 · Ghadiok · 2023 [cited by examiner]
US 11915499B1 · Kohler et al. · 2024 [cited by applicant]
US 20020072847A1 · Trajkovic et al. · 2002 [cited by applicant]
US 20100081200A1 · Rajala et al. · 2010 [cited by applicant]
US 20120148092A1 · Ni et al. · 2012 [cited by applicant]
US 20130266188A1 · Bulan et al. · 2013 [cited by applicant]
US 20140007762A1 · Gavish et al. · 2014 [cited by applicant]
US 20140036076A1 · Nerayoff et al. · 2014 [cited by applicant]
US 20140219561A1 · Nakamura · 2014 [cited by examiner]
US 20140311456A1 · Richter et al. · 2014 [cited by applicant]
US 20150054975A1 · Emmett · 2015 [cited by examiner]
US 20180121744A1 · Kim · 2018 [cited by examiner]
US 20180172454A1 · Ghadiok et al. · 2018 [cited by applicant]
US 20180240336A1 · Kareev et al. · 2018 [cited by applicant]
US 20190137280A1 · Ghadiok et al. · 2019 [cited by applicant]
US 20190197369A1 · Law et al. · 2019 [cited by applicant]
US 20190251369A1 · Popov · 2019 [cited by examiner]
US 20200063866A1 · Reinhart et al. · 2020 [cited by applicant]
US 20200177767A1 · Kelly et al. · 2020 [cited by applicant]
US 20200380270A1 · Cox et al. · 2020 [cited by applicant]
US 20200401833A1 · Popov · 2020 [cited by examiner]
US 20210027083A1 · Cohen · 2021 [cited by examiner]
US 20210166145A1 · Omari et al. · 2021 [cited by applicant]
US 20210209941A1 · Maheshwari et al. · 2021 [cited by applicant]
US 20210237737A1 · Al-Nuaimi et al. · 2021 [cited by applicant]
US 20210241003A1 · Seo · 2021 [cited by applicant]
US 20210342624A1 · Wang · 2021 [cited by examiner]
US 20220122351A1 · Chen · 2022 [cited by examiner]
US 20220147745A1 · Ghadiok et al. · 2022 [cited by applicant]
US 20230147685A1 · Koch · 2023 [cited by examiner]
US 20230162481A1 · Yuan · 2023 [cited by examiner]
US 20230186668A1 · Dong · 2023 [cited by examiner]
US 20230196710A1 · Pan · 2023 [cited by examiner]
US 20230245435A1 · Zhao · 2023 [cited by examiner]
CN 2277104 · 1998 [cited by applicant]
CN 101751785 · 2010 [cited by applicant]
CN 101751785A · 2010 [cited by examiner]
CN 101789080 · 2010 [cited by applicant]
CN 101789080A · 2010 [cited by examiner]
CN 103971097 · 2014 [cited by applicant]
CN 103971097A · 2014 [cited by examiner]
CN 106407981 · 2017 [cited by applicant]
CN 106407981A · 2017 [cited by examiner]
CN 106560861 · 2017 [cited by applicant]
CN 106650729 · 2017 [cited by applicant]
CN 106650729A · 2017 [cited by examiner]
CN 109858327 · 2019 [cited by applicant]
CN 109858327A · 2019 [cited by examiner]
CN 110197589 · 2019 [cited by applicant]
CN 110321823 · 2019 [cited by applicant]
CN 110717433 · 2020 [cited by applicant]
CN 111368687 · 2020 [cited by applicant]
CN 111492416 · 2020 [cited by applicant]
CN 111666853 · 2020 [cited by applicant]
JP 4805763 · 2008 [cited by applicant]
KR 20030009149 · 2003 [cited by applicant]
KR 20030009149A · 2003 [cited by examiner]
KR 100812397 · 2008 [cited by applicant]
KR 100812397B1 · 2008 [cited by examiner]
KR 101607912 · 2016 [cited by applicant]
KR 101607912B1 · 2016 [cited by examiner]
WO WO2010081200 · 2010 [cited by applicant]
WO WO2014007762 · 2014 [cited by applicant]
WO WO2020063866 · 2020 [cited by applicant]
WO WO2020177767 · 2020 [cited by applicant]
WO WO2022099237 · 2022 [cited by applicant]
License Plate Recognition From Still Images and Video Sequences: A Survey, Christos-Nikolaos E et al., IEEE, 2008, pp. 377-391 (Year: 2008). [cited by examiner]
Toward End-to-End Car License Plate Detection and Recognition With Deep Neural Networks, Hui Li et al., IEEE, 2019, pp. 1126-1136 (Year: 2019). [cited by examiner]
Vehicle License Plate Recognition Based on Extremal Regions and Restricted Boltzmann Machines, Chao Gou et al., IEEE, 2016, pp. 1096-1107 (Year: 2016). [cited by examiner]
MultiPath ViT OCR: A Lightweight Visual Transformer-based License Plate Optical Character Recognition, Alireza Azadbakht et al., ICCKE , 2022, pp. 092-095 (Year: 2022). [cited by examiner]
License Plate Character Recognition via Signature Analysis and Features Extraction, Lorita Angeline, et al., IEEE, 2012, pp. 1-6 (Year: 2012). [cited by examiner]
A robust license plate detection and recognition system based on DETR and CNN, Elsevier, 2022, pp. 1-10 (Year: 2022). [cited by examiner]
A Robust License Plate Recognition Model Based on Bi-LSTM, Yongjie Zou et al., IEEE, 2020, pp. 211630-211641 (Year: 2020). [cited by examiner]
License Plate Segmentation and Recognition of Chinese Vehicle Based on BPNN, Wang Naiguo, Zhu Xiangwei et al., Computer Society, 2016, pp. 403-406 (Year: 2016). [cited by examiner]
Research on License Plate Recognition Algorithms Based on Deep Learning in Complex Environment, Wang Weihong et al., IEEE, 2020, pp. 91661-91675 (Year: 2020). [cited by examiner]
RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition, Xiaoyu Yue et al., Springer, 2020, pp. 135-151 (Year: 2020). [cited by examiner]
Vehicle Plate Recognition for Wireless Traffic Control and Law Enforcement System, Francisco Alegria et al., IEEE, 2006, pp. 1800-1804 (Year: 2006). [cited by examiner]
Safe Fleet ClearLane, 2021, pp. 1-2 (Year: 2021). [cited by examiner]
ALPR—An Intelligent Approach Towards Detection and Recognition of License Plates in Uncontrolled Environments, Akshay Baksh et al., Springer, 2023, pp. 253-269 (Year: 2023). [cited by examiner]
Robust license plate detection and recognition with automatic rectification, Degui Xiao et al., JOEI, 2021, pp. 013002-1 to 013002-20 (Year: 2021). [cited by examiner]
Traffic Signal Violation Detection using Artificial Intelligence and Deep Learning, Ruben J. Franklin et al., IEEE, 2020, pp. 839-844 (Year: 2020). [cited by examiner]
Traffic Rules Violation Detection using Deep Learning, Aniruddha Tonge et al., IEEE, 2020, pp. 1250-1257 (Year: 2020). [cited by examiner]
Alireza Azadbakht et al., MultiPath Vit OCR: A Lightweight Visual Transformer-based License Plate Optical Character Recognition, ICCKE, 2022, pp. 029-095 (Year: 2022). [cited by applicant]
Chao Gou et al., Vehicle License Plate Recognition Based on Extremal Regions and Restricted Boltzmann Machines, IEEE, 2016, pp. 1096-1107 (Year: 2016). [cited by applicant]
Christos-Nikolaos E et al., License Plate Recognition From Still Images and Video Sequences: A Survey, IEEE, 2008, pp. 377-391 (Year: 2008). [cited by applicant]
Clearlane, “Automated Bus Lane Enforcement System”, [cited by applicant]
Elsevier, A robust license plate detection and recognition system based on DETR and CNN, pp. 1-10 (Year 2022). [cited by applicant]
Francisco Alegria et al., Vehicle Plate Recognition for Wireless Traffic Control and Law Enforcement System, IEEE, 2006, pp. 1800-1804 (Year: 2006). [cited by applicant]
Hui Li et al., Toward End-to-End Car License Plate Detection and Recognition With Deep Neural Networks, IEEE, 2019, pp. 1126-1136 (Year: 2019). [cited by applicant]
Lorita Angeline et al., License Plate Character Recognition via Signature Analysis and Features Extraction, IEEE, 2012, pp. 1-6 (Year: 2012). [cited by applicant]
Tongjie Zou et al., A Robust License Plate Recognition Model Based on Bi-LSTM, IEEE, 2020, p. 211630-211641 (Year: 2020). [cited by applicant]
Wang Weihong et al., Research on License Plate Recognition Algorithms Based on Deep Learning in Complex Environment, IEEE, 2020, pp. 91661-91675 (Year: 2020). [cited by applicant]
Xiaoyu Yue et al., RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition, Springer, 2020, pp. 135-151 (Year: 2020). [cited by applicant]
Zhu Ziangwei et al., License Plate Segmentation and Recognition of Chinese Vehicle Based on BPNN, Wang Naiguo, Computer Society, 2016, pp. 403-406 (Year: 2016). [cited by applicant]
“Consulting services in Computer Vision and AI,” accessed on May 8, 2023, online. [cited by applicant]
“Safety Vision Announces Smart Automated Bus Lane Enforcement (SABLE™) Solution,” [cited by applicant]
Bo, T. et al., “Common phase error estimation in coherent optical OFDM systems using best-fit bounding box,” [cited by applicant]
Bo, T. et al., “Common Phase Estimation in Coherent OFDM System Using Image Processing Technique,” [cited by applicant]
Bo, T. et al., “Image Processing Based Common Phase Estimation for Coherent Optical Orthogonal Frequency Division Multiplexing System,” [cited by applicant]
Canizo, M. et al. “Multi-Head CNN-RNN for multi time series anomaly detection: An industrial case study,” [cited by applicant]
Chen, S. et al., “A Dense Feature Pyramid Network-Based Deep Learning Model for Road Marking Instance Segmentation Using MLS Point Clouds,” [cited by applicant]
Chhaya, S. et al., “Basic Geometric Shape and Primary Colour Detection Using Image Processing on MATLAB,” [cited by applicant]
Clearlane, “The Safe Fleet Automated Bus Lane Enforcement (ABLE)”, [cited by applicant]
Evanko, K. “Siemens Mobility launches first-ever mobile bus lane enforcement solution in New York,” [cited by applicant]
Fan, Y. et al., “A Coarse-to-Fine Framework for Multiple Pedestrian Crossing Detection,” [cited by applicant]
Franklin, R. “Traffic Signal Violation Detection using Artificial Intelligence and Deep Learning,” [cited by applicant]
Github Repository, Our Camera, <https://github.com/Bellspringsteen/OurCamera> (last visited Sep. 25, 2023). [cited by applicant]
Glenn, J., “Adaptive Morphological Feature-Based Object Classifier for a Color Imaging System,” [cited by applicant]
Hsu, K., et al. “Augmented Multiple Instance Regression for Inferring Object Contours in Bounding Boxes,” [cited by applicant]
Huval, B. et al. “An Empirical Evaluation of Deep Learning on Highway Driving,” [cited by applicant]
Liu, X. “Vehicle-Related Scene Understanding Using Deep Learning,” [cited by applicant]
Nehemiah, A. et al. “Deep Learning for Automated Driving with MATLAB,” [cited by applicant]
Novak, L., “Vehicle Detection and Pose Estimation for Autonomous Driving,” [cited by applicant]
Oh, J. et al., “Context-based abnormal object detection using the fully-connected conditional random fields,” [cited by applicant]
Paquet, E. et al., “Description of shape information for 2-D and 3-D objects,” [cited by applicant]
Safe Fleet. “Whitepaper: Vendor Interoperability for ABLE,” [cited by applicant]
Sengupta, S., “Semantic Mapping of Road Scenes,” [cited by applicant]
Siemens Mobility Inc., “Ratification of Completed Procurement Actions,” New York City Transit and Siemens Mobility Inc., [cited by applicant]
Siemens Mobility Traffic Solutions, “Enforcement solutions for safe and efficient cities,” [cited by applicant]
Spencer, B. et al., “NYC extends Brooklyn bus lane enforcement,” [cited by applicant]
Sullivan, T., “Transit Bus Surveillance Solutions,” [cited by applicant]
Tonge, A. et al., “Traffic Rules Violation Detection using Deep Learning,” [cited by applicant]
Viorel, C., “Some Aspects Concerning Geometric Forms Automatically Find Images and Ordering Them Using Robot Studio Simulation,” [cited by applicant]
Wu, C. et al., “Adjacent Lane Detection and Lateral Vehicle Distance Measurement Using Vision-Based Neuro-Fuzzy Approaches,” [cited by applicant]
Zhao Z. et al., “Deep Reinforcement Learning Based Lane Detection and Localization,” [cited by applicant]
Zhou, C. et al., “Predicting the passenger demand on bus services for mobile users,” [cited by applicant]