IP Library Granted Patent US 12,347,143
Granted Patent B2
US 12,347,143 · App. 17/825,519 · Granted Jul 1, 2025

Reinforcement-learning based system for camera parameter tuning to improve analytics

Inventors: Kunal Rao (Monroe, NJ); Giuseppe Coviello (Robbinsville, NJ); Murugan Sankaradas (Dayton, NJ); Oliver Po (San Jose, CA); Srimat Chakradhar (Manalapan, NJ); Sibendu Paul (West Lafayette, IN)
Assignee: NEC Corporation
G06T7/80G06T7/0002G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,143
App. No.
17/825,519
Granted
Jul 1, 2025
Kind
B2
Abstract

A method for automatically adjusting camera parameters to improve video analytics accuracy during continuously changing environmental conditions is presented. The method includes capturing a video stream from a plurality of cameras, performing video analytics tasks on the video stream, the video analytics tasks defined as analytics units (AUs), applying image processing to the video stream to obtain processed frames, filtering the processed frames through a filter to discard low-quality frames and dynamically fine-tuning parameters of the plurality of cameras. The fine-tuning includes passing the filtered frames to an AU-specific proxy quality evaluator, employing State-Action-Reward-State-Action (SARSA) reinforcement learning (RL) computations to automatically fine-tune the parameters of the plurality of cameras, and based on the reinforcement computations, applying a new policy for an agent to take actions and learn to maximize a reward.

Claims (46)

1. A method for automatically adjusting camera parameters to improve video analytics accuracy during continuously changing environmental conditions, the method comprising:

capturing a video stream from a plurality of cameras;

performing video analytics tasks on the video stream, the video analytics tasks defined as analytics units (AUs);

applying image processing to the video stream to obtain processed frames;

filtering the processed frames through a filter to discard low-quality frames; and

dynamically fine-tuning parameters of the plurality of cameras by:

passing the filtered frames to an AU-specific proxy quality evaluator;

employing State-Action-Reward-State-Action (SARSA) reinforcement learning (RL) computations to automatically fine-tune the parameters of the plurality of cameras; and

based on the reinforcement computations, applying a new policy for an agent to take actions and learn to maximize a reward, wherein the SARSA RL computations include a state, an action, and a reward, the state being a vector including current brightness, contrast, sharpness and color parameter values of a camera of the plurality of cameras and a measure of brightness, contrast, sharpness and color, the action is an increase or decrease of one of the brightness, contrast, sharpness or color parameter values or taking no action at all, and the reward is a AU-specific quality evaluator's output.

2. The method of claim 1 , wherein a virtual-camera (VC) is employed to:

apply different environmental characteristics and evaluate different camera settings on an exact same scene captured from the video stream; and

rapidly and independently train and test RL algorithms, and various reward functions used for RL.

3. The method of claim 2 , wherein the VC includes an offline profiling phase and an online phase.

4. The method of claim 3 , wherein, during the offline profiling phase, a VC-camera table and a mapping function are generated, the mapping function mapping a particular time in a day to its corresponding brightness, contrast, color-saturation, and sharpness feature values observed during that time.

5. The method of claim 3 , wherein, during the online phase, a frame for a different environmental condition corresponding to a time of day other than its capture time is simulated.

6. The method of claim 1 , wherein an analytic quality (AQ) metric specific for each of the AUs is employed to train an AU-specific quality evaluator.

7. The method of claim 6 , wherein, in the absence of ground truth in real-world deployments, the AU-specific quality evaluator is used as a proxy to evaluate the AUs.

8. A non-transitory computer-readable storage medium comprising a computer-readable program for automatically adjusting camera parameters to improve video analytics accuracy during continuously changing environmental conditions, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of:

capturing a video stream from a plurality of cameras;

performing video analytics tasks on the video stream, the video analytics tasks defined as analytics units (AUS);

applying image processing to the video stream to obtain processed frames;

filtering the processed frames through a filter to discard low-quality frames; and

dynamically fine-tuning parameters of the plurality of cameras by:

passing the filtered frames to an AU-specific proxy quality evaluator;

employing State-Action-Reward-State-Action (SARSA) reinforcement learning (RL) computations to automatically fine-tune the parameters of the plurality of cameras; and

based on the reinforcement computations, applying a new policy for an agent to take actions and learn to maximize a reward, wherein a virtual-camera (VC) is employed to:

apply different environmental characteristics and evaluate different camera settings on an exact same scene captured from the video stream; and

rapidly and independently train and test RL algorithms, and various reward functions used for RL

wherein the VC includes an offline profiling phase and an online phase such that during the online phase, a frame for a different environmental condition corresponding to a time of day other than its capture time is simulated.

9. The non-transitory computer-readable storage medium of claim 8 , wherein, during the offline profiling phase, a VC-camera table and a mapping function are generated, the mapping function mapping a particular time in a day to its corresponding brightness, contrast, color-saturation, and sharpness feature values observed during that time.

10. The non-transitory computer-readable storage medium of claim 8 , wherein an analytic quality (AQ) metric specific for each of the AUs is employed to train an AU-specific quality evaluator.

11. The non-transitory computer-readable storage medium of claim 10 , wherein, in the absence of ground truth in real-world deployments, the AU-specific quality evaluator is used as a proxy to evaluate the AUs.

12. A system for automatically adjusting camera parameters to improve video analytics accuracy during continuously changing environmental conditions, the system comprising:

a memory; and

one or more processors in communication with the memory configured to:

capture a video stream from a plurality of cameras;

perform video analytics tasks on the video stream, the video analytics tasks defined as analytics units (AUs);

apply image processing to the video stream to obtain processed frames;

filter the processed frames through a filter to discard low-quality frames; and

dynamically fine-tune parameters of the plurality of cameras by:

passing the filtered frames to an AU-specific proxy quality evaluator;

employing State-Action-Reward-State-Action (SARSA) reinforcement learning (RL) computations to automatically fine-tune the parameters of the plurality of cameras; and

based on the reinforcement computations, applying a new policy for an agent to take actions and learn to maximize a reward, wherein a virtual-camera (VC) is employed to:

apply different environmental characteristics and evaluate different camera settings on an exact same scene captured from the video stream; and

rapidly and independently train and test RL algorithms, and various reward functions used for RL, wherein the VC includes an offline profiling phase and an online phase such that during the online phase, a frame for a different environmental condition corresponding to a time of day other than its capture time is simulated.

13. The system of claim 12 , wherein, during the offline profiling phase, a VC-camera table and a mapping function are generated, the mapping function mapping a particular time in a day to its corresponding brightness, contrast, color-saturation, and sharpness feature values observed during that time.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 071095/0825 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2022
From: RAO, KUNAL; COVIELLO, GIUSEPPE; SANKARADAS, MURUGAN; PO, OLIVER; CHAKRADHAR, SRIMAT; PAUL, SIBENDU
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 060028/0748 →
Continuity (2)
Provisional Application 63196399 · Jun 3, 2021
Related Publication 20220414935A1 · Dec 29, 2022
References Cited (19)
US 20180108121A1 · McCaughan · 2018 [cited by examiner]
US 20190122378A1 · Aswin · 2019 [cited by examiner]
US 20200221009A1 · Citerin · 2020 [cited by examiner]
US 20200342652A1 · Rowell · 2020 [cited by examiner]
Canel, C., Kim, T., Zhou, G., Li, C., Lim, H., Andersen, D. G., . . . & Dulloor, S. (Apr. 15, 2019). Scaling video analytics on constrained edge nodes. Proceedings of Machine Learning and Systems, 1, 406-417. [cited by applicant]
Dong, J., Frosio, I., & Kautz, J. (Jan. 28, 2018). Learning adaptive parameter tuning for image processing. Electronic Imaging, 2018(13), 196-1. [cited by applicant]
Du, K., Pervaiz, A., Yuan, X., Chowdhery, A., Zhang, Q., Hoffmann, H., & Jiang, J. (Jul. 30, 2020). Server-driven video streaming for deep learning inference. In Proceedings of the Annual conference of the ACM Special I… [cited by applicant]
Heide, F., Steinberger, M., Tsai, Y. T., Rouf, M., Pajk, D., Reddy, D., . . . & Pulli, K. (Nov. 19, 2014). Flexisp: A flexible camera image processing framework. ACM Transactions on Graphics (ToG), 33(6), 1-13. [cited by applicant]
Jiang, J., Ananthanarayanan, G., Bodik, P., Sen, S., & Stoica, I. (Aug. 7, 2018). Chameleon: scalable adaptation of video analytics. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Commun… [cited by applicant]
Kang, D., Emmons, J., Abuzaid, F., Bailis, P., & Zaharia, M. (Mar. 7, 2017). Noscope: optimizing neural network queries over video at scale. arXiv preprint arXiv:1703.02529. [cited by applicant]
Zhang, H., Ananthanarayanan, G., Bodik, P., Philipose, M., Bahl, P., & Freedman, M. J. (Mar. 27, 2017). Live Video Analytics at Scale with Approximation and {Delay-Tolerance}. In 14th USENIX Symposium on Networked Syste… [cited by applicant]
Zhang, B., Jin, X., Ratnasamy, S., Wawrzynek, J., & Lee, E. A. (Aug. 7, 2018). Awstream: Adaptive wide-area streaming analytics. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communicat… [cited by applicant]
Wu, C. T., Isikdogan, L. F., Rao, S., Nayak, B., Gerasimow, T., Sutic, A., . . . & Michael, G. (Sep. 22, 2019). VisionISP: Repurposing the image signal processor for computer vision applications. In 2019 IEEE Internatio… [cited by applicant]
Wang, Y., Wang, W., Zhang, J., Jiang, J., & Chen, K. (2019, July 8). Bridging the {Edge-Cloud} Barrier for Real-time Advanced Vision Analytics. In 11th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 19). [cited by applicant]
Schwartz, E., Giryes, R., & Bronstein, A. M. (Oct. 1, 2018). Deepisp: Toward learning an end-to-end image processing pipeline. IEEE Transactions on Image Processing, 28(2), 912-923. [cited by applicant]
Paul, S., Drolia, U., Hu, Y. C., & Chakradhar, S. T. (Dec. 14, 2021). Aqua: Analytical quality assessment for optimizing video analytics systems. In 2021 IEEE/ACM Symposium on Edge Computing (SEC) (pp. 135-147). IEEE. [cited by applicant]
Li, Y., Padmanabhan, A., Zhao, P., Wang, Y., Xu, G. H., & Netravali, R. (Jul. 30, 2020). Reducto: On-camera filtering for resource-efficient real-time video analytics. In Proceedings of the Annual conference of the ACM … [cited by applicant]
Chen, T. Y. H., Ravindranath, L., Deng, S., Bahl, P., & Balakrishnan, H. (Nov. 1, 2015). Glimpse: Continuous, real-time object recognition on mobile devices. In Proceedings of the 13th ACM Conference on Embedded Network… [cited by applicant]
Jang, S. Y., Lee, Y., Shin, B., & Lee, D. (Oct. 25, 2018). Application-aware loT camera virtualization for video analytics edge computing. In 2018 IEEE/ACM Symposium on Edge Computing (SEC) (pp. 132-144). IEEE. [cited by applicant]