IP Library Granted Patent US 12,248,949
Granted Patent B2
US 12,248,949 · App. 17/519,311 · Granted Mar 11, 2025

Media content enhancement based on user feedback of multiple variations

Inventors: Trisha Mittal (San Jose, CA); Viswanathan Swaminathan (Saratoga, CA); Ritwik Sinha (Cupertino, CA); Saayan Mitra (San Jose, CA); David Arbour (San Jose, CA); Somdeb Sarkhel (San Jose, CA)
Assignee: Adobe Inc.
G06Q30/0201G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,949
App. No.
17/519,311
Granted
Mar 11, 2025
Kind
B2
Abstract

Various disclosed embodiments are directed to using one or more algorithms or models to select a suitable or optimal variation, among multiple variations, of a given content item based on feedback. Such feedback guides the algorithm or model to arrive at suitable variation result such that the variation result is produced as the output for consumption by users. Further, various embodiments resolve tedious manual user input requirements and reduce computing resource consumption, among other things, as described in more detail below.

Claims (45)

1. A non-transitory computer readable medium storing computer-usable instructions that, when used by one or more processors, cause the one or more processors to perform operations comprising:

receiving a media content item that is an image of a plurality of pixel values; converting the media content item into a first feature vector representative of the plurality of pixel values in search space:

receiving an indication that a user has set a boundary or range of parameter values for which a model will generate variations from:

based on the boundary or range and the first feature vector, automatically generating a plurality of variations of the media content item by automatically changing in the search space, the first feature vector representative of a change in at least one of, vibrance, saturation, brightness, contrast, or sharpness of one or more of the plurality of pixel values, each variation, of the plurality of variations, being a different version of the image; receiving explicit user feedback for each variation of the plurality of variations, wherein the explicit user feedback corresponds to a scaled user rating of a respective variation of the plurality of variations according to an aesthetic preference for the respective variation; based on the explicit user feedback, scoring each variation of the plurality of variations according to the scaled user rating of the respective variation; based on the changing, in the search space, the first feature vector and the scoring of each variation, automatically generating, via a Bayesian Optimization Model a first variation of the image based on using a surrogate function that models an objective function representing a true distribution of user feedback for the plurality of variations by sampling, in the search space, at least a second feature vector representing the first variation based on minimizing a distance to the objective function evaluated at a maximum, wherein the first variation represents the maximum of the objective function and the maximum indicates a highest scoring variation according to the user feedback; based on the generating of the first variation, generating an output image of pixel values that represent the first variation; and based on the generating of the output image, causing

presentation, at a computing device associated with the user, of the output image.

2. The non-transitory computer readable medium of claim 1 , wherein the user feedback includes a user rating of a particular variation, of the plurality of variations, based on showing the particular variation to a second user and prompting the second user to rate the particular variation.

3. The non-transitory computer readable medium of claim 1 , wherein the generating of the first variation is based further on receiving additional user feedback includes user input at a user interface, and wherein the additional user feedback includes at least one of: a purchase of an item associated with the variation, a click of a button at a user interface page, and a view of the user interface page.

4. The non-transitory computer readable medium of claim 1 , wherein the media content item includes a video, and wherein the first variation includes a Graphics Interchange Format (GIF) of the video.

5. The non-transitory computer readable medium of claim 1 , wherein the one or more processors are caused to perform further operations comprising:

subsequent to the generating of the first variation, generating, via the model, a second set of variations that resemble the media content item; based on the generating of the second set of variations, receiving second user feedback for each variation of the second set of variations; based on the second user feedback, scoring each variation of the second set of variations; based on the scoring of the second set of variations, generating a second variation; and selecting the first variation instead of the second variation based on a score of the first variation being higher relative to the second variation, and wherein the causing presentation of the output image is further based on the selecting of the first variation.

6. The non-transitory computer readable medium of claim 1 , wherein the causing presentation includes causing the the output image to be produced at an application page as part of a digital marketing campaign.

7. The non-transitory computer readable medium of claim 1 , wherein each of the plurality of variations include a set pixels with different values relative to a corresponding set of pixels of the image.

8. The non-transitory computer readable medium of claim 1 , wherein each variation of plurality of variations include a set of pixels representing real world objects that are not included in the media content item.

9. The non-transitory computer readable medium of claim 1 , wherein the generating of the first variation is part of training or optimizing the model, and wherein the one or more processors further include operations comprising:

receiving, subsequent to the generating of the first variation, a first media content item; determining the plurality of variations; and changing first one or more image data-values of the first media content item to second one or more image data-values based on the scoring of each variation of the plurality of variations.

10. A computerized system, comprising:

at least one processor; and

at least one computer readable storage medium storing computer instructions that when executed by the at least one processor cause the at least one processor to perform operations comprising:

receiving a first media content item, the first media content item being an image or set of images that include a plurality of pixel values;

converting the first media content item into a first feature vector representative of the plurality of pixel values in search space:

generating a plurality of variations associated with the first media content item or a second media content item based at least in part on the converting and changing, in the search space, one or more values of the first feature vector, each variation, of the plurality of variations, resembling the image or set of images except each variation includes a change in at least one of, vibrance, saturation, brightness, contrast, or sharpness of one or more of the plurality of pixel values;

determining a score for each variation of the plurality of variations based on explicit user feedback for each variation, wherein the explicit user feedback corresponds to a scaled user rating of aesthetic preference for a respective variation of the plurality of variations according to one or more users;

changing first one or more pixel values of the plurality of pixel values to second one or more pixel values based on the explicit user feedback and using a surrogate function that models an objective function representing a true distribution of user feedback for the plurality of variations by sampling, in the search space, at least a second feature vector representing a first variation based on minimizing a distance to the objective function evaluated at a maximum;

based at least in part on the changing of the first one or more pixel values to the second one or more pixel values, generating an output image of pixel values that represents the first variation; and

causing presentation, at a computing device associated with a user, of the output image.

11. A computer-implemented method comprising:

receiving a media content item that is an image of a plurality of pixel values;

converting the media content item into a first feature vector representative of the plurality of pixel values in search space;

receiving an indication that a user has set a boundary or range of parameter values for which a model will generate variations from

based at least in part on the first feature vector and the boundary or range, generating a plurality of variations of media content item by changing one or more values of the first feature vector in the search space, each variation, of the plurality of variations, resembles the media content item except that each variation has one or more of, vibrance, saturation, brightness, contrast, or sharpness that are not included in the image or set of images;

receiving explicit user feedback for each variation, of the plurality of variations, wherein the explicit user feedback corresponds to a scaled user rating of aesthetic preference for a respective variation of the plurality of variations according to one or more users;

based on the explicit user feedback, scoring each variation of the plurality of variations according to the scaled user rating of the respective variation;

based on the scoring and the changing of the one or more values of the first feature vector, sampling, via the model and in the search space, a second feature vector representing a first variation of the image or set of images based at least in part on using a surrogate function that models an objective function representing a true distribution of user feedback for the plurality of variations and further based on minimizing a distance to the objective function evaluated at a maximum;

and based on the sampling of the first feature vector generating an output image of pixel values that represent the first variation.

12. The computer-implemented method of claim 11 , wherein the model is a Bayesian Optimization algorithm, that approximates a distribution of the objective function With the plurality of variations, and wherein the maximum indicating a highest scoring variant according to the explicit user feedback.

13. The computer-implemented method of claim 11 , wherein the explicit user feedback includes a user rating of a particular variation, of the plurality of variations, based on showing the particular variation to a second user and prompting the second user to rate the particular variation.

14. The computer-implemented method of claim 11 , wherein the media content item includes a video, and wherein the first variation includes a Graphics Interchange Format (GIF) of the video.

15. The computer-implemented method of claim 11 , further comprising:

in response to the generating of the first variation, determining a second set of variations of the plurality of variations, the determining of the second set of variations being part of a function that approximates a distribution associated with the user feedback;

based on the determining of the second set of variations, receiving second user feedback for each variation of the second set of variations;

based on at least a portion of the second user feedback, scoring each variation of the second set of variations;

based on the scoring of the second set of variations, determining a second variation; and

selecting the first variation instead of the second variation based on a score of the first variation being higher, and wherein the causing presentation of the first variation is further based on the selecting of the first variation.

16. The non-transitory computer readable medium of claim 10 , wherein the surrogate function that models the objective function representing a true distribution is based on using the Bayesian Optimization model that models the objective function associated with the user feedback includes generating the first variation using the model that approximates a ground truth distribution of user feedback scores, wherein the user feedback represents only a portion of the ground truth distribution of user feedback scores.

17. The computerized system of claim 10 , wherein the changing is based on using a model that has been trained or optimized, prior to the receiving of the first media content item, based on the user feedback, and wherein the plurality of variations are variations of the second media content item.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2021
From: MITTAL, TRISHA; SWAMINATHAN, VISWANATHAN; SARKHEL, SOMDEB; SINHA, RITWIK; ARBOUR, DAVID; MITRA, SAAYAN
To: ADOBE INC
Reel/Frame 058032/0597 →
Continuity (1)
Related Publication 20230139824A1 · May 4, 2023
References Cited (53)
US 11087178B2 · Naveh · 2021 [cited by examiner]
US 11537506B1 · Dasgupta · 2022 [cited by examiner]
US 20190311301A1 · Pyati · 2019 [cited by examiner]
US 20200334486A1 · Joseph · 2020 [cited by examiner]
US 20210334993A1 · Woodford · 2021 [cited by examiner]
US 20220019849A1 · Kim · 2022 [cited by examiner]
US 20220108138A1 · Kehler · 2022 [cited by examiner]
US 20220138511A1 · Xu · 2022 [cited by examiner]
KR 20200107389A · 2020 [cited by examiner]
Rudinac et al. (“Learning Crowdsourced User Preferences for Visual Summarization of Image Collections,” in IEEE Transactions on Multimedia, vol. 15, No. 6, pp. 1231-1243, Oct. 2013 (Year: 2013). [cited by examiner]
Balandat, M., et al., “Botorch: Programmable Bayesian Optimization in Pytorch”, arXiv:1910.06403v1, pp. 1-20 (Oct. 14, 2019). [cited by applicant]
Boerman, S. C., et al., “Online Behavioral Advertising: A Literature Review and Research Agenda”, Journal of Advertising, vol. 46, No. 3, pp. 363-376 (2017). [cited by applicant]
Bradski, G., and Kaehler, A., “Learning OpenCV”, O'Reilly Media, Inc., pp. 1-571 (2008). [cited by applicant]
Broder, A. Z., “Computational Advertising and Recommender Systems”, Proceedings of the ACM Conference on Recommender systems, p. 1 (2008). (Only Abstract Submitted). [cited by applicant]
Cai, S., et al., “Weakly-supervised Video Summarization using Variational Encoder-Decoder and Web Prior”, In Proceedings of the European Conference on Computer Vision (ECCV), pp. 1-17 (2018). [cited by applicant]
Crook, T., et al., “Seven Pitfalls to Avoid when Running Controlled Experiments on the Web”, In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1-9 (2009). [cited by applicant]
Dave, K., and Varma, V., “Computational Advertising: Techniques for Targeting Relevant Ads”, Foundations and Trends in Information Retrieval, vol. 8, No. 4-5, pp. 1-49 (2014). [cited by applicant]
Dixon, E., et al., “A/B Testing of a Webpage”, U.S. Pat. No. 7,975,000, pp. 1-10 (Jul. 5, 2011). [cited by applicant]
Dou, Q., et al., “Webthetics: Quantifying webpage aesthetics with deep learning”, International Journal of Human-Computer Studies, vol. 124, pp. 56-66 (2019). [cited by applicant]
Gardey, J. C., and Garrido, A., “User Experience Evaluation through Automatic A/B Testing”, In Proceedings of the 25th International Conference on Intelligent User Interfaces Companion, pp. 25-26 (Mar. 17-20, 2020). [cited by applicant]
Garnett, R., et al., “Bayesian Optimization for Sensor Set Selection”, In Proceedings of the 9th ACM/IEEE International Conference on Information Processing in Sensor Networks, pp. 1-11 (Apr. 12-16, 2010). [cited by applicant]
Gilotte, A, et al., “Offline A/B testing for Recommender Systems”, In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, arXiv:1801.07030v1, pp. 1-9 (Jan. 22, 2018). [cited by applicant]
Gu, H., and Swaminathan, V., “From Thumbnails to Summaries—A Single Deep Neural Network to Rule Them All”, In IEEE International Conference on Multimedia and Expo (ICME), pp. 1-6 (2018). [cited by applicant]
Gygli, M., and Soleymani, M., “Analyzing and Predicting GIF Interestingness”, In Proceedings of the 24th ACM International Conference on Multimedia, pp. 122-126 (Oct. 15-19, 2016). [cited by applicant]
Gygli, M., et al., “Video2GIF: Automatic Generation of Animated GIFs from Video”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1001-1009 (2016). [cited by applicant]
Hou, W., et al., “Blind Image Quality Assessment via Deep Learning”, IEEE Transactions on Neural Networks And Learning Systems, vol. 26, No. 6, pp. 1275-1286 (Jun. 2015). [cited by applicant]
Hu, Y., et al., “Exposure: A White-Box Photo Post-Processing Framework”, ACM Transactions on Graphics (TOG), vol. 37, No. 2, pp. 1-23 (2018). [cited by applicant]
Ji, Z., et al., “Video Summarization with Attention-Based Encoder-Decoder Networks”, IEEE Transactions on Circuits and Systems for Video Technology, arXiv:1708:09545v2, pp. 1-9 (Apr. 16, 2018). [cited by applicant]
Jones, D. R., et al., “Efficient Global Optimization of Expensive Black-Box Functions”, Journal of Global Optimization, vol. 13, pp. 455-492 (1998). [cited by applicant]
Kim, H., et al., “Exploiting Web Images for Video Highlight Detection With Triplet Deep Ranking”, IEEE Transactions on Multimedia, vol. 20, No. 9, pp. 2415-2426 (Sep. 2018). [cited by applicant]
Klein, A., et al., “Fast Bayesian Optimization of Machine Learning Hyperparameters on Large Datasets”, Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), pp. 1-9 (2017). [cited by applicant]
Kohavi, R., et al., “Controlled experiments on the web: survey and practical guide”, Data Mining and Knowledge Discovery, vol. 18, pp. 140-181 (2009). [cited by applicant]
Kohavi, R., et al., “Online Controlled Experiments at Large Scale”, In Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1-9 (Aug. 11-14, 2013). [cited by applicant]
Kohavi, R., et al., “Seven Rules of Thumb for Web Site Experimenters”, In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1-10 (Aug. 24-27, 2014). [cited by applicant]
Li, L., et al., “A Contextual-Bandit Approach to Personalized News Article Recommendation”, In Proceedings of the 19th International Conference on World Wide Web, pp. 661-670 (Apr. 26-30, 2010). [cited by applicant]
Martinez-Cantin, R., et al., “A Bayesian Exploration-Exploitation Approach for Optimal Online Sensing and Planning with a Visually Guided Mobile Robot”, Autonomous Robots, pp. 1-11 (Aug. 2009). [cited by applicant]
Mockus, J., “On Bayesian Methods for Seeking the Extremum”, In Optimization techniques IFIP Technical Conference, pp. 1-5 (1975). [cited by applicant]
Moran, S., et al., “DeepLPF: Deep Local Parametric Filters for Image Enhancement”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12826-12835 (2020). [cited by applicant]
Pan, H., et al., “Detection of Slow-Motion Replay Segments in Sports Video for Highlights Generation”, In IEEE International Conference on Acoustics, Speech, and Signal Processing, vol. 3, pp. 1-4 (2001). [cited by applicant]
Prabhakar, K. R., et al., “DeepFuse: A Deep Unsupervised Approach for Exposure Fusion with Extreme Exposure Image Pairs”, Proceedings of the IEEE International Conference on Computer Vision (ICCV), pp. 4714-4722 (2017). [cited by applicant]
Rahutomo, R., et al., “Improving Conversion Rates for Fashion e-Commerce with A/B Testing”, International Conference on Information Management and Technology (ICIMTech), IEEE, pp. 266-270 (Aug. 2020). [cited by applicant]
Shan, Q., et al., “The Visual Turing Test for Scene Reconstruction”, In International Conference on 3D Vision—3DV, IEEE, pp. 1-8 (2013). [cited by applicant]
Shen, Y., et al., “Deep Learning for Multimodal-Based Video Interestingness Prediction”, In EEE International Conference on Multimedia and Expo (ICME), pp. 1-7 (2017). [cited by applicant]
Shin, D., et al., “Enhancing Social Media Analysis with Visual Data Analytics: A Deep Learning Approach”, Article in MIS Quarterly, pp. 1-68 (Dec. 2020). [cited by applicant]
Sjoberg, A., et al., “Architecture-Aware Bayesian Optimization for Neural Network Tuning”, In International Conference on Artificial Neural Networks, pp. 220-231 (2019). [cited by applicant]
Talebi, H., and Milanfar, P., “Nima: Neural Image Assessment”, IEEE Transactions on Image Processing, vol. 27, No. 8, pp. 3998-4011 (Aug. 2018). [cited by applicant]
Tang, D., et al., “Overlapping Experiment Infrastructure: More, Better, Faster Experimentation”, In Proceedings of the 16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1-10 (Jul. 25-2… [cited by applicant]
Wang, S., et al., “Video Interestingness Prediction Based on Ranking Model”, In Proceedings of the Joint Workshop of the 4th Workshop on Affective Social Multimedia Computing and First Multi-Modal Affective Computing of… [cited by applicant]
Xiong, B., et al., “Less is More: Learning Highlight Detection from Video Duration”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1258-1267 (2019). [cited by applicant]
Xu, Y., et al., “From Infrastructure to Culture: A/B Testing Challenges in Large Scale Social Networks”, In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1-10 (2… [cited by applicant]
Yan, J., et al., “A Learning-to-Rank Approach for Image Color Enhancement”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-8 (2014). [cited by applicant]
Yuan, A., and Li, Y., “Modeling Human Visual Search Performance on Realistic Webpages Using Analytical and Deep Learning Methods”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, pp. 1-12 (Ap… [cited by applicant]
Zhou, K., et al., “Deep Reinforcement Learning for Unsupervised Video Summarization with Diversity-Representativeness Reward”, The Thirty-Second AAAI Conference on Artificial Intelligence, pp. 7582-7589 (2018). [cited by applicant]
Cited By (1)
US 12,488,581