IP Library Granted Patent US 12,469,262
Granted Patent B2
US 12,469,262 · App. 18/633,170 · Granted Nov 11, 2025

Subjective quality assessment tool for image/video artifacts

Inventors: Yuanyi Xue (Alameda, CA); Scott Labrozzi (Cary, NC); Wenhao Zhang (Beijing, CN); Christopher Richard Schroers (Uster, CH); Roberto Gerson De Albuquerque Azevedo (Zurich, CH); Xuchang Huangfu (Beijing, CN); Lemei Huang (Beijing, CN); Yang Zhang (Dübendorf, CH)
Assignees: Disney Enterprises, Inc.; Beijing YoJaJa Software Technology Development Co., Ltd.
G06V10/774G06T7/0002G06V10/26G06T2207/10016G06T2207/20081G06T2207/30168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,262
App. No.
18/633,170
Granted
Nov 11, 2025
Kind
B2
Abstract

In some embodiments, a method sends information for a sample of content, a first question, and a second question for output on an interface. The first question receives, from a subject, a first response for a sample level rating for an artifact that is perceived to be visible in the sample and the second question receives, from the subject, a second response for regions in the sample that are perceived to contain the artifact. The method receives the first response for the sample level rating and the second response for regions that are perceived to contain the artifact. First responses are combined from multiple subjects to generate an opinion score for the sample and second responses are combined to generate region scores for regions. The method generates training data from the opinion score and the region scores to train a process to perform an action based on the artifacts.

Claims (68)

1 . A method comprising:

sending information for a sample of content, a first question, and a second question for output on an interface, wherein the first question is configured to receive, from a subject, a first response for a sample level rating for an artifact that is perceived to be visible in the sample of content and the second question is configured to receive, from the subject, a second response for one or more regions in a plurality of regions in the sample of content that are perceived to contain the artifact;

receiving the first response for the sample level rating and the second response for one or more regions that are perceived to contain the artifact;

combining first responses for the first question from multiple subjects to generate an opinion score for the sample of content and combining second responses for the second question from the multiple subjects to generate region scores for regions in the plurality of regions; and

generating training data from the opinion score and the region scores to train a process to perform an action based on the artifacts in one or more regions in the sample of content, wherein generating training data comprises:

determining a cropped patch of a portion of the sample of content;

determining regions that are included in the cropped patch; and

determining a patch score based on region scores for the regions that are included in the cropped patch.

2 . The method of claim 1 , further comprising:

determining a dataset for a subjective assessment to be performed by the subject; and

retrieving the sample of content from the dataset.

3 . The method of claim 2 , further comprising:

sending multiple samples of content from the dataset for output on the interface;

receiving first responses and second responses for respective samples of content in the multiple samples of content; and

generating training data for the samples of content using the respective first responses and second responses.

4 . The method of claim 1 , wherein:

the first response comprises a value from a range for the sample level rating, and

the second response comprises identifiers for the one or more regions that are selected.

5 . The method of claim 1 , wherein combining the first responses from multiple subjects comprises:

generating a mean of the sample level ratings from the first responses.

6 . The method of claim 1 , wherein combining the second responses from multiple subjects comprises:

generating a value based on a number of subjects that select respective regions.

7 . The method of claim 6 , wherein the value is a percentage of subjects that selected respective regions.

8 . The method of claim 1 , wherein the region scores are represented in a heat map that displays the respective region scores in association with the regions.

9 . The method of claim 1 , wherein generating training data comprises:

associating the region scores for regions in the sample of content with the opinion score.

10 . The method of claim 1 , further comprising:

using the patch score to train the process.

11 . The method of claim 1 , wherein determining the patch score comprises:

using a weighted average of the region scores for the regions that are included in the cropped patch based on a proportion of area associated with each of the regions in the cropped patch.

12 . The method of claim 1 , wherein determining the patch score comprises:

selecting one of the region scores to be the patch score.

13 . The method of claim 1 , wherein determining the patch score comprises:

determining a proportion of area for a region that has a higher region score, wherein the higher region score indicates more artifacts were perceived to be visible;

when the proportion of area meets a threshold, selecting the patch score based a region score for the region that has the higher region score, and

when the proportion of area does not meet the threshold, selecting the patch score based on a region score for a region with a lower region score.

14 . The method of claim 1 , wherein determining the patch score comprises:

determining a proportion of area for a plurality of regions;

when a proportion of area for a first region that has a highest region score meets a threshold, selecting the patch score based the region score for the region that has the highest region score, wherein the higher region score indicates more artifacts were perceived to be visible,

when the proportion of area for two regions that have a highest region score meets the threshold, selecting the patch score based on region scores for two regions,

when the proportion of area for three regions that have a highest region score meets the threshold, selecting the patch score based on region scores for three regions, and

when the proportion of area for three regions that have the highest region score does not meet the threshold, selecting the patch score based on a region score for a region with a lowest region score, wherein the lowest region score indicates less artifacts were perceived to be visible.

15 . A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for:

sending information for a sample of content, a first question, and a second question for output on an interface, wherein the first question is configured to receive, from a subject, a first response for a sample level rating for an artifact that is perceived to be visible in the sample of content and the second question is configured to receive, from the subject, a second response for one or more regions in a plurality of regions in the sample of content that are perceived to contain the artifact;

receiving the first response for the sample level rating and the second response for one or more regions that are perceived to contain the artifact;

combining first responses for the first question from multiple subjects to generate an opinion score for the sample of content and combining second responses for the second question from the multiple subjects to generate region scores for regions in the plurality of regions; and

generating training data from the opinion score and the region scores to train a process to perform an action based on the artifacts in one or more regions in the sample of content, wherein generating training data comprises:

determining a cropped patch of a portion of the sample of content;

determining regions that are included in the cropped patch, and

determining a patch score based on region scores for the regions that are included in the cropped patch.

16 . A method comprising:

outputting a sample of content, a first question, and a second question on an interface;

receiving, from a subject, a first response to the first question for a sample level rating for an artifact that is perceived to be visible in the sample of content that is output on the interface;

receiving, from the subject, a second response to the second question for one or more regions in a plurality of regions in the sample of content that are perceived to contain the artifact; and

sending the first response and the second response to a server system, wherein first responses from multiple subjects are combined to generate an opinion score for the sample of content from the first responses and second responses from the multiple subjects are combined to generate region scores for regions in the plurality of regions, and training data is generated from the opinion score and the region scores to train a process to perform an action based on the artifacts in one or more regions in the sample of content, wherein generating training data comprises:

determining a cropped patch of a portion of the sample of content:

determining regions that are included in the cropped patch; and

determining a patch score based on region scores for the regions that are included in the cropped patch.

17 . The method of claim 16 , wherein:

the interface displays the sample and a window that lists information for the first question and the second question,

the first question lists a range of values for the sample level rating, and

the second question displays the plurality regions associated with the sample that can be selected as including artifacts.

18 . The method of claim 17 , wherein receiving the second response comprises:

receiving a value from the range of values; and

receiving selections of one or more regions in the plurality of regions.

19 . The method of claim 16 , wherein:

the interface highlights regions in the sample that are selected for the second question to allow the subject to view the artifacts in the region on the sample.

20 . The method of claim 16 , wherein the region scores are represented in a heat map that displays the respective region scores in association with the regions.

Assignments (5)
CHANGE OF NAME Recorded Sep 24, 2024
From: BEIJING HULU SOFTWARE TECHNOLOGY DEVELOPMENT CO., LTD.
To: BEIJING YOJAJA SOFTWARE TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 068684/0455 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2024
From: LABROZZI, SCOTT; XUE, YUANYI
To: DISNEY ENTERPRISES, INC.
Reel/Frame 067081/0958 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2024
From: ZHANG, WENHAO; HUANGFU, XUCHANG; HUANG, LEMEI
To: BEIJING HULU SOFTWARE TECHNOLOGY DEVELOPMENT CO., LTD.
Reel/Frame 067082/0024 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2024
From: SCHROERS, CHRISTOPHER RICHARD; DE ALBUQUERQUE AZEVEDO, ROBERTO GERSON; ZHANG, YANG
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 067082/0505 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 11, 2024
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 067082/0542 →
Continuity (2)
Provisional Application 63498642 · Apr 27, 2023
Related Publication 20240362896A1 · Oct 31, 2024
References Cited (51)
US 8532198B2 · Kumwilaisak et al. · 2013 [cited by applicant]
US 10949604B1 · Dwivedi et al. · 2021 [cited by applicant]
US 20070103551A1 · Kim et al. · 2007 [cited by applicant]
US 20100135575A1 · Guo · 2010 [cited by examiner]
US 20100309985A1 · Liu · 2010 [cited by examiner]
US 20130093768A1 · Lockerman et al. · 2013 [cited by applicant]
US 20170078706A1 · Van Der Vleuten · 2017 [cited by examiner]
US 20180034852A1 · Goldenberg · 2018 [cited by applicant]
US 20180167620A1 · Li et al. · 2018 [cited by applicant]
US 20190156459A1 · Chen · 2019 [cited by applicant]
US 20190261016A1 · Liu et al. · 2019 [cited by applicant]
US 20190340468A1 · Stumpe · 2019 [cited by examiner]
US 20200352518A1 · Lyman · 2020 [cited by examiner]
US 20220414402A1 · Sawkey · 2022 [cited by examiner]
US 20230098732A1 · Alemi · 2023 [cited by examiner]
US 20230131228A1 · Wang · 2023 [cited by examiner]
US 20230187072A1 · Neumann · 2023 [cited by examiner]
US 20230274818A1 · Etemadi · 2023 [cited by examiner]
US 20230282012A1 · Borges · 2023 [cited by applicant]
US 20240296535A1 · Bakunov · 2024 [cited by examiner]
CN 105551062A · 2016 [cited by applicant]
CN 109376731A · 2019 [cited by applicant]
CN 114170198A · 2022 [cited by applicant]
CN 117935180A · 2024 [cited by applicant]
EP 4456539A3 · 2024 [cited by applicant]
JP 4527127B · 2010 [cited by applicant]
WO 2019125026A1 · 2019 [cited by applicant]
WO 2023235730A1 · 2023 [cited by applicant]
Ying et al. (“Patch-VQ: ‘Patching Up’ the Video Quality Problem”; Jun. 20, 2021; IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE; pp. 14014-14024, XP034008989, DOI: 10.1109/CVPR46437.2021.013… [cited by examiner]
“SSIMWAVE”, IMAX Streaming and Consumer Technology (IMAX SCT), Retrieval date: Apr. 10, 2024. Retrieved from internet: https://www.imax.com/sct/product/streamsmart-on-demand. [cited by applicant]
“FFmpeg”, FFmpeg Developers, Retrieval date: Apr. 10, 2024. Retrieved from internet: https://ffmpeg.org/. [cited by applicant]
Campbell, Fergus W., and John G. Robson. “Application of Fourier analysis to the visibility of gratings.” The Journal of physiology 197, No. 3 (1968): 551. [cited by applicant]
ITU-T, Recommandation. “BT.500-14 Methodology For the Subjective Assessment of the Quality of Television Pictures.” International Telecommunication Union, Geneva. 2019. [cited by applicant]
ITU-T, Recommandation. “P910 Subjective video quality assessment methods for multimedia applications.” International Telecommunication Union, Geneva. 2022. [cited by applicant]
ITU-T, Recommendation. “P911 Subjective Audiovisual Quality Assessment Methods for Multimedia Applications.” International Telecommunication Union, Geneva. 1998. [cited by applicant]
Kapoor, Akshay, Jatin Sapra, and Zhou Wang. “Capturing banding in images: Database construction and objective assessment.” In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech andSignal Processing (ICA… [cited by applicant]
Mittal, Anish, Anush Krishna Moorthy, and Alan Conrad Bovik. “Noreference image quality assessment in the spatial domain.” IEEE Transactions on image processing 21, No. 12 (2012): 4695-4708. [cited by applicant]
Tandon, Pulkit, Mariana Afonso, Joel Sole, and Lukáš Krasula, “CAMBI: Contrast-aware multiscale banding index.” In 2021 Picture Coding Symposium (PCS), pp. 1-5. IEEE, 2021. [cited by applicant]
Tu, Zhengzhong, Jessie Lin, Yilin Wang, Balu Adsumilli, and Alan C. Bovik. “Adaptive debanding filter.” IEEE Signal Processing Letters 27 (2020): 1715-1719. [cited by applicant]
Extended European Search Report for EP App No. 24170914.6, dated Sep. 30, 2024, 16 pgs. [cited by applicant]
Tandon Pulkit et al: “CAMBI: Contrast-aware Multiscale Banding Index”, 2021 Picture Coding Symposium (PCS), IEEE, Jun. 29, 2021 (Jun. 29, 2021), pp. 1-5, XP033945096,DOI: 10.1109/PCS50896.2021.9477464 [retrieved on Jul.… [cited by applicant]
Testolina Michela et al: “Review of subjective quality assessment methodologies and standards for compressed images evaluation”, Proceedings of the SPIE, SPIE, US, val. 11842, Aug. 1, 2021 (Aug. 1, 2021), pp. 118420Y-11… [cited by applicant]
Tu Zhengzhong et al: “Bband Index: a No-Reference Banding Artifact Predictor”, ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, A,P May 4, 2020 (May 4, 2020), pp.… [cited by applicant]
Xiang Jie et al: “A Deep Learning-Based No-Reference Quality Metric for High-Definition Images Compressed With HEVC” I IEEE Transactions on Broadcasting, IEEE Service Center, Piscataway, NJ, US, val. 69, No. 3, Sep. 1, … [cited by applicant]
Xue Yuanyi et al: “Large-Scale Multi-Site 1-15 Subjective Assessment on Image Banding Artifacts”, 2023 15th International Conference on Quality of Multimedia Experience {QOMEX), IEEE, Jun. 20, 2023 (Jun. 20, 2023), pp. … [cited by applicant]
Ying Zhenqiang et al: “Patch-VQ: ‘Patching Up’ the Video Quality Problem”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE,Jun. 20, 2021 (Jun. 20, 2021), pp. 14014-14024, XP034008989, DO… [cited by applicant]
U.S. Appl. No. 18/653,592, filed Apr. 11, 2024, Inventor Xuchang Huangfu et al, Titled: “Banding Artifact Assessment Method”, Accessible via Patent Center. [cited by applicant]
Madhusudana, Pavan C. et al, “Image Quality Assessment Using Contrastive Learning.” IEEE Transactions on Image Processing, Oct. 25, 2021, 10 pages. [cited by applicant]
Mingyang Song et al., “A Generative Model for Digital Camera Noise Synthesis”, ETH Zurich, Switzerland; Disney Research Studios, Mar. 17, 2023, 18 pages. [cited by applicant]
Chen Zijian et al: “BAND-2k: Banding 1-15 INV. Artifact Noticeable Database for Banding G06T7/40 Detection and Quality Assessment”, IEEE Transactions on Circuits and Systems for Video Technology, IEEE, USA, vol. 34, No.… [cited by applicant]
Extended European Search Report, EP Application No. 25157981.9, mailed Jul. 1, 2025, 10 pages. [cited by applicant]