IP Library › Granted Patent US 12,579,220
Granted Patent B2
US 12,579,220 · App. 17/480,869 · Granted Mar 17, 2026

Visual attribute expansion via multiple machine learning models

Inventors: Pramod Kumar Sharma (Seattle, WA); Yijian Xiang (Redmond, WA); Yiran Li (Bunaby, CA); Paul Pangilinan Del Villar (Bothell, WA); Liang Du (Redmond, WA); Robin Abraham (Redmond, WA); Nilgoon Zarei (Kirkland, WA); Mandar Dilip Dixit (Seattle, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F18/2431G06F18/213G06N3/044G06T7/11G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,220
App. No.
17/480,869
Granted
Mar 17, 2026
Kind
B2
Abstract

A computer implemented method includes receiving an image that includes a type of object, segmenting the object into multiple segments via a trained segmentation machine learning model, and inputting the segments into multiple different attribute extraction models to extract different types of attributes from each of the multiple segments.

Claims (31)

1 . A computer implemented method comprising:

receiving an image that includes a type of object;

segmenting the image of the object into multiple segments, including multiple body part shape segments, via a trained segmentation machine learning model; and

inputting the segments into multiple different attribute extraction models to extract different types of attributes from each of the multiple segments.

2 . The method of claim 1 wherein the trained segmentation machine learning model has been trained on multiple images of the type of object that are labeled with multiple tagged segments identified by bounding boxes.

3 . The method of claim 2 wherein the type of object comprises a bottle and wherein tags comprise classes including bottle, neck, shoulder, body, top, and logo.

4 . The method of claim 3 wherein the top class includes multiple different top classes and wherein the logo class includes multiple different logo classes.

5 . The method of claim 1 wherein the trained segmentation machine learning model comprises a mask-recurrent convolutional neural network having classes corresponding to the multiple segments of the object.

6 . The method of claim 1 wherein the attribute extraction models include a shape attribute extraction model and a color attribute extraction model.

7 . The method of claim 6 wherein the color attribute extraction model generates image embedding on each salient region of the image and determines distance differences between the image embeddings and target color text embeddings in the same embedding space, to output target colors.

8 . The method of claim 7 wherein the color attribute extraction model is an unsupervised model and wherein a distance difference is compared to a threshold to determine that a color is present in a salient region corresponding to one of the segments.

9 . The method of claim 7 wherein the color attribute extraction model generates a list of colors in descending order of area covered by each color in a salient region.

10 . The method of claim 6 wherein the shape attribute model utilizes ratios of measurements of bounding boxes corresponding to segments to generate shape descriptions.

11 . The method of claim 6 wherein the attribute extraction models include a design elements attribute extraction model.

12 . The method of claim 11 wherein the design elements attribute extraction model is an unsupervised model that compares image embeddings to design element text embeddings in the same embedding space to select design element themes.

13 . The method of claim 12 wherein the type of object comprises a bottle and wherein the design elements capture a main design theme of a shape of the bottle as function of a highest score for each design theme.

14 . A machine-readable storage device having instructions for execution by a processor of a machine to cause the processor to perform operations to perform a method, the operations comprising:

receiving an image that includes a type of object;

segmenting the image of the object into multiple segments, including multiple body part shape segments, via a trained segmentation machine learning model; and

inputting the segments into multiple different attribute extraction models to extract different types of attributes from each of the multiple segments.

15 . The device of claim 14 wherein the trained segmentation machine learning model has been trained on multiple images of the type of object that are labeled with multiple tagged segments identified by bounding boxes.

16 . The device of claim 15 wherein the trained segmentation machine learning model comprises a mask-recurrent convolutional neural network having classes corresponding to the multiple segments of the object.

17 . The device of claim 15 wherein the type of object comprises a bottle and wherein tags comprise classes including bottle, neck, shoulder, body, top, and logo, wherein the top class includes multiple different top classes and wherein the logo class includes multiple different logo classes, and wherein the attribute extraction models include a shape attribute extraction model and a color attribute extraction model.

18 . The device of claim 14 wherein the attribute extraction models include a shape attribute extraction model and a color attribute extraction model, wherein the color attribute extraction model generates image embedding on each segment of the image and determines distance differences between the image embeddings and target color text embeddings in the same embedding space, to output target colors.

19 . The device of claim 18 wherein the color attribute extraction model is an unsupervised model and wherein a distance difference is compared to a threshold to determine that a color is present in a salient region corresponding to one of the segments and wherein the shape attribute model utilizes ratios of measurements of bonding boxes corresponding to segments to generate shape descriptions.

20 . A device comprising:

a processor; and

a memory device coupled to the processor and having a program stored thereon for execution by the processor to perform operations comprising:

receiving an image that includes a type of object;

segmenting the image of the object into multiple segments, including multiple body part shape segments, via a trained segmentation machine learning model; and

inputting the segments into multiple different attribute extraction models to extract different types of attributes from each of the multiple segments.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2021
From: SHARMA, PRAMOD KUMAR; XIANG, YIJIAN; LI, YIRAN; DEL VILLAR, PAUL PANGILINAN; DU, LIANG; ABRAHAM, ROBIN; ZAREI, NILGOON; DIXIT, MANDAR DILIP
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 057636/0968 →
Continuity (1)
Related Publication 20230088925A1 · Mar 23, 2023
References Cited (54)
US 10489683B1 · Koh · 2019 [cited by examiner]
US 10747807B1 · Garg · 2020 [cited by examiner]
US 11087742B1 · Levy · 2021 [cited by examiner]
US 11282221B1 · Yang · 2022 [cited by examiner]
US 11551096B1 · Mishra · 2023 [cited by examiner]
US 11887217B2 · Maheshwari · 2024 [cited by examiner]
US 20030108237A1 · Hirata · 2003 [cited by examiner]
US 20160019587A1 · Hueter · 2016 [cited by examiner]
US 20170301085A1 · Riklin Raviv · 2017 [cited by examiner]
US 20180150444A1 · Kasina · 2018 [cited by examiner]
US 20180204088A1 · Chen · 2018 [cited by examiner]
US 20190035431A1 · Attorre · 2019 [cited by examiner]
US 20190388005A1 · Averbuch · 2019 [cited by examiner]
US 20200082222A1 · Cohen · 2020 [cited by examiner]
US 20200145623A1 · Sadanand · 2020 [cited by examiner]
US 20200193591A1 · Kamiyama · 2020 [cited by examiner]
US 20200211692A1 · Kalafut · 2020 [cited by examiner]
US 20200302117A1 · Hu · 2020 [cited by examiner]
US 20200401835A1 · Zhao · 2020 [cited by examiner]
US 20210004577A1 · Amat Roldan · 2021 [cited by examiner]
US 20210027083A1 · Cohen · 2021 [cited by examiner]
US 20210027448A1 · Cohen · 2021 [cited by examiner]
US 20210027471A1 · Cohen · 2021 [cited by examiner]
US 20210103776A1 · Jiang · 2021 [cited by examiner]
US 20210192552A1 · Gugnani · 2021 [cited by examiner]
US 20210263962A1 · Chang · 2021 [cited by examiner]
US 20210340857A1 · Mohamed Shibly · 2021 [cited by examiner]
US 20220215201A1 · Dwivedi · 2022 [cited by examiner]
US 20220309275A1 · Kirsten · 2022 [cited by examiner]
US 20220383037A1 · Pham · 2022 [cited by examiner]
US 20230088925A1 · Sharma · 2023 [cited by examiner]
US 20230185839A1 · Frei · 2023 [cited by examiner]
US 20230274404A1 · George · 2023 [cited by examiner]
US 20250005819A1 · Xu · 2025 [cited by examiner]
US 20250147639A1 · Hoffer · 2025 [cited by examiner]
CN 110060239A · 2019 [cited by applicant]
CN 112949655A · 2021 [cited by applicant]
WO 2020026763A1 · 2020 [cited by applicant]
WO 2020129066A1 · 2020 [cited by applicant]
WO WO2021242551A1 · 2021 [cited by examiner]
Lin et al. “PAM: Understanding Product Images in Cross Product Category Attribute Extraction” ADS Track Paper, KDD 2021, Aug. 14-18, 2021, pp. 3262-3270 (Year: 2021). [cited by examiner]
Logan et al. “Multimodal Attribute Extraction” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, Ca. pp. 1-7 (Year: 2017). [cited by examiner]
Zhou et al. “Simplified DOM Trees for Transferable Attribute Extraction from the Web” IW3C2 Apr. 19-23, 2021, pp. 1-10 (Year: 2021). [cited by examiner]
“Anatomy of a Bottle”, Retrieved from: https://www.berlinpackaging.com/anatomy-of-a-bottle/, Oct. 20, 2019, 2 Pages. [cited by applicant]
“PASCAL-Part Dataset”, Retrieved from: https://web.archive.org/web/20200610005259/https:/www.cs.stanford.edu˜roozbeh/pascal-parts/pascal-parts.html, Jun. 10, 2020, 3 Pages. [cited by applicant]
Chen, et al., “Understanding the Impact of Label Granularity on CNN-based Image Classification”, In Repository of arXiv:1901.07012v1, Jan. 21, 2019, pp. 1-10. [cited by applicant]
Chitturi, et al., “The Influence of Color and Shape of Package Design on Consumer Preference: The Case of Orange Juice”, In International Journal of Innovation and Economic Development, vol. 5, Issue 2, Jun. 2019, pp. 4… [cited by applicant]
Diamond, Catherine, “Beverage Labeling”, Retrieved from: https://www.labelandnarrowweb.com/issues/2013-07/view_features/beverage-labeling/, Jul. 17, 2013, 11 Pages. [cited by applicant]
Na, et al., “Automatic Segmentation of Product Bottle Label Based on GrabCut Algorithm”, In International Journal of Contents, vol. 10, Issue 4, Dec. 2014, pp. 1-10. [cited by applicant]
Schoormans, et al., “Designing Packages that Communicate Product Attributes and Brand Values: An Exploratory Method”, In the Journal of Design, vol. 13, Issue 1, Mar. 2010, 19 Pages. [cited by applicant]
Wu, et al., “Detectron2 Model Zoo and Baselines”, Retrieved from: https://github.com/facebookresearch/detectron2/blob/master/MODEL_ZOO.md, Jul. 21, 2021, 9 Pages. [cited by applicant]
Communication under Rule 71(3) EPC, Received in European Patent Application No. 22761711.5, mailed on Oct. 10, 2025, 08 pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/038991”, Mailed Date: Nov. 17, 2022, 10 Pages. [cited by applicant]
Wang, et al., “Joint Object and Part Segmentation Using Deep Learned Potentials”, In Proceedings of IEEE International Conference on Computer Vision, Dec. 7, 2015, pp. 1573-1581. [cited by applicant]