IP Library › Granted Patent US 12,260,301
Granted Patent B2
US 12,260,301 · App. 17/023,766 · Granted Mar 25, 2025

Data generation and annotation for machine learning

Inventor: Mithilesh Kumar Singh (Ara, IN)
Assignee: SAP SE
G06N20/00G06T11/20G06V30/153G06T2210/12G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,301
App. No.
17/023,766
Filed
Sep 17, 2020
Granted
Mar 25, 2025
Kind
B2
Examiner
LU, HWEI-MIN
Art Unit
2142
USPC
706/12
Abstract

A data annotation server accesses a request from a machine learning server for annotated images of a user interface containing a specified user interface element. The data annotation server programmatically determines whether user interfaces generated by an application server include the specified user interface element. If so, an image of the user interface is stored and a location or bounding box of the user interface element is determined. The stored image of the user interface is annotated with the determined location of the user interface element. The image and the annotation are provided to the machine learning server, which uses the images and annotations to train a machine learning model.

Claims (52)

1. A method comprising:

receiving, by one or more processors, a request from a device for images for training a machine learning algorithm, the request specifying a user interface element and a number of images;

during a predetermined duration of time, receiving requests from other devices for user interfaces;

for each requested user interface, determining whether the requested user interface includes the specified user interface element;

for each requested user interface that is determined to include the specified user interface element, generating an image of the requested user interface that is determined to include the specified user interface element and annotation for the specified user interface element, the generating of the annotation comprising determining, by the one or more processors, coordinates of the specified user interface element in the requested user interface that is determined to include the specified user interface element; and

providing the generated images and annotation to the device for use in training the machine learning algorithm.

2. The method of claim 1 , wherein the generating of the annotation for the specified user interface element further comprises:

based on the coordinates, performing optical character recognition (OCR) on a portion of the image of the requested user interface that is determined to include the specified user interface element.

3. The method of claim 1 , wherein:

the determining of the coordinates of the specified user interface element in the requested user interface that is determined to include the specified user interface element comprises determining the coordinates using JavaScript to extract data from a Document Object Model (DOM) of the requested user interface that is determined to include the specified user interface element.

4. The method of claim 1 , wherein:

the providing of the annotation comprises providing a bounding box of the specified user interface element.

5. The method of claim 1 , wherein:

the determining that the requested user interface includes the specified user interface element comprises accessing an identifier of a type of the specified user interface element.

6. The method of claim 1 , wherein:

the determining that the requested user interface includes the specified user interface element comprises accessing an xpath.

7. A system comprising:

a memory that stores instructions; and

one or more processors configured by the instructions to perform operations comprising:

receiving a request from a device for images for training a machine learning algorithm, the request specifying a user interface element and a number of images;

during a predetermined duration of time, receiving requests from other devices for user interfaces;

for each requested user interface, determining whether the requested user interface includes the specified user interface element;

for each requested user interface that is determined to include the specified user interface element, generating an image of the requested user interface that is determined to include the specified user interface element and annotation for the specified user interface element, the generating of the annotation comprising determining, by the one or more processors, coordinates of the specified user interface element in the requested user interface that is determined to include the specified user interface element; and

providing the generated images and annotation to the device for use in training the machine learning algorithm.

8. The system of claim 7 , wherein the generating of the annotation for the specified user interface element further comprises:

based on the coordinates, performing optical character recognition (OCR) on a portion of the image of the requested user interface that is determined to include the specified user interface element.

9. The system of claim 7 , wherein:

the providing of the annotation comprises providing a bounding box of the specified user interface element.

10. The system of claim 7 , wherein the determining of the coordinates of the specified user interface element in the requested user interface that is determined to include the specified user interface element comprises determining the coordinates using JavaScript to extract data from a Document Object Model (DOM) of the requested user interface that is determined to include the specified user interface element.

11. The system of claim 7 , wherein:

the determined coordinates for each provided image are provided in an extended Markup Language (XML) file that corresponds to the provided image.

12. The system of claim 11 , wherein:

a prefix of the XML file is same as a prefix of a file containing the provided image.

13. The system of claim 7 , wherein:

the determining that the requested user interface includes the specified user interface element comprises accessing an xpath.

14. A non-transitory machine-readable medium that stores instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving a request from a device for images for training a machine learning algorithm, the request specifying a user interface element and a number of images;

during a predetermined duration of time, receiving requests from other devices for user interfaces;

for each requested user interface, determining whether the requested user interface includes the specified user interface element;

for each requested user interface that is determined to include the specified user interface element, generating an image of the requested user interface that is determined to include the specified user interface element and annotation for the specified user interface element, the generating of the annotation comprising determining, by the one or more processors, coordinates of the specified user interface element in the requested user interface that is determined to include the specified user interface element; and

providing the generated images and annotation to the device for use in training the machine learning algorithm.

15. The non-transitory machine-readable medium of claim 14 , wherein the determining of the coordinates of the specified user interface element in the requested user interface that is determined to include the specified user interface element comprises determining the coordinates using JavaScript to extract data from a Document Object Model (DOM) of the requested user interface that is determined to include the specified user interface element.

16. The non-transitory machine-readable medium of claim 14 , wherein:

the determined coordinates for each provided image are provided in an extended Markup Language (XML) file that corresponds to the provided image.

17. The non-transitory machine-readable medium of claim 16 , wherein:

a prefix of the XML file is same as a prefix of a file containing the provided image.

18. The non-transitory machine-readable medium of claim 14 , wherein the operations further comprise:

training the machine learning algorithm using the provided images and determined coordinates.

19. The non-transitory machine-readable medium of claim 14 , wherein the operations further comprise:

generating a portion of the requested images by modifying copies of images of requested interfaces.

20. The non-transitory machine-readable medium of claim 14 , wherein:

the determining that the requested user interface includes the specified user interface element comprises accessing an xpath.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2020
From: SINGH, MITHILESH K
To: SAP SE
Reel/Frame 053802/0297 →
Continuity (1)
Related Publication 20220083907A1 · Mar 17, 2022
References Cited (28)
US 10474564B1 · Konyshev · 2019 [cited by examiner]
US 10719301B1 · Dasgupta · 2020 [cited by examiner]
US 10846106B1 · Curic · 2020 [cited by examiner]
US 10970530B1 · Adam · 2021 [cited by examiner]
US 11176443B1 · Selva · 2021 [cited by examiner]
US 20050055633A1 · Ali · 2005 [cited by examiner]
US 20080195958A1 · Detiege · 2008 [cited by examiner]
US 20130124480A1 · Chua · 2013 [cited by examiner]
US 20140019843A1 · Schmidt · 2014 [cited by applicant]
US 20140337705A1 · Glover et al. · 2014 [cited by applicant]
US 20150220312A1 · Jemiolo · 2015 [cited by examiner]
US 20160283222A1 · Yaros · 2016 [cited by examiner]
US 20180203674A1 · Dayanandan · 2018 [cited by examiner]
US 20190087691A1 · Jelveh · 2019 [cited by examiner]
US 20190250891A1 · Kumar · 2019 [cited by examiner]
US 20200019418A1 · P K · 2020 [cited by examiner]
US 20200133644A1 · Hou · 2020 [cited by examiner]
US 20200167153A1 · Wang · 2020 [cited by examiner]
US 20200234025A1 · Cohen · 2020 [cited by examiner]
US 20200242154A1 · Haneda · 2020 [cited by examiner]
US 20210073977A1 · Carter · 2021 [cited by examiner]
US 20210192394A1 · McKay · 2021 [cited by examiner]
US 20210349587A1 · Bigham · 2021 [cited by examiner]
Zhou et al., “Applying machine learning to automated information graphics generation”, IBM Systems Journal, vol. 41, No. 3, 2002, pp. 504-523. (Year: 2002). [cited by examiner]
Hetherington et al., “Embodying and Extracting Data in Web3D Models of Proposed Building Developments”, Computer Graphics, Imaging and Visualisation (CGIV 2007), Aug. 2007, pp. 528-534. (Year: 2007). [cited by examiner]
Moran, et al., “Machine Learning-Based Prototyping of Graphical User Interfaces for Mobile Apps”, arXiv Article: 1802.02312, Feb. 7, 2018. (Year: 2018). [cited by examiner]
Moses, Olafenwa, “Building your private Cloud AI API”, [Online]. Retrieved from the Internet: <URL: https://medium.com/deepquestai/building-your-private-cloud-ai-api-50e93a83a6ce>, (Jun. 1, 2019), 10 pgs. [cited by applicant]
Moses, Olafenwa, “Object Detection Training—Preparing your custom dataset”, [Online]. Retrieved from the Internet: <URL: https://medium.com/deepquestai/object-detection-training-preparing-your-custom-dataset-6248679f0d1… [cited by applicant]