IP Library › Granted Patent US 12,242,824
Granted Patent B2
US 12,242,824 · App. 18/046,446 · Granted Mar 4, 2025

Creating user interface using machine learning

Inventors: Zifeng Huang (Emeryville, CA); Yang Li (Palo Alto, CA); Gang Li (Mountain View, CA); Xin Zhou (Mountain View, CA); John Francis Canny (Berkeley, CA)
Assignee: Google LLC
G06F8/38G06F3/0484G06F8/33G06F40/40G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,824
App. No.
18/046,446
Granted
Mar 4, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training and using machine learning models to generate graphical user interfaces from textual descriptions.

Claims (53)

1. A computer-implemented method, comprising:

receiving a training dataset comprising a plurality of training samples, each training sample comprising:

a graphical user interface that includes a plurality of graphical elements; and

a natural language textual description comprising a single phrase that describes the graphical user interface that includes the plurality of graphical elements;

generating, for each graphical user interface, graphical attribute data that describes, for each graphical element of the graphical user interface, an attribute type of the graphical element, and a position of the graphical element;

generating, for each natural language textual description, using a pre-trained word embedding model, an embedding vector of the natural language textual description; and

training a machine learning model, based on the graphical attribute data and the embedding vector for each training sample, to generate, as output, prediction data that is indicative of graphical elements in a graphical user interface, wherein the machine learning model comprises a transformer based model that includes:

an encoder that receives the embedding vector of the natural language textual description and processes the embedding vector to generate an output vector; and

a decoder that receives the output vector and the graphical attribute data generated from the training sample and is trained to generate the prediction data.

2. The computer-implemented method of claim 1 , wherein the prediction data comprises, for each of a plurality of graphical elements, a probability distribution of the graphical element being included in a graphical user interface, and a probability distribution of the positon of the graphical element in the graphical user interface.

3. The computer-implemented method of claim 1 , wherein the prediction data comprises, a plurality of graphical elements, and for each of the graphical elements, a positon of the graphical element in the graphical user interface.

4. The computer-implemented method of claim 1 , wherein the machine learning model comprises a transformer based model that includes an encoder that processes the embedding vector of the natural language textual description, and a decoder that processes the graphical attribute data.

5. The computer-implemented method of claim 1 , further comprising:

providing a natural language textual description of a graphical user interface as input to the machine learned model;

generating, by the machine learned model, based on the natural language textual description of graphical user interface prediction data that is indicative of graphical elements in a graphical user interface;

providing the prediction data to a graphical user interface renderer; and

generating, by the renderer, a graphical user interface based on the prediction data.

6. A computer storage medium encoded with a computer program, the program comprising instructions that when executed by data processing apparatus cause the data processing apparatus to perform the operations of:

receiving a training dataset comprising a plurality of training samples, each training sample comprising:

a graphical user interface that includes a plurality of graphical elements; and

a natural language textual description comprising a single phrase that describes the graphical user interface that includes the plurality of graphical elements;

generating, for each graphical user interface, graphical attribute data that describes, for each graphical element of the graphical user interface, an attribute type of the graphical element, and a position of the graphical element;

generating, for each natural language textual description, using a pre-trained word embedding model, an embedding vector of the natural language textual description; and

training a machine learning model, based on the graphical attribute data and the embedding vector for each training sample, to generate, as output, prediction data that is indicative of graphical elements in a graphical user interface, wherein the machine learning model comprises a transformer based model that includes:

an encoder that receives the embedding vector of the natural language textual description and processes the embedding vector to generate an output vector; and

a decoder that receives the output vector and the graphical attribute data generated from the training sample and is trained to generate the prediction data.

7. The computer storage medium of claim 6 , wherein the prediction data comprises, for each of a plurality of graphical elements, a probability distribution of the graphical element being included in a graphical user interface, and a probability distribution of the position of the graphical element in the graphical user interface.

8. The computer storage medium of claim 6 , wherein the prediction data comprises, a plurality of graphical elements, and for each of the graphical elements, a position of the graphical element in the graphical user interface.

9. The computer storage medium of claim 6 , wherein the machine learning model comprises a transformer based model that includes an encoder that processes the embedding vector of the natural language textual description, and a decoder that processes the graphical attribute data.

10. The computer storage medium of claim 6 , the operations further comprising:

providing a natural language textual description of a graphical user interface as input to the machine learned model;

generating, by the machine learned model, based on the natural language textual description of graphical user interface prediction data that is indicative of graphical elements in a graphical user interface;

providing the prediction data to a graphical user interface renderer; and

generating, by the renderer, a graphical user interface based on the prediction data.

11. A system, comprising:

a data processing apparatus; and

a computer storage medium encoded with a computer program, the program comprising instructions that when executed by the data processing apparatus cause the data processing apparatus to perform the operations of:

receiving a training dataset comprising a plurality of training samples, each training sample comprising:

a graphical user interface that includes a plurality of graphical elements; and

a natural language textual description comprising a single phrase that describes the graphical user interface that includes the plurality of graphical elements;

generating, for each graphical user interface, graphical attribute data that describes, for each graphical element of the graphical user interface, an attribute type of the graphical element, and a position of the graphical element;

generating, for each natural language textual description, using a pre-trained word embedding model, an embedding vector of the natural language textual description; and

training a machine learning model, based on the graphical attribute data and the embedding vector for each training sample, to generate, as output, prediction data that is indicative of graphical elements in a graphical user interface, wherein the machine learning model comprises a transformer based model that includes:

an encoder that receives the embedding vector of the natural language textual description and processes the embedding vector to generate an output vector; and

a decoder that receives the output vector and the graphical attribute data generated from the training sample and is trained to generate the prediction data.

12. The system of claim 11 , wherein the prediction data comprises, for each of a plurality of graphical elements, a probability distribution of the graphical element being included in a graphical user interface, and a probability distribution of the position of the graphical element in the graphical user interface.

13. The system of claim 11 , wherein the prediction data comprises, a plurality of graphical elements, and for each of the graphical elements, a position of the graphical element in the graphical user interface.

14. The system of claim 11 , wherein the machine learning model comprises a transformer based model that includes an encoder that processes the embedding vector of the natural language textual description, and a decoder that processes the graphical attribute data.

15. The system of claim 11 , the operations further comprising:

providing a natural language textual description of a graphical user interface as input to the machine learned model;

generating, by the machine learned model, based on the natural language textual description of graphical user interface prediction data that is indicative of graphical elements in a graphical user interface;

providing the prediction data to a graphical user interface renderer; and

generating, by the renderer, a graphical user interface based on the prediction data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2022
From: HUANG, ZIFENG; LI, YANG; LI, GANG; ZHOU, XIN; CANNY, JOHN FRANCIS
To: GOOGLE LLC
Reel/Frame 062137/0030 →
Continuity (2)
Provisional Application 63255366 · Oct 13, 2021
Related Publication 20230115185A1 · Apr 13, 2023
References Cited (41)
US 11740879B2 · Huang et al. · 2023 [cited by applicant]
US 20170011279A1 · Soldevila et al. · 2017 [cited by applicant]
US 20200193306A1 · Defiebre · 2020 [cited by examiner]
US 20200401716A1 · Yan et al. · 2020 [cited by applicant]
US 20210397942A1 · Collomosse · 2021 [cited by examiner]
US 20220222046A1 · Schoppe · 2022 [cited by examiner]
US 20230031702A1 · Li · 2023 [cited by examiner]
US 20230350651A1 · Huang et al. · 2023 [cited by applicant]
Ellawela, Chaveen, and K. B. N. Lakmali. “A Review about Voice and UI Design Driven Approaches to Identify UI Elements and Generate UI Designs.” 2021 International Conference on Intelligent Technologies (CONIT). IEEE, J… [cited by examiner]
Kolthoff, Kristian. “Automatic generation of graphical user interface prototypes from unrestricted natural language requirements.” 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEE… [cited by examiner]
Moran, Kevin, et al. “Machine learning-based prototyping of graphical user interfaces for mobile apps.” IEEE Transactions on Software Engineering 46.2 (2018): 196-221 (Year: 2018). [cited by examiner]
Biplab et al., “Rico: A mobile app dataset for building data-driven design applications.” Proceedings of the 30th annual ACM symposium on user interface software and technology. Oct. 2017, 845-854. [cited by applicant]
Bishop, Mixture density networks Feb. 1994, 26 pages. [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners.” Submitted on Jul. 2020, arXiv:2005.14165v4, 75 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of deep bidirectional transformers for language understanding.” Submitted on May 2019, arXiv:1810.04805v2, 16 pages. [cited by applicant]
Diego et al., “Variational Transformer Networks for Layout Generation” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13642-13652. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets” Submitted on Jun. 2014, arXiv:1406.2661v1, 9 pages. [cited by applicant]
Gupta et al., “Layout Generation and Completion with Self-attention” Submitted on Jun. 2020, arXiv:2006.14615, 17 pages. [cited by applicant]
Ha et al., “A neural representation of sketch drawings.” Submitted Apr. 2017, arXiv:1704.03477, 15 pages. [cited by applicant]
Huang et al., “Scones: towards conversational authoring of sketches.” Proceedings of the 25th International Conference on Intelligent User Interfaces, Mar. 2020, 11 pages. [cited by applicant]
Huang et al., “Swire: Sketch-based user interface retrieval.” Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, May 2019, 10 pages. [cited by applicant]
Kang et al., “MetaMap: Supporting visual metaphor ideation through multi-dimensional example-based exploration.” Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, May 2021, 15 pages. [cited by applicant]
Lasecki et al., “Apparition: Crowdsourced user interfaces that come to life as you sketch them.” Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, Apr. 2015, 10 pages. [cited by applicant]
Lee et al., “Neural design network: Graphic layout generation with constraints.” Submitted on Jul. 2020, arXiv:1912.09421v2, 16 pages. [cited by applicant]
Li et al., “LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators” Submitted on Jan. 2019, arXiv:1901.06767v1, 16 pages. [cited by applicant]
Li et al., “Screen2vec: Semantic embedding of gui screens and gui components.” Submitted on Jan. 2021, arXiv:2101.11103v1, 15 pages. [cited by applicant]
Li et al., “Mapping Natural Language Instructions to Mobile UI Action Sequences.” Submitted on Jun. 2020, arXiv:2005.03776v2, 13 pages. [cited by applicant]
Liu et al., “Learning design semantics for mobile apps.” Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology. Oct. 2018, 569-579. [cited by applicant]
Material.io [online], “Material Design at Google I/O 2021” May 6, 2021, retrieved on May 5, 2023, retrieved from URL <https://material.io/blog/material-google-io21>, 12 pages. [cited by applicant]
Norman, The Design of Everyday Things, Basic Books, Inc., USA, pp. 111. [cited by applicant]
Pandian et al., “UISketch: a large-scale dataset of UI element sketches.” Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, May 2021, 14 pages. [cited by applicant]
Parekh et al., “Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO.” Submitted on Mar. 2021, arXiv:2004.15020, 16 pages. [cited by applicant]
Patil et al., “Read: Recursive autoencoders for document layout generation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, 544-545. [cited by applicant]
Ramesh et al., “Zero-Shot Text-to-Image Generation” Submitted on Feb. 2021, arXiv:2102.12092, 20 pages. [cited by applicant]
Vaswani et al., “Attention is All You Need” Submitted on Dec. 2017, arXiv:1706.03762v5, 15 pages. [cited by applicant]
Wang et al., “Screen2words: Automatic mobile UI summarization with multimodal learning.” Submitted on Aug. 2021, arXiv:2108.03353v1, 13 pages. [cited by applicant]
Xia, “Crosspower: Bridging Graphics and Linguistics” In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology, Virtual Event, USA, Oct. 20-23, 2020, 13 pages. [cited by applicant]
Zhang et al., “Cross-modal contrastive learning for text-to-image generation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, 833-842. [cited by applicant]
Li et al., “A formal machine-learning approach to generating human-machine interfaces from task models.” IEEE Transactions on Human-Machine Systems 47.6, May 2017, 822-833. [cited by applicant]
Moran et al., “Machine learning-based prototyping of graphical user interfaces for mobile apps.” IEEE Transactions on Software Engineering 46.2, Jun. 2018, 196-221. [cited by applicant]
Rathnayake et al., “A framework for adaptive user interface generation based on user behavioural patterns.” 2019 Moratuwa Engineering Research Conference (MERCon). IEEE, Jul. 2019, 698-703. [cited by applicant]