IP Library › Granted Patent US 12,405,774
Granted Patent B2
US 12,405,774 · App. 18/348,191 · Granted Sep 2, 2025

Creating user interface using machine learning

Inventors: Zifeng Huang (Emeryville, CA); Yang Li (Palo Alto, CA); Xin Zhou (Mountain View, CA); Gang Li (Mountain View, CA); John Francis Canny (Berkeley, CA)
Assignee: Google LLC
G06F8/38G06F3/0484G06F8/33G06F40/40G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,774
App. No.
18/348,191
Filed
Jul 6, 2023
Granted
Sep 2, 2025
Kind
B2
Art Unit
2144
USPC
715/762
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for training and using machine learning models to generate graphical user interfaces from textual descriptions.

Claims (47)

1. A computer-implemented method of predicting a graphical user interface, comprising:

generating, for a natural language textual description, using a pre-trained word embedding model, an encoded representation of the natural language textual description;

providing the encoded representation of the natural language textual description to machine learned model that is configured to receive as input, the encoded representation and generate as output, a first embedding vector;

determining second embedding vectors, each second embedding vector determined from graphical attribute data describing a corresponding user interface image in a training dataset;

selecting one of the corresponding user interface images based on the first embedding vector and the second embedding vectors; and

providing the selected corresponding user interface images as output in response to the natural language text description;

wherein the machine learned model comprises a first encoder and a second encoder trained on a training dataset comprising a plurality of training samples, and wherein:

each training sample including:

a user interface image that includes a plurality of graphical elements; and

a natural language textual description of the user interface image;

the first encoder receives, as input, the encoded representation of the natural language textual description and generates, as output, the first embedding vector; and

the second encoder receives, as input, the graphical attribute data and generates, as output, the second embedding vector.

2. The computer-implemented method of claim 1 , wherein determining second embedding vectors comprises accessing a data store storing second embedding vectors, where each second embedding vector is a generated by a second encoder model that generates the embedding vector from graphical attribute data describing a corresponding user interface image.

3. The computer-implemented method of claim 1 , wherein selecting one of the corresponding user interface images based on the first embedding vector and the second embedding vectors comprises generating a dot product of each second embedding vector and the first embedding vector.

4. The computer-implemented method of claim 1 , wherein the graphical attribute data that describes a corresponding user interface image is graphical attribute data that, for each graphical element of the user interface image, describes an attribute type of the graphical element, and a position of the graphical element.

5. A computer storage medium encoded with a computer program, the program comprising instructions that when executed by data processing apparatus cause the data processing apparatus to perform the operations of:

generating, for a natural language textual description, using a pre-trained word embedding model, an encoded representation of the natural language textual description;

providing the encoded representation of the natural language textual description to machine learned model that is configured to receive as input, the encoded representation and generate as output, a first embedding vector;

determining second embedding vectors, each second embedding vector determined from graphical attribute data describing a corresponding user interface image in a training dataset;

selecting one of the corresponding user interface images based on the first embedding vector and the second embedding vectors; and

providing the selected corresponding user interface images as output in response to the natural language text description;

wherein the machine learned model comprises a first encoder and a second encoder trained on a training dataset comprising a plurality of training samples, and wherein:

each training sample including:

a user interface image that includes a plurality of graphical elements; and

a natural language textual description of the user interface image;

the first encoder receives, as input, the encoded representation of the natural language textual description and generates, as output, the first embedding vector; and

the second encoder receives, as input, the graphical attribute data and generates, as output, the second embedding vector.

6. The computer storage medium of claim 5 , wherein determining second embedding vectors comprises accesses a data store storing second embedding vectors, where each second embedding vector is a generated by a second encoder model that generates the embedding vector from graphical attribute data describing a corresponding user interface image.

7. The computer storage medium of claim 5 , wherein selecting one of the corresponding user interface images based on the first embedding vector and the second embedding vectors comprises generating a dot product of each second embedding vector and the first embedding vector.

8. The computer storage medium of claim 5 , wherein the graphical attribute data that describes a corresponding user interface image is graphical attribute data that, for each graphical element of the graphical user interface image, describes an attribute type of the graphical element, and a position of the graphical element.

9. A system, comprising:

a data processing apparatus; and

a computer storage medium encoded with a computer program, the program comprising instructions that when executed by the data processing apparatus cause the data processing apparatus to perform the operations of:

generating, for a natural language textual description, using a pre-trained word embedding model, an encoded representation of the natural language textual description;

providing the encoded representation of the natural language textual description to machine learned model that is configured to receive as input, the encoded representation and generate as output, a first embedding vector;

determining second embedding vectors, each second embedding vector determined from graphical attribute data describing a corresponding user interface image in a training dataset;

selecting one of the corresponding user interface images based on the first embedding vector and the second embedding vectors; and

providing the selected corresponding user interface images as output in response to the natural language text description;

wherein the machine learned model comprises a first encoder and a second encoder trained on a training dataset comprising a plurality of training samples, and wherein:

each training sample including:

a user interface image that includes a plurality of graphical elements; and

a natural language textual description of the user interface image;

the first encoder receives, as input, the encoded representation of the natural language textual description and generates, as output, the first embedding vector; and

the second encoder receives, as input, the graphical attribute data and generates, as output, the second embedding vector.

10. The system of claim 9 , wherein determining second embedding vectors comprises accesses a data store storing second embedding vectors, where each second embedding vector is a generated by a second encoder model that generates the embedding vector from graphical attribute data describing a corresponding user interface image.

11. The system of claim 9 , wherein selecting one of the corresponding user interface images based on the first embedding vector and the second embedding vectors comprises generating a dot product of each second embedding vector and the first embedding vector.

12. The system of claim 9 , wherein the graphical attribute data that describes a corresponding user interface image is graphical attribute data that, for each graphical element of the graphical user interface image, describes an attribute type of the graphical element, and a position of the graphical element.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2023
From: HUANG, ZIFENG; LI, YANG; ZHOU, XIN; LI, GANG; CANNY, JOHN FRANCIS
To: GOOGLE LLC
Reel/Frame 064723/0662 →
Continuity (3)
Continuation 18046428 · Oct 13, 2022
Provisional Application 63255366 · Oct 13, 2021
Related Publication 20230350651A1 · Nov 2, 2023
References Cited (42)
US 20170011279A1 · Soldevila · 2017 [cited by examiner]
US 20200193306A1 · Defiebre · 2020 [cited by applicant]
US 20200401716A1 · Yan · 2020 [cited by examiner]
US 20210397942A1 · Collomosse · 2021 [cited by examiner]
US 20220222046A1 · Schoppe et al. · 2022 [cited by applicant]
US 20230031702A1 · Li · 2023 [cited by applicant]
US 20230115185A1 · Huang et al. · 2023 [cited by applicant]
Li, Toby Jia-Jun, et al. “Screen2vec: Semantic embedding of gui screens and gui components.” Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 2021 (Year: 2021). [cited by examiner]
Juárez-Ramírez, Reyes, Carlos Huertas, and Sergio Inzunza. “Automated generation of user-interface prototypes based on controlled natural language description.” 2014 IEEE 38th International Computer Software and Applica… [cited by examiner]
Kolthoff, Kristian. “Automatic generation of graphical user interface prototypes from unrestricted natural language requirements.” 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEE… [cited by examiner]
Kolthoff, “Automatic generation of graphical user interface prototypes from unrestricted natural language requirements.” 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2019, 1… [cited by applicant]
Ellawela et al., “A Review about Voice and UI Design Driven Approaches to Identify UI Elements and Generate UI Designs.” 2021 International Conference on Intelligent Technologies (CON IT), IEEE, Jun. 25-27, 2021, 4 page… [cited by applicant]
Biplab et al., “Rico: A mobile app dataset for building data-driven design applications.” Proceedings of the 30th annual ACM symposium on user interface software and technology. Oct. 2017, 845-854. [cited by applicant]
Bishop, Mixture density networks, Feb. 1994, 26 pages. [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners.” CoRR, Submitted on Jul. 2020, arXiv:2005.14165v4, 75 pages. [cited by applicant]
Devlin et al., “BERT: Pre-training of deep bidirectional transformers for language understanding.” CoRR, Submitted on May 2019, arXiv:1810.04805v2, 16 pages. [cited by applicant]
Diego et al., “Variational Transformer Networks for Layout Generation” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13642-13652. [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets” CoRR, Submitted on Jun. 2014, arXiv:1406.2661v1, 9 pages. [cited by applicant]
Gupta et al., “Layout Generation and Completion with Self-attention” CoRR, Submitted on Jun. 2020, arXiv:2006.14615, 17 pages. [cited by applicant]
Ha et al., “A neural representation of sketch drawings.” CoRR, Submitted Apr. 2017, arXiv:1704.03477, 15 pages. [cited by applicant]
Huang et al., “Scones: towards conversational authoring of sketches.” Proceedings of the 25th International Conference on Intelligent User Interfaces, Mar. 2020, 11 pages. [cited by applicant]
Huang et al., “Swire: Sketch-based user interface retrieval.” Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, May 2019, 10 pages. [cited by applicant]
Kang et al., “MetaMap: Supporting visual metaphor ideation through multi-dimensional example-based exploration.” Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, May 2021, 15 pages. [cited by applicant]
Lasecki et al., “Apparition: Crowdsourced user interfaces that come to life as you sketch them.” Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, Apr. 2015, 10 pages. [cited by applicant]
Lee et al., “Neural design network: Graphic layout generation with constraints.” CoRR, Submitted on Jul. 2020, arXiv:1912.09421v2, 16 pages. [cited by applicant]
Li et al., “A formal machine-learning approach to generating human-machine interfaces from task models.” IEEE Transactions on Human-Machine Systems 47.6, May 2017, 822-833. [cited by applicant]
Li et al., “LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators” CoRR, Submitted on Jan. 2019, arXiv:1901.06767v1, 16 pages. [cited by applicant]
Li et al., “Screen2vec: Semantic embedding of gui screens and gui components.” CoRR, Submitted on Jan. 2021, arXiv:2101.11103v1, 15 pages. [cited by applicant]
Li et al., “Mapping Natural Language Instructions to Mobile UI Action Sequences.” CoRR, Submitted on Jun. 2020, arXiv:2005.03776v2, 13 pages. [cited by applicant]
Liu et al., “Learning design semantics for mobile apps.” Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology. Oct. 2018, 569-579. [cited by applicant]
Material.io [online], “Material Design at Google I/O 2021” May 6, 2021, retrieved on May 5, 2023, retrieved from URL <https://material.io/blog/material-google-io21>, 12 pages. [cited by applicant]
Moran et al., “Machine learning-based prototyping of graphical user interfaces for mobile apps.” IEEE Transactions on Software Engineering 46.2, Jun. 2018, 196-221. [cited by applicant]
Norman, The Design of Everyday Things, Basic Books, Inc., USA, pp. 111. [cited by applicant]
Pandian et al., “UISketch: a large-scale dataset of UI element sketches.” Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, May 2021, 14 pages. [cited by applicant]
Parekh et al., “Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO.” CoRR, Submitted on Mar. 2021, arXiv:2004.15020, 16 pages. [cited by applicant]
Patil et al., “Read: Recursive autoencoders for document layout generation.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, 544-545. [cited by applicant]
Ramesh et al., “Zero-Shot Text-to-Image Generation” CoRR, Submitted on Feb. 2021, arXiv:2102.12092, 20 pages. [cited by applicant]
Rathnayake et al., “A framework for adaptive user interface generation based on user behavioural patterns.” 2019 Moratuwa Engineering Research Conference (MERCon). IEEE, Jul. 2019, 698-703. [cited by applicant]
Vaswani et al., “Attention is All You Need” CoRR, Submitted on Dec. 2017, arXiv:1706.03762v5, 15 pages. [cited by applicant]
Wang et al., “Screen2words: Automatic mobile UI summarization with multimodal learning.” CoRR, Submitted on Aug. 2021, arXiv:2108.03353v1, 13 pages. [cited by applicant]
Xia, “Crosspower: Bridging Graphics and Linguistics” In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology, Virtual Event, USA, Oct. 20-23, 2020, 13 pages. [cited by applicant]
Zhang et al., “Cross-modal contrastive learning for text-to-image generation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, 833-842. [cited by applicant]