IP Library › Granted Patent US 12,272,001
Granted Patent B2
US 12,272,001 · App. 17/937,418 · Granted Apr 8, 2025

Rapid generation of 3D heads with natural language

Inventors: Joseph Logan Olson (San Mateo, CA); Mager Kamel Aquino (Montreal, CA); Jade Raymond (Montreal, CA)
Assignee: Sony Interactive Entertainment LLC
G06T17/20G06F40/40G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,001
App. No.
17/937,418
Granted
Apr 8, 2025
Kind
B2
Abstract

Two dimensional images are converted to a 3D neural radiance field (NeRF), which is modified based on text input to resemble the type of character demanded by the text. An open-source “CLIP” model scores how well an image matches a line of text to produce a final 3D NeRF, which may be converted to a polygonal mesh and imported into a computer simulation such as a computer game.

Claims (15)

1. A device comprising:

at least one computer storage that is not a transitory signal and that comprises instructions executable by at least one processor system to:

generate a base three dimensional (3D) neural radiance field (NeRF) from plural images;

use text input to a Contrastive Language-Image Pre-training (CLIP) model to generate a modified NeRF from the base 3D NeRF; and

convert the modified NeRF to a polygonal mesh representing a virtual human head for presentation of the virtual human head in at least one computer simulation; wherein the instructions are executable to:

use a machine learning (ML) model on the base 3D NeRF to minimize a loss indication in matching the text;

train the ML model on a chain of causality from initial image parameters that control vertices of an object to pixels of the object rendered onscreen.

2. The device of claim 1 , wherein the CLIP model rates an image match to the text.

3. The device of claim 2 , wherein the CLIP model is trained on image-text pairs using cosine similarity to score a goodness of match.

4. The device of claim 1 , wherein the ML model comprises at least one fully connected (non-convolutional) deep network.

5. The device of claim 1 , wherein input to the ML model comprises values representing three spatial dimensions and two viewing dimensions.

6. The device of claim 5 , wherein output of the ML model comprises volume density and view-dependent emitted radiance.

7. The device of claim 1 , wherein the instructions are executable to:

generate the text from a starting phrase using learned ensuing phrases.

8. The device of claim 1 , comprising the at least one processor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2022
From: OLSON, JOSEPH LOGAN; AQUINO, MAGER KAMEL; RAYMOND, JADE
To: SONY INTERACTIVE ENTERTAINMENT LLC
Reel/Frame 061414/0011 →
Continuity (1)
Related Publication 20240112403A1 · Apr 4, 2024
References Cited (15)
US 20190279411A1 · Mitchell et al. · 2019 [cited by applicant]
US 20220036153A1 · O'Malia et al. · 2022 [cited by applicant]
US 20220269867A1 · Yang et al. · 2022 [cited by applicant]
US 20220284662A1 · Bond · 2022 [cited by examiner]
US 20220414959A1 · Peng · 2022 [cited by examiner]
US 20230052645A1 · Keller · 2023 [cited by examiner]
US 20230334071A1 · Whitehead, Jr. · 2023 [cited by examiner]
US 20230334754A1 · Kirchmayer · 2023 [cited by examiner]
“International Search Report and Written Opinion”, dated Dec. 18, 2023, from the counterpart PCT application PCT/US23/74151. [cited by applicant]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever, “Proceedings of the 38th Internationa… [cited by applicant]
Dellaert, Frank, “NeRF at CVPR 2022”, Jun. 1, 2022, retrieved from https://dellaert.github.io/NeRF22/. [cited by applicant]
Jain et al., “Zero-Shot Text-Guided Object Generation with Dream Fields”, CVPR 2022. 13 pages. Website: https://ajayj.com/dreamfield. [cited by applicant]
Olson et al., “Hyper-Personalized Game Items”, file history of related U.S. Appl. No. 17/938,322, filed Oct. 5, 2022. (1275-052). [cited by applicant]
Poole et al., “Dreamfusion: TEXT-to-3D Using 2D Diffusion”, Google Research, UC Berkeley, Sep. 29, 2022. [cited by applicant]
Tewari et al. “Advances in Neural Rendering”, Mar. 30, 2022, Retrieved from the Internet, entire document. [cited by applicant]