IP Library Granted Patent US 12711702
Granted Patent B2
US 12711702 · App. 18/525,530 · Granted Aug 18, 2026

Augmented video generation with dental modifications

Inventors: Philipp Kopp (Zürich, CH); Dmitry Yurievich Chekh (Moscow, RU); Niko Benjamin Huber (Zug, CH); Shipra Jain (Zürich, CH); Christopher E. Cramer (Durham, NC); Chad Clayton Brown (Cary, NC); Vladislav Andreevich Miryaha (Dolgoprudniy, RU); Boris Aleksandrovich Vysokanov (Moscow, RU); Eric Paul Meyer (Zürich, CH); Maik Gerth (Seeheim-Jugenheim, DE); Sinan Ibrahim Bayraktar (Zürich, CH); Michael Seeber (Zürich, CH); Ritika Chakraborty (Kirchdorf, CH); Andreea Maria Radoescu (Baden, CH); Doruk Cetin (Zürich, CH)
Assignee: Align Technology, Inc.
G06T17/00G06T7/0012G06T7/11G06T7/13G06T7/174G06T7/90G06T11/00G06T19/20G06T2207/10016G06T2207/20081G06T2207/30036G06T2207/30168G06T2219/2012G06T2219/2016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711702
App. No.
18/525,530
Granted
Aug 18, 2026
Kind
B2
Abstract

A method includes receiving a video comprising a face of an individual, the video comprising a current condition of a dental site of the individual. The method includes determining or receiving an altered condition of the dental site and modifying the video by replacing the current condition of the dental site with the altered condition of the dental site in the video.

Claims (76)

1 . A system comprising:

a memory; and

a processor operatively coupled to the memory, the processor to:

receive a video comprising a face of an individual, the video comprising a current condition of a dental site of the individual;

receive or determine an altered condition of the dental site; and

modify the video by replacing the current condition of the dental site with the altered condition of the dental site in the video, wherein modifying the video comprises:

determining an inner mouth area of the face in at least one frame of the video;

performing segmentation on the inner mouth area of the at least one frame using one of the following techniques:

a) inputting the inner mouth area of the at least one frame and inner mouth areas of one or more previous frames of the video into a trained machine learning model that segments the inner mouth area into a plurality of individual teeth, or

b) determining an optical flow between the at least one frame and the one or more previous frames, and inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model,

wherein the trained machine learning model segments the inner mouth area of the at least one frame into the plurality of individual teeth in a manner that is temporally consistent with the one or more previous frames; and

replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the dental site.

2 . The system of claim 1 , wherein the dental site comprises one or more teeth, and wherein the one or more teeth in the modified video are different from the one or more teeth in an original version of the video and are temporally stable and consistent between frames of the modified video.

3 . The system of claim 1 , wherein the processor is further to:

identify one or more frames of the modified video that fail to satisfy one or more image quality criteria; and

remove the one or more frames of the modified video that failed to satisfy the one or more image quality criteria.

4 . The system of claim 3 , wherein the processor is further to:

generate replacement frames for the removed one or more frames of the modified video.

5 . The system of claim 1 , wherein the altered condition of the dental site comprises an estimated future condition of the dental site, wherein the dental site comprises one or more teeth, and wherein determining the estimated future condition of the dental site comprises:

generating or receiving a first three-dimensional (3D) model of a dental arch comprising the current condition of the one or more teeth; and

generating or receiving a second 3D model of the dental arch comprising a post-treatment condition of the one or more teeth, the second 3 D model having been generated based on modifying the first 3D model of the dental arch, wherein the post-treatment condition of the one or more teeth corresponds to the estimated future condition of the one or more teeth.

6 . The system of claim 1 , wherein receiving the video of the face of the individual comprises capturing the video using one or more image sensors while the individual views a display, and wherein the processor is further to:

output the modified video to the display while the individual views the display.

7 . The system of claim 1 , wherein determining the inner mouth area for the at least one frame comprises:

inputting the at least one frame into a trained machine learning model, wherein the trained machine learning model outputs a position of the inner mouth area for the at least one frame.

8 . The system of claim 1 , wherein the processor is further to perform the following prior to determining the inner mouth area:

determine a plurality of landmarks for a plurality of frames of the video using a trained machine learning model, wherein the at least one frame is one of the plurality of frames of the video; and

perform smoothing of the plurality of landmarks between the plurality of frames, wherein the inner mouth area is determined based on the plurality of landmarks.

9 . The system of claim 1 , wherein performing the segmentation of the at least one frame comprises inputting the inner mouth area of the at least one frame and inner mouth areas of the one or more previous frames into the trained machine learning model.

10 . The system of claim 1 , wherein the processor is to:

determine the optical flow between the at least one frame and the one or more previous frames;

wherein performing the segmentation of the at least one frame comprises inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model.

11 . The system of claim 1 , wherein the processor is further to:

determine color information for the inner mouth area in the at least one frame;

determine contours of the altered condition of the dental site; and

input at least one of the color information, the determined contours, the at least one frame or information on the inner mouth area into a generative model, wherein the generative model outputs an altered version of the at least one frame.

12 . The system of claim 1 , wherein modifying the video comprises performing the following for at least one frame of the video:

determining an area of interest corresponding to a dental condition in the at least one frame; and

replacing initial data for the area of interest with replacement data determined from the altered condition of the dental site.

13 . The system of claim 1 , wherein the video comprises a plurality of frames, and wherein modifying the video comprises performing the following for one or more frames of the plurality of frames:

inputting data from the one or more frames and the altered condition of the dental site into a trained generative model, wherein the trained generative model outputs a modified version of the one or more frames.

14 . The system of claim 1 , wherein the processor is further to:

receive a three-dimensional (3D) model of the dental site generated based on intraoral scanning of an oral cavity of the individual;

determine the altered condition based on modifying the 3D model of the dental site; and

for each frame of the video, project the modified 3D model of the dental site onto a plane associated with the frame of the video.

15 . A non-transitory computer readable medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising:

receiving a video comprising a face of an individual, the video comprising a current condition of dentition of the individual;

receiving or determining an altered condition of the dentition; and

modifying the video by replacing the current condition of the dentition with the altered condition of the dentition in the video, wherein modifying the video comprises:

determining an inner mouth area of the face in at least one frame of the video;

performing segmentation on the inner mouth area of the at least one frame using one of the following techniques:

a) inputting the inner mouth area of the at least one frame and inner mouth areas of one or more previous frames of the video into a trained machine learning model that segments the inner mouth area into a plurality of individual teeth, or

b) determining an optical flow between the at least one frame and the one or more previous frames, and inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model,

wherein the trained machine learning model segments the inner mouth area of the at least one frame into the plurality of individual teeth in a manner that is temporally consistent with the one or more previous frames; and

replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the dentition.

16 . The non-transitory computer readable medium of claim 15 , wherein the altered condition of the dentition in the modified video is temporally stable and consistent between frames of the modified video.

17 . The non-transitory computer readable medium of claim 15 , the operations further comprising:

identifying one or more frames of the modified video that fail to satisfy one or more image quality criteria;

removing the one or more frames of the modified video that failed to satisfy the one or more image quality criteria; and

generating replacement frames for the removed one or more frames of the modified video.

18 . The non-transitory computer readable medium of claim 15 , wherein the altered condition of the dentition comprises an estimated future condition of the dentition, and wherein determining the estimated future condition of the dentition comprises:

generating or receiving a first three-dimensional (3D) model of a dental arch comprising the current condition of the dentition; and

generating or receiving a second 3D model of the dental arch comprising a post-treatment condition of the dentition, the second 3D model having been generated based on modifying the first 3D model of the dental arch, wherein the post-treatment condition of the dentition corresponds to the estimated future condition of the dentition.

19 . The non-transitory computer readable medium of claim 15 , wherein modifying the video comprises performing the following for at least one frame of the video:

determining an inner mouth area of the face in the at least one frame; and

replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the dentition.

20 . A method comprising:

receiving a video comprising a face of an individual, the video comprising a current condition of one or more teeth of the individual;

receiving or determining an altered condition of the one or more teeth; and

modifying the video by replacing the current condition of the one or more teeth with the altered condition of the one or more teeth in the video, wherein modifying the video comprises:

determining an inner mouth area of the face in at least one frame of the video;

performing segmentation on the inner mouth area of the at least one frame using one of the following techniques:

a) inputting the inner mouth area of the at least one frame and inner mouth areas of one or more previous frames of the video into a trained machine learning model that segments the inner mouth area into a plurality of individual teeth, or

b) determining an optical flow between the at least one frame and the one or more previous frames, and inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model,

wherein the trained machine learning model segments the inner mouth area of the at least one frame into the plurality of individual teeth in a manner that is temporally consistent with the one or more previous frames; and

replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the one or more teeth.