Augmented video generation with dental modifications
A method includes receiving a video comprising a face of an individual, the video comprising a current condition of a dental site of the individual. The method includes determining or receiving an altered condition of the dental site and modifying the video by replacing the current condition of the dental site with the altered condition of the dental site in the video.
1 . A system comprising:
a memory; and
a processor operatively coupled to the memory, the processor to:
receive a video comprising a face of an individual, the video comprising a current condition of a dental site of the individual;
receive or determine an altered condition of the dental site; and
modify the video by replacing the current condition of the dental site with the altered condition of the dental site in the video, wherein modifying the video comprises:
determining an inner mouth area of the face in at least one frame of the video;
performing segmentation on the inner mouth area of the at least one frame using one of the following techniques:
a) inputting the inner mouth area of the at least one frame and inner mouth areas of one or more previous frames of the video into a trained machine learning model that segments the inner mouth area into a plurality of individual teeth, or
b) determining an optical flow between the at least one frame and the one or more previous frames, and inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model,
wherein the trained machine learning model segments the inner mouth area of the at least one frame into the plurality of individual teeth in a manner that is temporally consistent with the one or more previous frames; and
replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the dental site.
2 . The system of claim 1 , wherein the dental site comprises one or more teeth, and wherein the one or more teeth in the modified video are different from the one or more teeth in an original version of the video and are temporally stable and consistent between frames of the modified video.
3 . The system of claim 1 , wherein the processor is further to:
identify one or more frames of the modified video that fail to satisfy one or more image quality criteria; and
remove the one or more frames of the modified video that failed to satisfy the one or more image quality criteria.
4 . The system of claim 3 , wherein the processor is further to:
generate replacement frames for the removed one or more frames of the modified video.
5 . The system of claim 1 , wherein the altered condition of the dental site comprises an estimated future condition of the dental site, wherein the dental site comprises one or more teeth, and wherein determining the estimated future condition of the dental site comprises:
generating or receiving a first three-dimensional (3D) model of a dental arch comprising the current condition of the one or more teeth; and
generating or receiving a second 3D model of the dental arch comprising a post-treatment condition of the one or more teeth, the second 3 D model having been generated based on modifying the first 3D model of the dental arch, wherein the post-treatment condition of the one or more teeth corresponds to the estimated future condition of the one or more teeth.
6 . The system of claim 1 , wherein receiving the video of the face of the individual comprises capturing the video using one or more image sensors while the individual views a display, and wherein the processor is further to:
output the modified video to the display while the individual views the display.
7 . The system of claim 1 , wherein determining the inner mouth area for the at least one frame comprises:
inputting the at least one frame into a trained machine learning model, wherein the trained machine learning model outputs a position of the inner mouth area for the at least one frame.
8 . The system of claim 1 , wherein the processor is further to perform the following prior to determining the inner mouth area:
determine a plurality of landmarks for a plurality of frames of the video using a trained machine learning model, wherein the at least one frame is one of the plurality of frames of the video; and
perform smoothing of the plurality of landmarks between the plurality of frames, wherein the inner mouth area is determined based on the plurality of landmarks.
9 . The system of claim 1 , wherein performing the segmentation of the at least one frame comprises inputting the inner mouth area of the at least one frame and inner mouth areas of the one or more previous frames into the trained machine learning model.
10 . The system of claim 1 , wherein the processor is to:
determine the optical flow between the at least one frame and the one or more previous frames;
wherein performing the segmentation of the at least one frame comprises inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model.
11 . The system of claim 1 , wherein the processor is further to:
determine color information for the inner mouth area in the at least one frame;
determine contours of the altered condition of the dental site; and
input at least one of the color information, the determined contours, the at least one frame or information on the inner mouth area into a generative model, wherein the generative model outputs an altered version of the at least one frame.
12 . The system of claim 1 , wherein modifying the video comprises performing the following for at least one frame of the video:
determining an area of interest corresponding to a dental condition in the at least one frame; and
replacing initial data for the area of interest with replacement data determined from the altered condition of the dental site.
13 . The system of claim 1 , wherein the video comprises a plurality of frames, and wherein modifying the video comprises performing the following for one or more frames of the plurality of frames:
inputting data from the one or more frames and the altered condition of the dental site into a trained generative model, wherein the trained generative model outputs a modified version of the one or more frames.
14 . The system of claim 1 , wherein the processor is further to:
receive a three-dimensional (3D) model of the dental site generated based on intraoral scanning of an oral cavity of the individual;
determine the altered condition based on modifying the 3D model of the dental site; and
for each frame of the video, project the modified 3D model of the dental site onto a plane associated with the frame of the video.
15 . A non-transitory computer readable medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising:
receiving a video comprising a face of an individual, the video comprising a current condition of dentition of the individual;
receiving or determining an altered condition of the dentition; and
modifying the video by replacing the current condition of the dentition with the altered condition of the dentition in the video, wherein modifying the video comprises:
determining an inner mouth area of the face in at least one frame of the video;
performing segmentation on the inner mouth area of the at least one frame using one of the following techniques:
a) inputting the inner mouth area of the at least one frame and inner mouth areas of one or more previous frames of the video into a trained machine learning model that segments the inner mouth area into a plurality of individual teeth, or
b) determining an optical flow between the at least one frame and the one or more previous frames, and inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model,
wherein the trained machine learning model segments the inner mouth area of the at least one frame into the plurality of individual teeth in a manner that is temporally consistent with the one or more previous frames; and
replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the dentition.
16 . The non-transitory computer readable medium of claim 15 , wherein the altered condition of the dentition in the modified video is temporally stable and consistent between frames of the modified video.
17 . The non-transitory computer readable medium of claim 15 , the operations further comprising:
identifying one or more frames of the modified video that fail to satisfy one or more image quality criteria;
removing the one or more frames of the modified video that failed to satisfy the one or more image quality criteria; and
generating replacement frames for the removed one or more frames of the modified video.
18 . The non-transitory computer readable medium of claim 15 , wherein the altered condition of the dentition comprises an estimated future condition of the dentition, and wherein determining the estimated future condition of the dentition comprises:
generating or receiving a first three-dimensional (3D) model of a dental arch comprising the current condition of the dentition; and
generating or receiving a second 3D model of the dental arch comprising a post-treatment condition of the dentition, the second 3D model having been generated based on modifying the first 3D model of the dental arch, wherein the post-treatment condition of the dentition corresponds to the estimated future condition of the dentition.
19 . The non-transitory computer readable medium of claim 15 , wherein modifying the video comprises performing the following for at least one frame of the video:
determining an inner mouth area of the face in the at least one frame; and
replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the dentition.
20 . A method comprising:
receiving a video comprising a face of an individual, the video comprising a current condition of one or more teeth of the individual;
receiving or determining an altered condition of the one or more teeth; and
modifying the video by replacing the current condition of the one or more teeth with the altered condition of the one or more teeth in the video, wherein modifying the video comprises:
determining an inner mouth area of the face in at least one frame of the video;
performing segmentation on the inner mouth area of the at least one frame using one of the following techniques:
a) inputting the inner mouth area of the at least one frame and inner mouth areas of one or more previous frames of the video into a trained machine learning model that segments the inner mouth area into a plurality of individual teeth, or
b) determining an optical flow between the at least one frame and the one or more previous frames, and inputting the inner mouth area of the at least one frame and the optical flow into the trained machine learning model,
wherein the trained machine learning model segments the inner mouth area of the at least one frame into the plurality of individual teeth in a manner that is temporally consistent with the one or more previous frames; and
replacing initial data for the inner mouth area of the face with replacement data determined from the altered condition of the one or more teeth.