Systems and methods for surgical video de-identification
An improved approach is described herein wherein an automated de-identification system is provided to process the raw captured data. The automated de-identification system utilizes specific machine learning data architectures and transforms the raw captured data into processed captured data by modifying, replacing, or obscuring various identifiable features. The processed captured data can include transformed video or audio data.
1 . A system for de-identifying raw data including audio or video from one or more input video streams having one or more visual or audible identifiable artifacts, the system comprising:
a computer processor coupled to a computer memory and one or more non-transitory computer readable media, the computer processor configured to:
maintain a machine learning model architecture configured to track one or more audio or visual artifacts present in the one or more input video streams and to identify the one or more audio or visual identifiable artifacts as objects for de-identification;
transform portions of frames of the one or more input video streams to replace the one or more audio or visual identifiable artifacts with corresponding obfuscated replacement audio or visual artifacts; and
generate one or more output video streams including the transformed portions of the frames of the one or more input video streams having the obfuscated replacement audio or visual artifacts;
wherein momentum-based interpolation is utilized to establish one or more boundaries providing a temporal, visual, or frequency-based margin around the one or more audio or visual identifiable artifacts for generating the corresponding obfuscated replacement audio or visual artifacts; and
wherein the temporal, visual, or frequency-based margin is utilized to include an additional region for expanding the corresponding obfuscated replacement audio or visual artifacts;
wherein the momentum-based interpolation includes using head tracking based on an average displacement of a bounding box between the frames of the one or more input video streams, extending a blurred field around an individual as the individual moves with a degree of extension positively correlated with a speed of the individual.
2 . The system of claim 1 , wherein the computer processor is configured to transform portions of the frames of the one or more input video streams by replacing detected visual objects corresponding to bodies of practitioners with contour maps established around a pixel boundary corresponding to the bodies of the practitioners.
3 . The system of claim 1 , wherein the computer processor is configured to transform portions of the frames of the one or more input video streams by replacing detected visual objects corresponding to heads of practitioners with blurred visual blocks established around a pixel boundary corresponding to the heads of the practitioners.
4 . The system of claim 1 , wherein an inference subnet is configured to use the momentum-based interpolation during head detection, and the head detection is only run on sampled frames and intervening frames which are required to be smoothed or interpolated.
5 . The system of claim 1 , wherein a size of the bounding box is dynamically modified responsive to a detected inconsistency between the frames.
6 . The system of claim 1 , wherein the machine learning model architecture is tunable through modification of one or more hyperparameters to modify a balance between false positives and false negatives.
7 . The system of claim 1 , wherein the one or more output video streams are transmitted to a downstream analytic computing system configured for replay of the one or more output video streams having one or more automatically appended predictive annotations indicative of one or more key sections of the one or more output video streams for analysis.
8 . The system of claim 7 , wherein for the one or more key sections, a different balance between false positives and false negatives is applied relative to one or more non-key sections of the one or more output video streams.
9 . The system of claim 1 , wherein the one or more output video streams are generated for real or near-real time analysis, and one or more refined output video streams are generated with different parameters for batch analysis.
10 . The system of claim 9 , wherein the one or more refined output video streams include less obfuscated audio or video features relative to the one or more output video streams generated for the real or near-real time analysis.
11 . A method for de-identifying raw data including audio or video from one or more input video streams having one or more visual or audible identifiable artifacts, the method comprising:
maintaining a machine learning model architecture configured to track one or more audio or visual artifacts present in the one or more input video streams and to identify the one or more audio or visual identifiable artifacts as objects for de-identification;
transforming portions of frames of the one or more input video streams to replace the one or more audio or visual identifiable artifacts with corresponding obfuscated replacement audio or visual artifacts; and
generating one or more output video streams including the transformed portions of the frames of the one or more input video streams having the obfuscated replacement audio or visual artifacts;
wherein momentum-based interpolation is utilized to establish one or more boundaries providing a temporal, visual, or frequency-based margin around the one or more audio or visual identifiable artifacts for generating the corresponding obfuscated replacement audio or visual artifacts; and
wherein the temporal, visual, or frequency-based margin is utilized to include an additional region for expanding the corresponding obfuscated replacement audio or visual artifacts;
wherein the momentum-based interpolation includes using head tracking based on an average displacement of a bounding box between the frames of the one or more input video streams, extending a blurred field around an individual as the individual moves with a degree of extension positively correlated with a speed of the individual.
12 . The method of claim 11 , wherein the computer processor is configured to transform portions of the frames of the one or more input video streams by replacing detected visual objects corresponding to bodies of practitioners with contour maps established around a pixel boundary corresponding to the bodies of the practitioners.
13 . The method of claim 11 , wherein the computer processor is configured to transform portions of the frames of the one or more input video streams by replacing detected visual objects corresponding to heads of practitioners with blurred visual blocks established around a pixel boundary corresponding to the heads of the practitioners.
14 . The method of claim 11 , wherein an inference subnet is configured to use the momentum-based interpolation during head detection, and the head detection is only run on sampled frames and intervening frames which are required to be smoothed or interpolated.
15 . The method of claim 11 , wherein a size of the bounding box is dynamically modified responsive to a detected inconsistency between the frames.
16 . The method of claim 11 , wherein the machine learning model architecture is tunable through modification of one or more hyperparameters to modify a balance between false positives and false negatives.
17 . The method of claim 11 , wherein the one or more output video streams are transmitted to a downstream analytic computing system configured for replay of the one or more output video streams having one or more automatically appended predictive annotations indicative of one or more key sections of the one or more output video streams for analysis.
18 . The method of claim 17 , wherein for the one or more key sections, a different balance between false positives and false negatives is applied relative to one or more non-key sections of the one or more output video streams.
19 . The method of claim 11 , wherein the one or more output video streams are generated for real or near-real time analysis, and one or more refined output video streams are generated with different parameters for batch analysis.
20 . A non-transitory computer readable medium storing machine interpretable instructions, which when executed by a processor, cause the processor to perform a method for de-identifying raw data including audio or video from one or more input video streams having one or more visual or audible identifiable artifacts, the method comprising:
maintaining a machine learning model architecture configured to track one or more audio or visual artifacts present in the one or more input video streams and to identify the one or more audio or visual identifiable artifacts as objects for de-identification;
transforming portions of frames of the one or more input video streams to replace the one or more audio or visual identifiable artifacts with corresponding obfuscated replacement audio or visual artifacts; and
generating one or more output video streams including the transformed portions of the frames of the one or more input video streams having the obfuscated replacement audio or visual artifacts;
wherein momentum-based interpolation is utilized to establish one or more boundaries providing a temporal, visual, or frequency-based margin around the one or more audio or visual identifiable artifacts for generating the corresponding obfuscated replacement audio or visual artifacts; and
wherein the temporal, visual, or frequency-based margin is utilized to include an additional region for expanding the corresponding obfuscated replacement audio or visual artifacts;
wherein the momentum-based interpolation includes using head tracking based on an average displacement of a bounding box between the frames of the one or more input video streams, extending a blurred field around an individual as the individual moves with a degree of extension positively correlated with a speed of the individual.