IP Library › Granted Patent US 12,013,968
Granted Patent B2
US 12,013,968 · App. 17/076,896 · Granted Jun 18, 2024

Data anonymization for data labeling and development purposes

Inventor: Sascha Lange (Vila do Conde, PT)
Assignee: Robert Bosch GmbH
G06F21/6254G06F16/48G06T5/70G06T7/70G06T11/00G10L21/007G10L21/0232G10L25/57G06T2207/10016G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,013,968
App. No.
17/076,896
Granted
Jun 18, 2024
Kind
B2
Abstract

A method and system are disclosed for anonymizing data for labeling and development purposes. A data storage backend has a database of non-anonymous data that is received from a data source. An anonymization engine of the data storage backend generates anonymized data by removing personally identifiable information from the non-anonymous data. These anonymized data are made available to human labelers who manually provide labels based on the anonymized data using a data labeling tool. These labels are then stored in association with the corresponding non-anonymous data, which can then be used for training one or more machine learning models. In this way, non-anonymous data having personally identifiable information can be manually labelled for development purposes without exposing the personally identifiable information to any human labelers.

Claims (72)

1. A method for labeling data, the method comprising:

storing, in a memory of a server, a plurality of data files, the data files including at least one of videos, images, and audio clips, at least some of the plurality of data files including personally identifiable information of individual people;

generating, with a processor of the server, a plurality of anonymized data files by removing the personally identifiable information from each of the plurality of data files, each of the plurality of anonymized data files being generated based on a corresponding one of the plurality of data files; and

labeling the plurality of data files by, for each respective anonymized data file in the plurality of anonymized data files:

transmitting, with a transceiver of the server, the respective anonymized data file to a respective client device in a plurality of client devices;

receiving, with the transceiver, at least one label for the respective anonymized data file from the respective client device; and

storing, in the memory, the at least one label in association with a respective data file in the plurality of data files that corresponds to the respective anonymized data file,

wherein the generating the plurality of anonymized data files includes, for a first audio clip in the plurality of data files, (i) generating, with the processor, an isolated audio clip by isolating a voice of a first person in the first audio clip and (ii) determining, with the processor, a first anonymized audio clip by converting the isolated audio clip into a frequency domain.

2. The method according to claim 1 , the generating the plurality of anonymized data files further comprising, for each respective frame of a first video in the plurality of data files:

determining, with the processor, a location of a face of each person in the respective frame; and

generating, with the processor, a respective anonymized frame of a first anonymized video by blurring portions of the respective frame corresponding to the location of the face of each person.

3. The method according to claim 2 , the generating the plurality of anonymized data files further comprising, for each respective frame of the first video:

determining, with the processor, whether an error occurred in the determination the location of the face of each person; and

generating, with the processor, the respective anonymized frame of the first anonymized video by blurring portions of the respective frame corresponding to an entirety of each person in the respective frame, in response to determining that the error occurred.

4. The method according to claim 1 , the generating the plurality of anonymized data files further comprising:

determining, with the processor, a location of a face of each person in a first image in the plurality of data files; and

generating, with the processor, an anonymized image by blurring portions of the first image corresponding to the location of the face of each person.

5. The method according to claim 1 , the generating the plurality of anonymized data files further comprising:

generating, with the processor, a distorted audio clip by distorting a voice of each person in a first audio clip in the plurality of data files; and

segmenting, with the processor, the distorted audio clip into a first plurality of anonymized audio clip segments.

6. The method according to claim 5 , the labeling the plurality of data files further comprising, for the first plurality of anonymized audio clip segments:

transmitting, with the transceiver, each respective anonymized audio clip segment in the first plurality of anonymized audio clip segments to a different client device in the plurality of client devices;

receiving, with the transceiver, at least one label for each respective anonymized audio clip segment in the first plurality of anonymized audio clip segments from the different client device in the plurality of client devices; and

storing, in the memory, the at least one label received for each respective anonymized audio clip segment in the first plurality of anonymized audio clip segments in association with the first audio clip.

7. The method according to claim 1 , the generating the isolated audio clip further comprising at least one of:

removing, with the processor, background noise from the first audio clip; and

removing, with the processor, a voice of a second person from the first audio clip.

8. The method according to claim 1 , the generating the plurality of anonymized data files further comprising, for a first video in the plurality of data files that was captured during a same time period as the first audio clip:

detecting, with the processor, a face of the first person in the first video;

determining, with the processor, a classification of an emotional state of the first person based on the face of the first person; and

generating, with the processor, a first anonymized video by adding graphical elements to the first video, the graphical elements being configured to indicate the classification of the emotional state of the first person.

9. The method according to claim 8 , the generating the first anonymized video further comprising at least one of:

adding, with the processor, the graphical elements to the first video such that the graphical elements obscure the face of the first person; and

blurring portions of the first video corresponding to the face of the first person in the first video.

10. The method according to claim 8 , the labeling the plurality of data files further comprising, for the first anonymized audio clip and the first anonymized video:

transmitting, with the transceiver, the first anonymized audio clip and the first anonymized video to the respective client device in the plurality of client devices;

receiving, with the transceiver, at least one label for the first anonymized audio clip and the first anonymized video from the respective client device; and

storing, in the memory, the at least one label in association with the first audio clip and the first video.

11. The method according to claim 10 , the labeling the plurality of data files further comprising, for the first anonymized audio clip and the first anonymized video:

displaying, with a display device of the respective client device, the first anonymized video; and

displaying, with the display device, the first anonymized audio clip as a frequency domain graph.

12. The method according to claim 1 , the labeling the plurality of data files by further comprising, for each respective anonymized data file in the plurality of anonymized data files:

outputting, with an output device of the respective client device, the respective anonymized data file; and

receiving, via a user interface of the respective client device, inputs indicating the at least one label for the respective anonymized data file.

13. The method according to claim 1 , the labeling the plurality of data files further comprising, for a respective frame of a first anonymized video in the plurality of anonymized data files:

displaying, with a display of the respective client device, the respective frame of the first anonymized video; and

receiving, via a user interface of the respective client device, inputs indicating the at least one label for the respective frame of the first anonymized video.

14. The method according to claim 1 , the labeling the plurality of data files further comprising, for a respective frame of the first anonymized video in the plurality of anonymized data files:

transmitting, with the transceiver, the respective frame of the first anonymized video and a respective frame of a first video of the plurality of data files that corresponds to the respective frame of the first anonymized video to the respective client device, the respective frame of the first video including a face of a person and the respective frame of the first anonymized video including a blurred face of the person;

displaying, with a display of the respective client device, a movable cursor that is movable on the display via a user interface of the respective client device; and

displaying, with the display, a blended combination of the respective frame of the first anonymized video and the respective frame of the first video, the blended combination being such that a portion of the blended combination around the movable cursor displays the respective frame of the first video and remaining portions of the blended combination display the respective frame of the first anonymized video.

15. The method according to claim 1 , the labeling the plurality of data files further comprising, for a first image in the plurality of anonymized data files:

displaying, with a display of the respective client device, the first image; and

receiving, via a user interface of the respective client device, inputs indicating the at least one label for the first image.

16. The method according to claim 1 , the labeling the plurality of data files further comprising, for a first audio clip in the plurality of anonymized data files:

outputting, with a speaker of the respective client device, the first audio clip; and

receiving, via a user interface of the respective client device, inputs indicating the at least one label for the first audio clip.

17. The method according to claim 1 further comprising:

training, with the processor, a machine learning model using the plurality of data files and the at least one label stored in association with each respective data file in the plurality of data files; and

processing, with the processor, new data files received from a data source based using the trained machine learning model.

18. A system for labeling data, the system comprising:

a transceiver configured to communicate with a plurality of client devices;

a memory configured to store a plurality of data files, the data files including at least one of videos, images, and audio clips, at least some of the plurality of data files including personally identifiable information of individual people; and

a processor operably connected to the transceiver and the memory, the processor configured to:

generate a plurality of anonymized data files by removing the personally identifiable information from each of the plurality of data files, each of the plurality of anonymized data files being generated based on a corresponding one of the plurality of data files; and

label the plurality of data files by, for each respective anonymized data file in the plurality of anonymized data files, (i) operating the transceiver to transmit the respective anonymized data file to a respective client device in the plurality of client devices, (ii) operating the transceiver to receive at least one label for the respective anonymized data file from the respective client device, (iii) and writing, to the memory, the at least one label in association with a respective data file in the plurality of data files that corresponds to the respective anonymized data file,

wherein the generating the plurality of anonymized data files includes, for a first audio clip in the plurality of data files, (i) generating, with the processor, an isolated audio clip by isolating a voice of a first person in the first audio clip and (ii) determining, with the processor, a first anonymized audio clip by converting the isolated audio clip into a frequency domain.

19. A non-transitory computer-readable medium for labeling data, the computer-readable medium storing program instructions that, when executed by a processor, cause the processor to:

read, from a memory, a plurality of data files, the data files including at least one of videos, images, and audio clips, at least some of the plurality of data files including personally identifiable information of individual people

generate a plurality of anonymized data files by removing the personally identifiable information from each of the plurality of data files, each of the plurality of anonymized data files being generated based on a corresponding one of the plurality of data files; and

label the plurality of data files by, for each respective anonymized data file in the plurality of anonymized data files, (i) operating a transceiver to transmit the respective anonymized data file to a respective client device in a plurality of client devices, (ii) operating the transceiver to receive at least one label for the respective anonymized data file from the respective client device, (iii) and writing, to the memory, the at least one label in association with a respective data file in the plurality of data files that corresponds to the respective anonymized data file,

wherein the labeling the plurality of data files further includes, for a respective frame of the first anonymized video in the plurality of anonymized data files, (i) transmitting, with the transceiver, the respective frame of the first anonymized video and a respective frame of a first video of the plurality of data files that corresponds to the respective frame of the first anonymized video to the respective client device, the respective frame of the first video including a face of a person and the respective frame of the first anonymized video including a blurred face of the person, (ii) displaying, with a display of the respective client device, a movable cursor that is movable on the display via a user interface of the respective client device; and (iii) displaying, with the display, a blended combination of the respective frame of the first anonymized video and the respective frame of the first video, the blended combination being such that a portion of the blended combination around the movable cursor displays the respective frame of the first video and remaining portions of the blended combination display the respective frame of the first anonymized video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2020
From: LANGE, SASCHA
To: ROBERT BOSCH GMBH
Reel/Frame 054379/0073 →
Continuity (1)
Related Publication 20220129582A1 · Apr 28, 2022