IP Library Granted Patent US 12,260,331
Granted Patent B2
US 12,260,331 · App. 18/225,656 · Granted Mar 25, 2025

Distributed labeling for supervised learning

Inventors: Abhishek Bhowmick (Santa Clara, CA); Ryan M. Rogers (Sunnyvale, CA); Umesh S. Vaishampayan (Santa Clara, CA); Andrew H. Vyrros (San Francisco, CA)
Assignee: Apple Inc.
G06N3/08G06N3/04G06N3/044G06N3/045G06N3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,331
App. No.
18/225,656
Granted
Mar 25, 2025
Kind
B2
Abstract

Embodiments described herein provide a technique to crowdsource labeling of training data for a machine learning model while maintaining the privacy of the data provided by crowdsourcing participants. Client devices can be used to generate proposed labels for a unit of data to be used in a training dataset. One or more privacy mechanisms are used to protect user data when transmitting the data to a server. The server can aggregate the proposed labels and use the most frequently proposed labels for an element as the label for the element when generating training data for the machine learning model. The machine learning model is then trained using the crowdsourced labels to improve the accuracy of the model.

Claims (36)

1. A method comprising:

receiving, by an electronic device, an unlabeled set of data from a server;

generating, by the electronic device and using a trained machine learning model, proposed labels for elements of the unlabeled set of data; and

transmitting, by the electronic device, a privatized version of one of the proposed labels to the server.

2. The method of claim 1 , further comprising:

generating, by the electronic device, the privatized version of the one of the proposed labels via a privacy-preserving encoding.

3. The method of claim 1 , wherein the unlabeled set of data includes one or more of text data, image data, application activity data, and device activity data.

4. The method of claim 1 , wherein the trained machine learning model is configured as a convolutional neural network.

5. The method of claim 1 , wherein the trained machine learning model is configured as a recurrent neural network.

6. The method of claim 1 , wherein the proposed labels are encoded to mask individual contributors via a privacy-preserving encoding algorithm.

7. The method of claim 6 , wherein the privacy-preserving encoding algorithm is a differential privacy algorithm.

8. The method of claim 1 , further comprising:

processing the proposed labels to determine an estimate of a most frequent proposed label for the unlabeled set of data by generating a sketch matrix of received proposed labels to aggregate proposed label data.

9. The method of claim 1 , further comprising:

processing the proposed labels to determine an estimate of a most frequent proposed label for the unlabeled set of data by generating a histogram of received proposed labels to aggregate proposed label data.

10. A device comprising:

a memory; and

at least one processor configured to:

receive an unlabeled set of data from a server;

generate, using a trained machine learning model, proposed labels for elements of the unlabeled set of data; and

transmit a privatized version of one of the proposed labels to the server.

11. The device of claim 10 , wherein the at least one processor is further configured to:

generate the privatized version of the one of the proposed labels via a privacy-preserving encoding.

12. The device of claim 10 , wherein the unlabeled set of data includes one or more of text data, image data, application activity data, and device activity data.

13. The device of claim 10 , wherein the trained machine learning model is configured as at least one of a convolutional neural network or a recurrent neural network.

14. The device of claim 10 , wherein the proposed labels are encoded to mask individual contributors via a privacy-preserving encoding algorithm.

15. The device of claim 14 , wherein the privacy-preserving encoding algorithm is a differential privacy algorithm.

16. A non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

receiving, by an electronic device, an unlabeled set of data from a server;

generating, by the electronic device and using a trained machine learning model, proposed labels for elements of the unlabeled set of data; and

transmitting, by the electronic device, a privatized version of one of the proposed labels to the server.

17. The non-transitory machine-readable medium of claim 16 , wherein the operations further comprise:

generating, by the electronic device, the privatized version of the one of the proposed labels via a privacy-preserving encoding.

18. The non-transitory machine-readable medium of claim 16 , wherein the unlabeled set of data includes one or more of text data, image data, application activity data, and device activity data.

19. The non-transitory machine-readable medium of claim 16 , wherein the trained machine learning model is configured as a convolutional neural network.

20. The non-transitory machine-readable medium of claim 16 , wherein the proposed labels are encoded to mask individual contributors via a privacy-preserving encoding algorithm.

Continuity (3)
Continuation 16556066 · Aug 29, 2019
Provisional Application 62738990 · Sep 28, 2018
Related Publication 20240028890A1 · Jan 25, 2024
References Cited (116)
US 4829577A · Kuroda · 1989 [cited by examiner]
US 8417648B2 · Hido · 2013 [cited by examiner]
US 9275347B1 · Harada · 2016 [cited by applicant]
US 9355359B2 · Welinder · 2016 [cited by applicant]
US 9594741B1 · Thakurta et al. · 2017 [cited by applicant]
US 9792562B1 · Chen et al. · 2017 [cited by applicant]
US 10063582B1 · Feng · 2018 [cited by examiner]
US 10275690B2 · Chen · 2019 [cited by applicant]
US 10540578B2 · Madani · 2020 [cited by applicant]
US 10554738B1 · Ren · 2020 [cited by applicant]
US 10726356B1 · Zarandioon · 2020 [cited by applicant]
US 10979461B1 · Cervantez · 2021 [cited by examiner]
US 11093818B2 · Li · 2021 [cited by applicant]
US 11120337B2 · Wu · 2021 [cited by applicant]
US 11194846B2 · Stenneth · 2021 [cited by applicant]
US 11321629B1 · Rowan · 2022 [cited by examiner]
US 11354578B2 · Baker · 2022 [cited by examiner]
US 11379695B2 · Desai · 2022 [cited by examiner]
US 11710035B2 · Bhowmick · 2023 [cited by examiner]
US 11755915B2 · Fenchel · 2023 [cited by examiner]
US 12014426B2 · Myers · 2024 [cited by examiner]
US 20020007334A1 · Dicks · 2002 [cited by applicant]
US 20120203716A1 · Baker · 2012 [cited by examiner]
US 20130066818A1 · Assadollahi et al. · 2013 [cited by applicant]
US 20130346356A1 · Welinder · 2013 [cited by applicant]
US 20140115715A1 · Pasdar · 2014 [cited by applicant]
US 20140279716A1 · Cormack · 2014 [cited by applicant]
US 20150178596A1 · Bengio · 2015 [cited by applicant]
US 20150213360A1 · Venanzi · 2015 [cited by applicant]
US 20150254573A1 · Abu-Mostafa et al. · 2015 [cited by applicant]
US 20150356459A1 · Ghosh · 2015 [cited by applicant]
US 20160247501A1 · Kim · 2016 [cited by applicant]
US 20160277215A1 · Zhou · 2016 [cited by applicant]
US 20160277435A1 · Salajegheh et al. · 2016 [cited by applicant]
US 20170011306A1 · Kim · 2017 [cited by applicant]
US 20170116519A1 · Johnson · 2017 [cited by applicant]
US 20170220949A1 · Feng et al. · 2017 [cited by applicant]
US 20170270244A1 · Murugesan · 2017 [cited by examiner]
US 20170277906A1 · Camenisch et al. · 2017 [cited by applicant]
US 20170316348A1 · Ghosh · 2017 [cited by examiner]
US 20170353855A1 · Joy · 2017 [cited by examiner]
US 20170353865A1 · Li · 2017 [cited by applicant]
US 20170364810A1 · Gusev · 2017 [cited by applicant]
US 20180018590A1 · Szeto · 2018 [cited by applicant]
US 20180053071A1 · Chen · 2018 [cited by applicant]
US 20180101697A1 · Rane · 2018 [cited by applicant]
US 20180114101A1 · Desai · 2018 [cited by applicant]
US 20180114334A1 · Desai · 2018 [cited by applicant]
US 20180137433A1 · Devarakonda · 2018 [cited by applicant]
US 20180336486A1 · Chu · 2018 [cited by applicant]
US 20180349620A1 · Bhowmick · 2018 [cited by applicant]
US 20180357292A1 · Rai · 2018 [cited by examiner]
US 20180357556A1 · Rai · 2018 [cited by examiner]
US 20180357595A1 · Rai · 2018 [cited by examiner]
US 20190050713A1 · Nirihira · 2019 [cited by applicant]
US 20190196795A1 · Cavalier · 2019 [cited by examiner]
US 20190227980A1 · McMahan · 2019 [cited by applicant]
US 20190244138A1 · Bhowmick · 2019 [cited by examiner]
US 20190279085A1 · Umeda · 2019 [cited by examiner]
US 20190311298A1 · Kopp · 2019 [cited by applicant]
US 20190318261A1 · Deng · 2019 [cited by examiner]
US 20190354810A1 · Samel · 2019 [cited by examiner]
US 20190370334A1 · Bhowmick · 2019 [cited by applicant]
US 20200043467A1 · Qian · 2020 [cited by examiner]
US 20200065706A1 · Kao · 2020 [cited by applicant]
US 20200104705A1 · Bhowmick · 2020 [cited by examiner]
US 20200334979A1 · Goncalves · 2020 [cited by applicant]
US 20200410134A1 · Bhowmick · 2020 [cited by applicant]
US 20220129641A1 · Jezewski · 2022 [cited by examiner]
US 20240028890A1 · Bhowmick · 2024 [cited by examiner]
CN 105868773A · 2016 [cited by applicant]
CN 106295697 · 2017 [cited by applicant]
CN 107077487A · 2017 [cited by applicant]
CN 107085585A · 2017 [cited by applicant]
CN 107251060 · 2017 [cited by applicant]
CN 107292330 · 2017 [cited by applicant]
CN 107316049 · 2017 [cited by applicant]
CN 107851213A · 2018 [cited by applicant]
KR 20140117156 · 2014 [cited by applicant]
WO WO2016064576A1 · 2016 [cited by applicant]
WO WO2017052709A2 · 2017 [cited by applicant]
Beaulieu-Jones, “Machine Learning Methods to Identify Hidden Phenotypes in the Electronic Health Record,” Dissertation, University of Pennsylvania, 2017, 178 pages. [cited by applicant]
Chang, et al., “Learning Representations of Emotional Speech With Deep Convolutional Generative Adversarial Networks,” Apr. 22, 2017, retrieved from https://arxiv.org/pdf/1705.02394.pdf, 5 pages. [cited by applicant]
Choi, et al., “Generating Multi-label Discrete Patient Records using Generative Adversarial Networks,” Machine Leaning for Healthcare Conference, PMLR, 2017, 20 pages. [cited by applicant]
Creswell, et al., “Generative Adversarial Networks: An Overview,” Submitted to IEEE—SPM, Apr. 2017, 14 pages. [cited by applicant]
Dwork, et al., “Pan-Private Streaming Algorithms,” Jan. 1, 2010, XP055474103, 32 pages. [cited by applicant]
Github, Github mean-count-sketch, Jul. 15, 2017, retrieved from https://gitbub.com/barrust/pyprobables/blob . . . , 22 pages. [cited by applicant]
Goodfellow, “NIPS 2016 Tutorial: Generative Adversarial Networks,” Apr. 3, 2017, retrieved from https://arxiv.org/pdf/1701.00160.pdf, 57 pages. [cited by applicant]
Goodfellow, et al., “Generative Adversarial Nets,” Jun. 10, 2014, retrieved from https://arxiv.org/pdf/1406.2661.pdf, 9 pages. [cited by applicant]
Google, Google Download Instructions, May 6, 2016, downloaded from the wayback machine, https://web.archive.org/web/20160506185913/https://support.google.com/drive/answer/24 . . . , 1 page. [cited by applicant]
Goyal, et al., “Sketch Algorithms for Estimating Point Queries in NLP,” Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, 2012, pp.… [cited by applicant]
Graepel, et al., “ML Confidential: Machine Learning on Encrypted Data,” Nov. 28, 2012, Information Security and Cryptology ICISC 2012, 21 pages. [cited by applicant]
Hitaj, et al., “Deep Models Under the GAN: Information Leakage from Collaborative Deep Learning,” Feb. 24, 2017, retrieved from https://arxiv.org/pdf/1702.07464.pdf, 14 pages. [cited by applicant]
Kingman, et al., “Semi-supervised Learning with Deep Generative Models,” Oct. 31, 2014, retrieved from https://arxiv.org/pdf/1406.5298.pdf, 9 pages. [cited by applicant]
Kumar, et al., “Improved Semi-supervised Learning with GANs using Manifold Invariances,” May 24, 2017, retrieved from https://arxiv.org/pdf/1705.08850.pdf. [cited by applicant]
Liu, et al., “Cloud-Enabled Privacy-Preserving Collaborative Learning for Mobile Sensing,” Nov. 6, 2012, SenSys '12, pp. 57-70. [cited by applicant]
Liu, et al., “Secure Multi-label Classification over Encrypted Data in Cloud,” Oct. 17, 2017, International Conference on Computer Analysis of Images and Patterns, CAIP 2017, pp. 57-73. [cited by applicant]
Lu, et al., “Poster: A Unified Framework of Differentially Private Synthetic Data Release with Generative Adversarial Network,” CCS '17, Oct. 30-Nov. 3, 2017, Dallas, TX, pp. 2547-2549. [cited by applicant]
Melis, et al., “Efficient Private Statistics with Succinct Sketches,” Jan. 6, 2016, Proceedings of 2016 Network and Distributed System Security Symposium, 15 pages. [cited by applicant]
Nguyen, et al., “Collecting and Analyzing Data from Smart Device Users with Local Differential Privacy,” Jun. 16, 2016, retrieved from https://arxiv.org/pdf/1606.05053.pdf, 11 pages. [cited by applicant]
Papernot, et al., “Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data,” Mar. 3, 2017, retrieved from https://arxiv.org/pdf/1610.05755.pdf, 16 pages. [cited by applicant]
Premachandran, et al., “Unsupervised Learning Using Generative Adversarial Training and Clustering,” ICLR 2017 conference paper, 10 pages. [cited by applicant]
Qin, et al., “Heavy Hitter Estimation over Set-Valued Data with Local Differential Privacy,” Oct. 2016, CCS '16, 12 pages. [cited by applicant]
Radford, et al., “Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks,” Jan. 7, 2016, retrieved from https://arxiv.org/pdf/1511.06434.pdf, 16 pages. [cited by applicant]
Springenberg, “Unsupervised and Semi-supervised Learning with Categorical Generative Adversarial Networks,” Apr. 30, 2016, retrieved from https://arxiv.org/pdf/1511.06390.pdf, 20 pages. [cited by applicant]
Weiss, et al., “An Empirical Comparison of Pattern Recognition, Neural Nets, and Machine Learning Classification Methods,” 1989, IJCAI, vol. 89, pp. 781-787. [cited by applicant]
Zhang, et al., “Differentially Private Releasing via Deep Generative Model (technical report),” retrieved from arXiv preprint arXiv:1801.01594, 2018, 11 pages. [cited by applicant]
Australian Examination Report from Australian Patent Application No. 2019200896, dated Feb. 3, 2020, 4 pages. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 201910106947.3, dated Nov. 24, 2022. [cited by applicant]
European Office Action from European Patent Application No. 19768964.9, dated Feb. 8, 2023, 7 pages. [cited by applicant]
European Search Report from European Patent Application No. 19153349.6, dated Oct. 30, 2019, 16 pages. [cited by applicant]
Partial European Search Report from European Patent Application No. 19153349.6, dated Jun. 21, 2019, 15 pages. [cited by applicant]
Chinese Notice of Allowance from Chinese Patent Application No. 201910106947.3, dated Oct. 16, 2023, 7 pages including machine-generated translation. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 201980053510.6, dated May 29, 2024, 19 pages including English language translation. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 201980053510.6, dated Jan. 9, 2025, 21 pages with English language translation. [cited by applicant]
European Office Action from European Patent Application No. 19768964.9, dated Jan. 29, 2025, 6 pages. [cited by applicant]