IP Library › Granted Patent US 12,537,994
Granted Patent B2
US 12,537,994 · App. 18/263,597 · Granted Jan 27, 2026

Selective content masking for collaborative computing

Inventors: Dhandapani Shanmugam (Bangalore, IN); Sreenivas Makam (Bangalore, IN)
Assignee: GOOGLE LLC
H04N21/45455H04L65/403
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,537,994
App. No.
18/263,597
Filed
Jul 31, 2023
Granted
Jan 27, 2026
Kind
B2
Art Unit
2693
USPC
348/14.07
Abstract

A machine-learned sharing system and methods are provided for sharing content with users while masking sensitive information. The system receives a content stream for display to one or more users, converts the content stream into image data representative of at least a portion of the content stream, inputs the image data into a machine-learned model configured for masking sensitive content within shared content, receives from the machine-learned model a first mask indicative of a region within the first content stream that contains sensitive content, and renders a display of the content stream that masks the sensitive content based at least in part on the first mask indicative of the region of the first content stream having the sensitive content.

Claims (81)

1 . A computer-implemented method for sharing content within a videoconferencing application, comprising:

receiving, by a computing system comprising one or more computing devices, a request from a first participant in a video conference to share with one or more additional participants of the video conference a first content stream within the videoconferencing application;

converting, by the computing system, at least a portion of the first content stream into image data representative of a display of the first content stream;

inputting, by the computing system, the image data representative of the display of the first content stream into a machine-learned model configured for masking sensitive content within shared content;

processing, by using the machine-learned model, the image data representative of the display of the first content stream to generate a first mask indicative of a region within the first content stream that contains sensitive content; and

rendering, by the computing system, for the video conference a display of the first content stream that masks the sensitive content based at least in part on the first mask indicative of the region of the first content stream having the sensitive content.

2 . The computer-implemented method of claim 1 , wherein:

said converting the first content stream into image data comprises converting the first content stream into a plurality of image frames, each image frame including image data representing at least a portion of the first content stream;

said inputting the image data representative of the display of the first content stream into the machine-learned model comprises inputting the plurality of image frames into the machine-learned model; and

the first mask is indicative of a first region within at least a first frame of the plurality of image frames that includes sensitive content.

3 . The computer-implemented method of claim 2 , further comprising:

partitioning, by the computing system using the machine-learned model, each frame into a set of logical partitions;

wherein the first mask identifies at least one logical partition of the first frame of the plurality of image frames as containing sensitive content;

wherein rendering for the video conference the display of the first content stream comprises masking the at least one logical partition of the first frame.

4 . The computer-implemented method of claim 2 , further comprising:

determining, by the computing system with the machine-learned model, that a second frame of the plurality of image frames includes sensitive content; and

receiving, by the computing system from the machine-learned model, a second mask indicative of a second region within the second frame that contains sensitive content;

wherein the first region and the second region are at different locations within a display area of the first content stream.

5 . The computer-implemented method of claim 1 , further comprising:

analyzing, by the computing system, the image data representative of the display of the first content stream to determine representative content of the first content stream;

wherein rendering for the video conference the display of the first content stream that masks the sensitive content comprises rendering the representative content for the region within the first content stream that contains the sensitive content.

6 . The computer-implemented method of claim 1 , wherein:

the computing system includes a host computing device, a first client computing device associated with the first participant and at least one additional client computing device associated with the one or more additional participants; and

the machine-learned model is configured at the first client computing device.

7 . The computer-implemented method of claim 6 , wherein:

the videoconferencing application is implemented at least partially by a web browser at the first client computing device.

8 . The computer-implemented method of claim 1 , wherein:

the computing system includes a host computing device, a first client computing device associated with the first participant and at least one additional client computing device associated with the at least one additional participant; and

the machine-learned model is configured at the host computing device.

9 . The computer-implemented method of claim 1 , wherein:

the first content stream is generated by a first application of a client device associated with the first participant.

10 . The computer-implemented method of claim 1 , wherein:

the first content stream is a multimedia content stream including a plurality of content types;

said converting at least a portion of the first content stream into image data representative of the display of the first content stream comprises converting the multimedia content stream into image data representative of the plurality of content types.

11 . The computer-implemented method of claim 1 , further comprising:

obtaining, by the computing system, data descriptive of the machine-learned model;

obtaining, by the computing system, one or more sets of training data comprising image data labeled to indicate sensitive content; and

training, by the computing system, the machine-learned model based on the one or more sets of training data, wherein training the machine-learned model comprises:

determining one or more parameters of a loss function based on the one or more sets of training data; and

modifying at least a portion of the machine-learned model based at least in part on the one or more parameters of the loss function.

12 . A computing system, comprising:

one or more processors; and

one or more non-transitory, computer-readable media that store instructions that when executed by the one or more processors cause the computing system to perform operations, the operations comprising:

receiving a request from a first participant in a video conference to share with one or more additional participants of the video conference a first content stream within a videoconferencing application;

converting at least a portion of the first content stream into image data representative of a display of the first content stream;

inputting the image data representative of the display of the first content stream into a machine-learned model configured for masking sensitive content within shared content;

processing, by using the machine-learned model, the image data representative of the display of the first content stream to generate a first mask indicative of a region within the first content stream that contains sensitive content; and

receiving from the machine-learned model a first mask indicative of a region within the first content stream that contains sensitive content; and

rendering for the video conference a display of the first content stream that masks the sensitive content based at least in part on the first mask indicative of the region of the first content stream having the sensitive content.

13 . The computing system of claim 12 , wherein:

said converting the first content stream into image data comprises converting the first content stream into a plurality of image frames, each image frame including image data representing at least a portion of the first content stream;

said inputting the image data representative of the display of the first content stream into the machine-learned mode comprises inputting the plurality of image frames into the machine-learned model; and

the first mask is indicative of a first region within at least a first frame of the plurality of image frames that includes sensitive content.

14 . The computing system of claim 13 , wherein the operations further comprise:

partitioning, by the machine-learned model, each frame into a set of logical partitions;

wherein the first mask identifies at least one logical partition of the first frame of the plurality of image frames as containing sensitive content; and

wherein rendering for the video conference the display of the first content stream comprises masking the at least one logical partition of the first frame.

15 . The computing system of claim 13 , wherein the operations further comprise:

determining with the machine-learned model that a second frame of the plurality of image frames includes sensitive content; and

receiving from the machine-learned model a second mask indicative of a second region within the second frame that contains sensitive content;

wherein the first region and the second region are at different locations within a display area of the first content stream.

16 . The computing system of claim 15 , wherein:

the videoconferencing application is implemented at least partially by a web browser at a client computing device.

17 . One or more non-transitory computer-readable media that store instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising:

receiving a first content stream for display to one or more users;

converting the first content stream into a plurality of image frames including image data representative of at least a portion of the first content stream;

inputting the plurality of image frames into a machine-learned model configured for masking sensitive content within shared content;

processing, by using the machine-learned model, the plurality of image frames to generate a first mask indicative of a region within the first content stream that contains sensitive content; and

receiving from the machine-learned model a first mask indicative of a region within the first content stream that contains sensitive content; and

rendering a display of the first content stream that masks the sensitive content based at least in part on the first mask indicative of the region of the first content stream having the sensitive content.

18 . The one or more non-transitory computer-readable media of claim 17 , further comprising:

partitioning, by the machine-learned model, each frame into a set of logical partitions;

wherein the first mask identifies at least one logical partition of a first frame of the plurality of image frames as containing sensitive content; and

wherein rendering for the display of the first content stream comprises masking the at least one logical partition of the first frame.

19 . The one or more non-transitory computer-readable media of claim 17 , wherein the first mask is indicative of a first region within at least a first frame of the plurality of image frames that includes sensitive content, the operations further comprising:

determining with the machine-learned model that a second frame of the plurality of image frames includes sensitive content; and

receiving from the machine-learned model a second mask indicative of a second region within the second frame that contains sensitive content;

wherein the first region and the second region are at different locations within a display area of the first content stream.

20 . The one or more non-transitory computer-readable media of claim 17 , wherein:

the first content stream is a multimedia content stream including a plurality of content types; and

said converting the first content stream into a plurality of image frames representative comprises converting the multimedia content stream into image data representative of the plurality of content types.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2023
From: SHANMUGAM, DHANDAPANI; MAKAM, SREENIVAS
To: GOOGLE LLC
Reel/Frame 064491/0839 →
Continuity (1)
Related Publication 20240089537A1 · Mar 14, 2024
References Cited (17)
US 8826322B2 · Bliss et al. · 2014 [cited by applicant]
US 9762855B2 · Browne et al. · 2017 [cited by applicant]
US 10713794B1 · He et al. · 2020 [cited by applicant]
US 11006077B1 · Truong · 2021 [cited by examiner]
US 20110314387A1 · Gold et al. · 2011 [cited by applicant]
US 20150032686A1 · Kuchoor · 2015 [cited by applicant]
US 20150371049A1 · Xavier · 2015 [cited by examiner]
US 20200159958A1 · Kochura et al. · 2020 [cited by applicant]
US 20200218961A1 · Kanazawa et al. · 2020 [cited by applicant]
US 20200314483A1 · Rakshit et al. · 2020 [cited by applicant]
US 20210182430A1 · Negi · 2021 [cited by examiner]
WO WO2020243059 · 2020 [cited by applicant]
WO WO2021033853 · 2021 [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2022/031976, mailed Dec. 14, 2023, 9 pages. [cited by applicant]
Perrigo, “Chrome is Finally Rolling Out the Ability to Hide Your Notifications While Screen Sharing.”, Jan. 27, 2021, https://chromeunboxed.com/chrome-auto-hide-notifications-while-screensharing, retrieved on May 6, 202… [cited by applicant]
Sultana et al., “Evolution of Image Segmentation Using Deep Convolutional Neural Network: A Survey”, Knowledge-Based Systems, Elsevier, vol. 201, May 26, 2020, Amsterdam, Netherlands, XP086179027, 38 pages. [cited by applicant]
International Search Report for Application No. PCT/US2022/031976, Sep. 9, 2022, 4 pages. [cited by applicant]