IP Library Granted Patent US 12,189,676
Granted Patent B2
US 12,189,676 · App. 18/158,416 · Granted Jan 7, 2025

Intent-based copypasta filtering

Inventors: Shilpi Agrawal (Bengaluru, IN); Sumit Srivastava (Bengaluru, IN); Ashish Tripathy (Bengaluru, IN); Grace W. Tang (Los Altos, CA); Hitesh Manwani (Bengaluru, IN)
Assignee: Microsoft Technology Licensing, LLC
G06F16/435G06F16/45
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,676
App. No.
18/158,416
Granted
Jan 7, 2025
Kind
B2
Abstract

Embodiments of copypasta filtering system technologies cluster digital content items into copypasta clusters, extract a first feature set from the digital content items in the copypasta clusters, apply a first set of filters to the first feature set, and based on output of the first set of filters, divide the copypasta clusters into first intent copypasta clusters and possible second intent copypasta clusters. A second feature set is extracted from the digital content items in the possible second intent copypasta clusters. A second set of filters is applied to the second feature set. Based on output of the second set of filters, second intent copypasta clusters are created.

Claims (95)

1. A method comprising:

clustering digital content items distributed by an online system into copypasta clusters;

extracting a first feature set from the digital content items in the copypasta clusters;

applying a first set of filters to the first feature set;

based on output of the first set of filters, dividing the copypasta clusters into first intent copypasta clusters and second intent copypasta clusters;

extracting a second feature set from the digital content items in the second intent copypasta clusters;

applying a second set of filters different from the first set of filters to the second feature set;

based on output of the second set of filters, creating third intent copypasta clusters;

executing a first downstream action on the first intent copypasta clusters, wherein the first downstream action comprises allowing distribution of the first intent copypasta clusters through the online system; and

executing a second downstream action different from the first downstream action on the third intent copypasta clusters.

2. The method of claim 1 , wherein executing the first downstream action comprises:

labeling the first intent copypasta clusters with a first intent label; and

based on the first intent label, distributing content items in the first intent copypasta clusters to a first portion of the online system.

3. The method of claim 2 , wherein executing the second downstream action comprises:

labeling the third intent copypasta clusters with a second intent label different from the first intent label; and

based on the second intent label, distributing content items in the third intent copypasta clusters to a second portion of the online system different from the first portion of the online system.

4. The method of claim 1 , wherein:

executing the first downstream action comprises labeling the first intent copypasta clusters with a first intent label;

executing the second downstream action comprises labeling the third intent copypasta clusters with a second intent label different from the first intent label; and

the method further comprises creating training data comprising the first intent copypasta clusters labeled with the first intent label and the third intent copypasta clusters labeled with the second intent label, and training a machine learning model based on the training data.

5. The method of claim 1 , wherein:

executing the first downstream action comprises labeling the first intent copypasta clusters with a first intent label, and scoring content items in the first intent copypasta clusters based on the first intent label; and

executing the second downstream action comprises labeling the third intent copypasta clusters with a second intent label different from the first intent label, and (ii) scoring content items in the third intent copypasta clusters based on the second intent label.

6. The method of claim 1 , wherein executing the second downstream action comprises sending the third intent copypasta clusters to a content moderation system.

7. The method of claim 1 , wherein:

(i) extracting the first feature set comprises extracting first attribute data from content items within a copypasta cluster; and

(ii) the first attribute data comprises at least one of:

a count of authors of the content items within the copypasta cluster;

a count of organizations associated with the authors of the content items within the copypasta cluster;

a count of the content items within the copypasta cluster that are re-shares of an original content item within the copypasta cluster; or

intent labels output by an intent model for the content items within the copypasta cluster.

8. The method of claim 1 , wherein the first set of filters comprises at least one of:

a first filter that associates copypasta clusters that have an author count equal to an author count threshold with the first intent copypasta clusters;

a second filter that associates copypasta clusters that have an organization count equal to an organization count threshold with the first intent copypasta clusters;

a third filter that associates copypasta clusters that have a re-share count equal to a re-share count threshold with the first intent copypasta clusters; or

a fourth filter that associates copypasta clusters that have intent labels in a set of approved intent labels with the first intent copypasta clusters.

9. The method of claim 1 , wherein:

(i) extracting the second feature set comprises extracting second attribute data from content items within a copypasta cluster; and

(ii) the second attribute data comprises at least one of:

a count of the content items within the copypasta cluster that have an associated user-submitted report;

a count of views of an original content item within the copypasta cluster; or

a count of signal keywords in comments associated with the content items within the copypasta cluster.

10. The method of claim 1 , wherein the second set of filters comprises at least one of:

a first filter that moves, from the second intent copypasta clusters to the third intent copypasta clusters, copypasta clusters that have a user-submitted report count greater than a user-submitted report threshold;

a second filter that moves, from the second intent copypasta clusters to the third intent copypasta clusters, copypasta clusters that have a signal keyword count greater than a signal keyword count threshold; or

a third filter that moves, from the second intent copypasta clusters to the third intent copypasta clusters, copypasta clusters that have an original view count greater than an original view count threshold.

11. A system comprising:

at least one processor;

at least one memory coupled to the at least one processor;

the at least one memory comprises instructions that when executed by the at least one processor, cause the at least one processor to perform operations comprising:

using a first set of filters, dividing copypasta clusters into first intent copypasta clusters and second intent copypasta clusters;

applying a second set of filters to the second intent copypasta clusters;

based on output of the second set of filters, creating third intent copypasta clusters;

executing a first downstream action on the first intent copypasta clusters, wherein the first downstream action comprises allowing distribution of the first intent copypasta clusters through an online system; and

executing a second downstream action different from the first downstream action on the third intent copypasta clusters.

12. The system of claim 11 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

extracting a first feature set from digital content items in the copypasta clusters;

applying the first set of filters to the first feature set;

extracting a second feature set from digital content items in the second intent copypasta clusters; and

applying the second set of filters to the second feature set.

13. The system of claim 11 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

extracting first attribute data from content items within a copypasta cluster, wherein the first attribute data comprises at least one of:

a count of authors of the content items within the copypasta cluster;

a count of organizations associated with the authors of the content items within the copypasta cluster;

a count of the content items within the copypasta cluster that are re-shares of an original content item within the copypasta cluster; or

intent labels output by an intent model for the content items within the copypasta cluster.

14. The system of claim 11 , wherein the first set of filters comprises at least one of:

a first filter that associates copypasta clusters that have an author count equal to an author count threshold with the first intent copypasta clusters;

a second filter that associates copypasta clusters that have an organization count equal to an organization count threshold with the first intent copypasta clusters;

a third filter that associates copypasta clusters that have a re-share count equal to a re-share count threshold with the first intent copypasta clusters; or

a fourth filter that associates copypasta clusters that have intent labels in a set of approved intent labels with the first intent copypasta clusters.

15. The system of claim 11 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

extracting second attribute data from content items within a copypasta cluster, wherein the second attribute data comprises at least one of:

a count of the content items within the copypasta cluster that have an associated user-submitted report;

a count of views of an original content item within the copypasta cluster; or

a count of signal keywords in comments associated with the content items within the copypasta cluster.

16. The system of claim 11 , wherein the second set of filters comprises at least one of:

a first filter that moves, from the second intent copypasta clusters to the third intent copypasta clusters, copypasta clusters that have a user-submitted report count greater than a user-submitted report threshold;

a second filter that moves, from the second intent copypasta clusters to the third intent copypasta clusters, copypasta clusters that have a signal keyword count greater than a signal keyword count threshold; or

a third filter that moves, from the second intent copypasta clusters to the third intent copypasta clusters, copypasta clusters that have an original view count greater than an original view count threshold.

17. The system of claim 11 , wherein;

executing the first downstream action comprises:

labeling the first intent copypasta clusters with a first intent label; and

based on the first intent label, distributing content items in the first intent copypasta clusters to a first portion of an online system; and

executing the second downstream action comprises:

labeling the third intent copypasta clusters with a second intent label different from the first intent label; and

based on the second intent label, distributing content items in the third intent copypasta clusters to a second portion of the online system different from the first portion of the online system.

18. The system of claim 11 , wherein:

executing the first downstream action comprises labeling the first intent copypasta clusters with a first intent label;

executing the second downstream action comprises labeling the third intent copypasta clusters with a second intent label different from the first intent label; and

the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising (a) creating training data comprising the first intent copypasta clusters labeled with the first intent label and the third intent copypasta clusters labeled with the second intent label, and (b) using the training data, training a machine learning model to classify copypasta clusters.

19. The system of claim 11 , wherein:

executing the first downstream action comprises labeling the first intent copypasta clusters with a first intent label, and scoring content items in the first intent copypasta clusters based on the first intent label; and

executing the second downstream action comprises labeling the third intent copypasta clusters with a second intent label different from the first intent label, and scoring content items in the third intent copypasta clusters based on the second intent label.

20. The system of claim 11 , wherein executing the third downstream action comprises sending the second intent copypasta clusters to a content moderation system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2023
From: AGRAWAL, SHILPI; SRIVASTAVA, SUMIT; TRIPATHY, ASHISH; TANG, GRACE W.; MANWANI, HITESH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063467/0001 →
Continuity (1)
Related Publication 20240248926A1 · Jul 25, 2024
References Cited (3)
US 20160259859A1 · Sathish · 2016 [cited by examiner]
US 20200320978A1 · Chatterjee · 2020 [cited by examiner]
“Copypasta”, Retrieved From: https://en.wikipedia.org/wiki/Copypasta#:˜:text=A%20copypasta%20is%20a%20block,users%20and%20disrupt%20online%20discourse, Dec. 14, 2022, 3 Pages. [cited by applicant]