IP Library › Granted Patent US 11,586,731
Granted Patent B2
US 11,586,731 · App. 16/583,256 · Granted Feb 21, 2023

Risk-aware entity linking

Inventors: Juan Pablo Bottaro (Dublin, IE); Daria Bogdanova (Dublin, IE); Maria Laura Jedrzejowska (Dublin, IE)
Assignee: Microsoft Technology Licensing, LLC
G06F21/562G06F16/217G06F16/243G06F16/24575G06N20/00G06F2221/032
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,586,731
App. No.
16/583,256
Granted
Feb 21, 2023
Kind
B2
Abstract

In an embodiment, the disclosed technologies include identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity; inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data; based on the qualitative data, determining link risk data; based on the link risk data and the semantic similarity data, and determining whether to generate a link between the content item and the target entity.

Claims (82)

1. A method, comprising:

identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity;

inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data;

based on the qualitative data, determining link risk data;

the link risk data indicates a likelihood of a negative link decision; and

based on the link risk data and the semantic similarity data, determining whether to generate a link between the content item and the target entity;

wherein the method is performed by at least one computing device.

2. The method of claim 1 , further comprising:

using a model of the set of digital models, performing a sentiment analysis on the at least one feature of the content item to determine and output a sentiment score; and

based on the sentiment score, determining the link risk data.

3. The method of claim 1 , further comprising:

using a model of the set of digital models, performing a topic analysis on the at least one feature of the content item to determine and output a topic descriptor associated with the content item; and

based on the topic descriptor and a sentiment score, determining the link risk data.

4. The method of claim 1 , further comprising:

using a model of the set of digital models, analyzing a position feature of the candidate entity within the content item and a frequency of occurrence of the candidate entity within the content item to determine and output an aboutness score; and

based on the aboutness score, determining the link risk data.

5. The method of claim 1 , further comprising:

using a model of the set of digital models, analyzing a sentiment feature of the content item relative to the candidate entity to determine and output an entity sentiment score that associates a sentiment with the candidate entity; and

based on the entity sentiment score, determining the link risk data.

6. The method of claim 1 , further comprising:

using a first set of digital models to analyze at least one feature of the content item to determine and output at least one content score;

using a second set of digital models to analyze at least one feature of the target entity to determine and output at least one entity score; and

combining the at least one content score and the at least one entity score to determine the link risk data.

7. The method of claim 1 , further comprising:

tuning at least one threshold used to determine the link risk data in response to an input from a user of a software application that is operatively coupled to the second digital data source.

8. A method, comprising:

identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity;

inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data;

using a model of the set of digital models, performing a malware analysis on the at least one feature of the content item and the at least one feature of the target entity;

based on the qualitative data and output of the malware analysis, determining link risk data; and

based on the link risk data and the semantic similarity data, determining whether to generate a link between the content item and the target entity.

9. A method, comprising:

identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity;

inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data;

performing a quantitative analysis of engagement data associated with the target entity;

based on the qualitative data and output of the quantitative analysis, determining link risk data; and

based on the link risk data and the semantic similarity data, determining whether to generate a link between the content item and the target entity.

10. A method, comprising:

identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity;

inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data;

based on the qualitative data, determining link risk data;

based on the link risk data and the semantic similarity data, determining whether to generate a link between the content item and the target entity; and

causing at least a portion of the content item to be included in a digital notification transmitted to a user of a software application that is operatively coupled to the second digital data source.

11. The method of claim 10 , further comprising:

determining and analyzing engagement data indicative of user engagement with the digital notification; and

using the engagement data to tune at least one threshold used to determine the link risk data.

12. One or more storage media storing instructions which, when executed by one or more processors, cause:

identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity;

inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data;

based on the qualitative data, determining link risk data;

the link risk data indicates a likelihood of a negative link decision; and

based on the link risk data and the semantic similarity data, determining whether to generate a link between the content item and the target entity.

13. The one or more storage media of claim 12 , wherein the instructions, when executed by the one or more processors, further cause:

using a model of the set of digital models, performing a sentiment analysis on the at least one feature of the content item to determine and output a sentiment score; and

based on the sentiment score, determining the link risk data.

14. The one or more storage media of claim 12 , wherein the instructions, when executed by the one or more processors, further cause:

using a model of the set of digital models, performing a topic analysis on the at least one feature of the content item to determine and output a topic descriptor associated with the content item; and

based on the topic descriptor and a sentiment score, determining the link risk data.

15. The one or more storage media of claim 12 , wherein the instructions, when executed by the one or more processors, further cause:

using a model of the set of digital models, analyzing a position feature of the candidate entity within the content item and a frequency of occurrence of the candidate entity within the content item to determine and output an aboutness score; and

based on the aboutness score, determining the link risk data.

16. The one or more storage media of claim 12 , wherein the instructions, when executed by the one or more processors, further cause:

using a model of the set of digital models, analyzing a sentiment feature of the content item relative to the candidate entity to determine and output an entity sentiment score that associates a sentiment with the candidate entity; and

based on the entity sentiment score, determining the link risk data.

17. The one or more storage media of claim 12 , wherein the instructions, when executed by the one or more processors, further cause:

using a first set of digital models to analyze at least one feature of the content item to determine and output at least one content score;

using a second set of digital models to analyze at least one feature of the target entity to determine and output at least one entity score; and

combining the at least one content score and the at least one entity score to determine the link risk data.

18. The one or more storage media of claim 12 , wherein the instructions, when executed by the one or more processors, further cause:

tuning at least one threshold used to determine the link risk data in response to an input from a user of a software application that is operatively coupled to the second digital data source.

19. One or more storage media storing instructions which, when executed by one or more processors, cause:

identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity;

inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data;

using a model of the set of digital models, performing a malware analysis on the at least one feature of the content item and the at least one feature of the target entity;

based on the qualitative data and output of the malware analysis, determining link risk data; and

based on the link risk data and the semantic similarity data, determining whether to generate a link between the content item and the target entity.

20. One or more storage media storing instructions which, when executed by one or more processors, cause:

identifying a content item of a first digital data source as a candidate for linking with a target entity of a second digital data source by matching a candidate entity mentioned in the content item to the target entity in accordance with semantic similarity data computed between the candidate entity and the target entity;

inputting at least one feature of the content item and at least one feature of the target entity to a set of digital models that analyze the at least one feature of the content item and the at least one feature of the target entity and determine and output qualitative data;

performing a quantitative analysis of engagement data associated with the target entity;

based on the qualitative data and output of the quantitative analysis, determining the link risk data; and

based on the link risk data and the semantic similarity data, determining whether to generate a link between the content item and the target entity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2019
From: BOTTARO, JUAN PABLO; BOGDANOVA, DARIA; JEDRZEJOWSKA, MARIA LAURA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 050560/0894 →
Continuity (1)
Related Publication 20210097178A1 · Apr 1, 2021