IP Library › Granted Patent US 10,289,676
Granted Patent B2
US 10,289,676 · App. 15/114,648 · Granted May 14, 2019

Method for semantic analysis of a text

Inventor: Jean-Pierre Malle (Saint Michel S/Orge, FR)
Assignee: DEADIA
G06F17/271G06F17/2785G06F17/3071G06F17/30684G06F17/27G06F17/28G06F17/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,289,676
App. No.
15/114,648
Granted
May 14, 2019
Kind
B2
Abstract

The present invention relates to the field of computer-based semantic understanding. Specifically, it relates to a method for semantic analysis of a natural-language text by data-processing means with a view to the classification thereof.

Claims (46)

1. A method for semantic analysis of a text in natural language received by a piece of equipment from input means, the method being characterized in that it comprises performing, by data processing means of the piece of equipment, steps for:

(a) Syntactically parsing at least one text portion into a plurality of words;

(b) Filtering words of said text portion with respect to a plurality of a list of reference words stored on data storage means of the piece of equipment, each being associated with a theme, so as to identify:

The set of the words of said text portion associated with at least one theme,

The set of the themes of said text portion;

(c) Constructing a plurality of subsets of the set of the words of said text portion associated with at least one theme;

(d) For each of said subsets and for each identified theme, computing:

a coverage coefficient of the theme and/or a relevance coefficient of the theme depending on the occurrences in said text portion, of reference words associated with the theme;

at least one orientation coefficient of the theme from the words of said text portion not belonging to the subset;

(e) For each of said subsets and for each identified theme, computing a semantic coefficient representative of a meaning level borne by a sub-group depending on said coverage, relevance and/or orientation coefficients of the theme;

(f) Selecting according to the semantic coefficients at least one subset/theme pair;

(g) Classifying the text according to said at least one selected subset/theme pair,

wherein the method comprises a preliminary step (a0) for parsing the text into a plurality of propositions, each being a text portion for which the steps (a) to (d) of the method are repeated so as to obtain for each proposition a set of coverage, relevance, and/or orientation coefficients associated with the proposition, the method comprising before step (e) a computing step (e0) for each of said subsets and for each identified theme, for at least one proposition of the text of a global coverage coefficient of the theme and/or of a global relevance coefficient of the theme, and of at least one global orientation coefficient of the theme depending on the set of said coefficients associated with a proposition.

2. The method according to claim 1 , wherein a coverage coefficient of a theme is computed in step (d) like the number N of reference words associated with the theme comprised in said subset.

3. The method according to claim 1 , wherein a relevance coefficient of a theme is computed in step (d) with the formula N*(1+ln(R)), wherein N is the number of reference words associated with the theme comprised in the subset and R is the total number of occurrences in said text portion of reference words associated with the theme.

4. The method according to claim 1 , wherein two orientation coefficients of the theme are computed in step (c), including a certainty coefficient of the theme and a grade coefficient of the theme.

5. The method according to claim 4 , wherein a certainty coefficient of a theme is computed in step (d) as having the value:

1 if the words not belonging to the subset are representative of an affirmative proximity with the theme;

−1 if the words not belonging to the subset are representative of a negative proximity with the theme;

0 if the words not belonging to the subset are representative of an uncertain proximity with the theme.

6. The method according to claim 4 , wherein a grade coefficient of a theme is a positive scalar greater than 1 when the words not belonging to the subset are representative of an amplification of the theme, and a positive scalar of less than 1 when the words not belonging to the subset are representative of an attenuation of the theme.

7. The method according to claim 1 , wherein a global coverage coefficient of a theme is computed in step (e0) as the sum of the coverage coefficients of the theme associated with a proposition less the number of reference words of the theme present in at least two propositions.

8. The method according to claim 1 , wherein a global relevance coefficient of a theme is computed in step (e0) as the sum of the relevance coefficients of the theme associated with a proposition.

9. The method according to claim 1 , wherein a global orientation coefficient of a theme is computed in step (e0) as the average of the orientation coefficients of the theme associated with a proposition weighted by the associated coverage coefficients of the theme.

10. The method according to claim 1 , wherein step (e0) comprises for each of said subsets and for each theme, the computation of a global divergence coefficient of the theme corresponding to the standard deviation of the distribution of the products of the orientation coefficients by the coverage coefficients associated with each proposition.

11. The method according to claim 10 , wherein a semantic coefficient of a subset A for a theme T is computed in step (e) with the formula M(A,T)=relevance coefficient (A,T)*orientation coefficient (A,T)*√[1+divergence coefficient (A,T) 2 ].

12. The method according to claim 1 , wherein the subset/theme pairs selected in step (f) are those such that for any partition of the subset into a plurality of portions of said subset, the semantic coefficient of the subset for the theme is greater than the sum of the semantic coefficients of the sub-portions of the subset for the theme.

13. The method according to claim 1 , wherein groups of subset/reference theme pairs are stored on the data storage means, step (g) comprising the determination of group(s) comprising at least one subset/theme pair selected in step (f).

14. The method according to claim 13 , wherein the step (g) comprises the generation of a new group if no group of subset/reference theme pairs contains at least one subset/theme pair selected for the text.

15. The method according to claim 13 , wherein each subset/reference theme pair is associated with a score stored on the data storage means, the score of a subset/reference theme pair decreasing over time but increasing every time this subset/theme pair is selected for a text.

16. The method according to claim 15 , comprising a step (h) for suppressing a subset/reference theme pair of a group if the score of said pair passes below a first threshold, or modification on the data storage means of said plurality of lists associated with the themes if the score of said pair passes above a second threshold.

17. The method according to claim 13 , wherein step (g) comprises for each group of subset/reference theme pairs the computation of a dilution coefficient representing the number of occurrences in said text portion of reference words associated with themes of the subset/reference theme pairs present in the text relatively to the total number of reference words associated with said themes.

18. The method according to claim 1 , wherein all the subsets of the set of the words of said text portion associated with at least one theme are constructed in step (c).

19. A piece of equipment comprising data processing means configured for performing, following reception of a text in natural language, steps of:

(a) Syntactically parsing at least one text portion into a plurality of words;

(b) Filtering words of said text portion with respect to a plurality of a list of reference words stored on data storage means ( 12 ) of the piece of equipment ( 1 ), each being associated with a theme, so as to identify:

The set of the words of said text portion associated with at least one theme,

The set of the themes of said text portion;

(c) Constructing a plurality of subsets of the set of the words of said text portion associated with at least one theme;

(d) For each of said subsets and for each identified theme, computing:

a coverage coefficient of the theme and/or a relevance coefficient of the theme depending on the occurrences in said text portion, of reference words associated with the theme;

at least one orientation coefficient of the theme from the words of said text portion not belonging to the subset;

(e) For each of said subsets and for each identified theme, computing a semantic coefficient representative of a meaning level borne by a sub-group depending on said coverage, relevance and/or orientation coefficients of the theme;

(f) Selecting according to the semantic coefficients at least one subset/theme pair;

(g) Classifying the text according to said at least one selected subset/theme pair,

wherein the method comprises a preliminary step (a0) for parsing the text into a plurality of propositions, each being a text portion for which the steps (a) to (d) of the method are repeated so as to obtain for each proposition a set of coverage, relevance, and/or orientation coefficients associated with the proposition, the method comprising before step (e) a computing step (e0) for each of said subsets and for each identified theme, for at least one proposition of the text of a global coverage coefficient of the theme and/or of a global relevance coefficient of the theme, and of at least one global orientation coefficient of the theme depending on the set of said coefficients associated with a proposition.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2016
From: MALLE, JEAN-PIERRE
To: DEADIA
Reel/Frame 039739/0466 →
Priority Claims (1)
FR 1400201 · Jan 28, 2014 · national
Continuity (1)
Related Publication 20160350277A1 · Dec 1, 2016