IP Library Granted Patent US 11,640,501
Granted Patent B2
US 11,640,501 · App. 17/048,531 · Granted May 2, 2023

Method and device for verifying the author of a short message

Inventors: Guy Genilloud (Neyruz, CH); Alexandre-Pierre Cotty (Aproz, CH); Antoine Jover (Burgistein, CH); Adrien Donnet-Monay (Puidoux, CH); Florent Devillard (Lausanne, CH); Constanze Andel Rimensberger (Geneva, CH); Valentin Roten (Blonay, CH); Stefan Codrescu (Ecublens, CH); Alain Favre (Nods, CH); Luc-Olivier Pochon (Cormondrèche, CH); Lionel Pousaz (Boston, MA); Claire Roten (Vevey, CH); Stéphanie Riand (Sion, CH); Serge Nicollerat (Voluntary, RO); Myriam Eugster (Vevey, CH); Jean-Luc Buhlmann (Saint-Julien-en-Genevois, FR); Léonard Andrè Henri Studer (Villeneuve, CH); Claude-Alain Roten (Vevey, CH)
Assignee: Orphanalytics SA
G06F40/216G06F40/253G06K9/6219
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,640,501
App. No.
17/048,531
Granted
May 2, 2023
Kind
B2
Abstract

A method for verifying whether a queried text of less than 500 characters has been compiled by an author, comprising the following steps: multivariate statistical analysis of the queried text, for example, PCA or PCoA, in order to generate a matrix of coordinates in a space with N dimensions; hierarchical clustering of the points of this space that can be represented by a dendrogram; verification of the author of the queried text on the basis of this clustering.

Claims (38)

1. A method allowing to verify the authorship of a queried text of less than 500 characters, comprising the following steps:

obtaining an electronic copy of at least one queried text and at least one reference text;

normalizing the queried text and at least one reference text by means of a window-splitting module;

cutting up at least one reference text into a plurality of windows by means of said window-splitting module;

determining the number of occurrences of predefined patterns in said queried text by means of a stylistic module, said predefined patterns comprising exclusively intra and/or inter-word letter patterns, and analysing said numbers of occurrences

provide a multivariate statistical analysis of said occurrences of the queried text, in such a manner as to generate a matrix of coordinates in an N-dimensional space;

hierarchical clustering of the points of this space representable by a dendrogram by means of module for calculating stylometric distance;

verification of the authorship of the queried text on the basis of this clustering.

2. The method as claimed in claim 1 , said clustering comprising a UPGMA, Minimum Variance, WPGMA, or NJ method.

3. The method as claimed in claim 1 , comprising the establishment of a measurement of robustness of the dendrogram by means of a cophenetic correlation coefficient.

4. The method as claimed in claim 1 , comprising a step for determining whether the structure of the dendrogram is perfect, almost-perfect or nested.

5. The method as claimed in claim 1 , comprising the comparison of the queried text with texts from several authors, and the assignment of the most probable author to the queried text.

6. The method as claimed in claim 5 , comprising:

calculation of the distance of the queried text (Q) with at least two other groups of texts (A and B) from known authors;

for each pair of groups (QQ, QA, QB, AA, AB and BB), calculation of the average of the distances between the fragments of texts from the two groups of the pair, with their standard deviation;

for each group, calculation of a confidence interval, which is the distance on either side of the average which contains a given proportion of the fragments of text from this group.

7. The method as claimed in claim 5 , comprising a clustering of the fragments of queried texts into several groups of queried texts associated with several authors.

8. The method as claimed in claim 1 , said multivariate statistical analysis and/or said clustering comprising the calculation of a Boolean distance between two texts.

9. The method as claimed in claim 1 , said patterns corresponding

to trigrams; and/or

to bigrams with n intercalator letters; and/or

to bigrams at the start of words, in the middle of words or at the end of words, or to inter-word bigrams.

10. The method as claimed in claim 1 , said patterns comprising occurrences of n-gram multigrams, with or without n intercalator letters.

11. The method as claimed in claim 1 , said patterns comprising linking bigrams between two words, with or without intercalator word.

12. The method as claimed in claim 1 , wherein said

normalization of the queried text comprises eliminating the punctuation marks, replacing the upper case letters with lower case ones, and replacing the accented letters or other variations of the basic letters with the main form of the corresponding letters.

13. The method as claimed in claim 1 , wherein said

automatic cutting up of the queried text provides a plurality of windows, at least two windows intersecting, said windows being offset from one another by t characters, certain windows comprising a portion of the end of the text and a portion of the start of the text.

14. The method as claimed in claim 1 , wherein

said automatic cutting up of a reference text provides a plurality of windows, at least two windows intersecting, said windows being offset from one another by t characters, certain windows comprising a portion of the end of the text and a portion of the start of the text.

15. The method as claimed in claim 1 , said analysis being based on a measurement of distance to the barycenters.

16. The method as claimed in claim 15 , in which several queried texts are compared one after the other with texts from at least two reference authors.

17. The method as claimed in claim 1 , in which:

it is tested first of all whether a group of queried texts is far from two other groups of reference texts, from known authors, with which it is compared;

if the group of queried texts is sufficiently far from the other two reference text groups, two sub-clusters of queried texts are created starting from the group of queried texts, according to their distance to one of said reference text groups, and the difference between the average of the cophenetic distances between the fragments of each sub-cluster with a reference text group is determined in order to determine whether the two sub-clusters come or do not come from the same author.

18. A data processing storage medium comprising a computer program designed to be executed by a processor in order to cause it to execute the method as claimed in claim 1 , said data storage medium comprising a memory having a portion for the application programs comprising a window-spitting module, a module for determination stylistic parameters, a module for calculating stylistic distance and a module for identifying ruptures of styles.

19. The method of claim 1 , wherein said multivariate statistical analysis of the queried text is selected from PCA or PCoA.

20. The method of claim 1 , wherein said stylometric distance between the points resulting from the multivariate statistical analysis is an angular distance cos θ.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2023
From: GENILLOUD, GUY; COTTY, ALEXANDRE-PIERRE; JOVER, ANTOINE; DONNET-MONAY, ADRIEN; DEVILLARD, FLORENT; ANDEL RIMENSBERGER, CONSTANZE; ROTEN, VALENTIN; CODRESCU, STEFAN; FAVRE, ALAIN; POCHON, LUC-OLIVIER; POUSAZ, LIONEL; ROTEN, CLAIRE; RIAND, STÉPHANE; NICOLLERAT, SERGE; EUGSTER, MYRIAM; BUHLMANN, JEAN-LUC; STUDER, LÉONARD ANDRÉ HENRI; ROTEN, CLAUDE-ALAIN
To: ORPHANALYTICS SA
Reel/Frame 062688/0954 →
Priority Claims (2)
CH 00510/18 · Apr 20, 2018 · national
CH 00835/18 · Jul 4, 2018 · national
Continuity (1)
Related Publication 20210174017A1 · Jun 10, 2021