IP Library Granted Patent US 12664363
Granted Patent B2
US 12664363 · App. 18/583,186 · Granted Jun 23, 2026

Method and system for evaluating non-fiction narrative text documents

Inventors: Sachin Sharad Pawar (Pune, IN); Girish Keshav Palshikar (Pune, IN); Ankita Jain (Indore, IN); Mahesh Prasad Singh (Noida, IN); Mahesh Rangarajan (Bangalore, IN); Aman Agarwal (Noida, IN); Kumar Karan Singh (Noida, IN); Hetal Jani (Gandhinagar, IN); Vishal Kumar (Noida, IN)
Assignee: TATA CONSULTANCY SERVICES LIMITED
G06F40/253G06F40/279
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664363
App. No.
18/583,186
Granted
Jun 23, 2026
Kind
B2
Abstract

Existing approaches for processing and evaluation of documents containing non-fiction narrative texts have the disadvantage that they are comparatively less studied in linguistics, and hence do not provide sufficient data required for evaluations. Method and system are for evaluating non-fiction narrative text documents are provided. The system processes a plurality of non-fiction narrative text documents and computes a plurality of corpus statistics. The plurality of corpus statistics is then used for evaluation of any non-fiction narrative text document that may or may not be collected as real-time input.

Claims (57)

1 . A processor implemented method for evaluating non-fiction narrative text document, comprising:

receiving, via one or more hardware processors, a plurality of non-fiction narrative text documents as input, wherein the plurality of non-fiction narrative text documents are collected as real-time input;

identifying, via the one or more hardware processors, one or more categories associated with each of a plurality of sentences in each of the plurality of non-fiction narrative text documents by building a classifier, wherein the classifier is built by representations of a text of each sentence and representations for each word in each sentence;

extracting, via the one or more hardware processors, a plurality of entities associated with a plurality of knowledge markers from the plurality of non-fiction narrative text documents, wherein the plurality of knowledge markers is specific to contents of each of the narrative text documents including skills, tasks, roles, and domain-specific concepts, wherein mentions of the skill are extracted by performing a lookup in a large gazette of known skill names, mentions of the tasks are extracted using linguistic rules, mentions of the roles are extracted by performing lookup in a gazette, and domain specific concepts are extracted upon computing domain relevance scores for all noun phrases and selecting only noun phrases that are above a threshold;

calculating, via the one or more hardware processors, one or more knowledge quality metrics for each of the plurality of non-fiction narrative text documents based on the knowledge markers and the identified one or more categories;

calculating, via the one or more hardware processors, a relative position of sentences of the one or more categories, in each of the plurality of non-fiction narrative text documents;

computing, via the one or more hardware processors, a plurality of corpus statistics for each of the one or more knowledge quality metrics and the relative position of the sentences of the one or more categories, wherein the plurality of corpus statistics is used to evaluate a non-fiction narrative text document received as an input document, wherein evaluating the non-fiction narrative text document comprises:

obtaining the non-fiction narrative text document as input;

processing the non-fiction narrative text document using the plurality of corpus statistics, comprising:

identifying one or more categories of each sentence in the input document;

extracting a plurality of entities associated with a plurality of knowledge markers from the non-fiction narrative text document;

calculating one or more knowledge quality metrics of the non-fiction narrative text document based on the plurality of knowledge markers and the identified one or more categories;

determining a quality of the non-fiction narrative text document based on comparison of the one or more knowledge quality metrics of the non-fiction narrative text document with the plurality of corpus statistics;

calculating the relative position of sentences of the one or more categories, in the input document;

computing a sentence categories flow metric of the non-fiction narrative text document based on a determined sentence categories flow in the plurality of non-fiction narrative text documents, by comparing with the corpus statistics of the relative position of sentences of the one or more categories;

generating at least one suggestion for improvement of the non-fiction narrative text document automatically, over the determined quality, based on the sentence categories flow metric and the one or more knowledge quality metrics; and

recommending, via the one or more hardware processors, the generated at least one suggestion to a user and displayed via a suitable user interface.

2 . The processor implemented method of claim 1 , wherein the plurality of non-fiction narrative documents are documents that are reviewed and re-written incorporating one or more expert suggestions, wherein the one or more expert suggestions are further reviewed by one or more reviewers for knowledge contents, readability, narration flow and grammar in iterative manner.

3 . A system for evaluating non-fiction narrative text document, comprising:

one or more hardware processors;

a communication interface; and

a memory storing a plurality of instructions, wherein the plurality of instructions cause the one or more hardware processors to:

receive a plurality of non-fiction narrative text documents, wherein the plurality of non-fiction narrative text documents are collected as real-time input;

identify one or more categories associated with each of a plurality of sentences in each of the plurality of non-fiction narrative text documents by building a classifier, wherein the classifier is built by representations of a text of each sentence, and representations for each word in each sentence;

extract a plurality of entities associated with a plurality of knowledge markers from the plurality of non-fiction narrative text documents, wherein the plurality of knowledge markers is specific to contents of each of the narrative text documents including skills, tasks, roles, and domain-specific concepts, wherein mentions of the skill are extracted by performing a lookup in a large gazette of known skill names, mentions of the tasks are extracted using linguistic rules, mentions of the roles are extracted by performing lookup in a gazette, and domain specific concepts are extracted upon computing domain relevance scores for all noun phrases and selecting only noun phrases that are above a threshold;

calculate one or more knowledge quality metrics for each of the plurality of non-fiction narrative text documents based on the knowledge markers and the identified one or more categories;

calculate the relative position of sentences of the one or more categories, in each of the plurality of non-fiction narrative text documents; and

compute a plurality of corpus statistics for each of the one or more knowledge quality metrics and the relative position of the sentences of the one or more categories, wherein the one or more hardware processors are configured to use the plurality of corpus statistics to evaluate a non-fiction narrative text document received as an input document, by:

obtaining the non-fiction narrative text document as input; and

processing the non-fiction narrative text document using the plurality of corpus statistics, comprising:

identifying one or more categories of each sentence in the input document;

extracting a plurality of entities associated with a plurality of knowledge markers from the non-fiction narrative text document;

calculating one or more knowledge quality metrics of the non-fiction narrative text document based on the plurality of knowledge markers and the identified one or more categories;

determining a quality of the non-fiction narrative text document based on comparison of the one or more knowledge quality metrics of the non-fiction narrative text document with the plurality of corpus statistics;

calculating the relative position of sentences of the one or more categories, in the input document;

computing a sentence categories flow metric of the non-fiction narrative text document based on a determined sentence categories flow in the plurality of non-fiction narrative text documents, by comparing with the corpus statistics of the relative position of sentences of the one or more categories;

generating at least one suggestion for improvement of the non-fiction narrative text document automatically, over the determined quality, based on the sentence categories flow metric and the one or more knowledge quality metrics; and

recommend the generated at least one suggestion to a user and displayed via a suitable user interface.

4 . The system of claim 3 , wherein the one or more hardware processors are configured to use documents that are reviewed and re-written incorporating one or more expert suggestions, as the plurality of non-fiction narrative, wherein the one or more expert suggestions are further reviewed by one or more reviewers for knowledge contents, readability, narration flow and grammar in iterative manner.

5 . One or more non-transitory machine-readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors cause:

receiving a plurality of non-fiction narrative text documents wherein the plurality of non-fiction narrative text documents are collected as real-time input;

identifying one or more categories associated with each of a plurality of sentences in each of the plurality of non-fiction narrative text documents by building a classifier, wherein the classifier is built by representations of a text of each sentence, and representations for each word in each sentence;

extracting a plurality of entities associated with a plurality of knowledge markers from the plurality of non-fiction narrative text documents, wherein the knowledge markers is specific to contents of each of the narrative text documents including skills, tasks, roles, and domain-specific concepts, wherein mentions of the skill are extracted by performing a lookup in a large gazette of known skill names, mentions of the tasks are extracted using linguistic rules, mentions of the roles are extracted by performing lookup in a gazette, and domain specific concepts are extracted upon computing domain relevance scores for all noun phrases and selecting only noun phrases that are above a threshold;

calculating one or more knowledge quality metrics for each of the plurality of non-fiction narrative text documents based on the knowledge markers and the identified one or more categories;

calculating the relative position of sentences of the one or more categories, in each of the plurality of non-fiction narrative text documents; and

computing a plurality of corpus statistics for each of the one or more knowledge quality metrics and the relative position of the sentences of the one or more categories, wherein the plurality of corpus statistics is used to evaluate a non-fiction narrative text document received as an input document, wherein evaluating the non-fiction narrative text document comprises:

obtaining the non-fiction narrative text document as input; and

processing the non-fiction narrative text document using the plurality of corpus statistics, comprising:

identifying one or more categories of each sentence in the input document;

extracting a plurality of entities associated with a plurality of knowledge markers from the non-fiction narrative text document;

calculating one or more knowledge quality metrics of the non-fiction narrative text document based on the plurality of knowledge markers and the identified one or more categories;

determining a quality of the non-fiction narrative text document based on comparison of the one or more knowledge quality metrics of the non-fiction narrative text document with the plurality of corpus statistics;

calculating the relative position of sentences of the one or more categories, in the input document;

computing a sentence categories flow metric of the non-fiction narrative text document based on a determined sentence categories flow in the plurality of non-fiction narrative text documents, by comparing with the corpus statistics of the relative position of sentences of the one or more categories;

generating at least one suggestion for improvement of the non-fiction narrative text document automatically, over the determined quality, based on the sentence categories flow metric and the one or more knowledge quality metrics; and

recommending the generated at least one suggestion to a user and displayed via a suitable user interface.

6 . The one or more non-transitory machine-readable information storage mediums of claim 5 , wherein the plurality of non-fiction narrative documents are documents that are reviewed and re-written incorporating one or more expert suggestions, wherein the one or more expert suggestions are further reviewed by one or more reviewers for knowledge contents, readability, narration flow and other aspects like grammar in iterative manner.