IP Library Patent Application 18993171
Patent Application
App. No. 18/993,171

Visual Structure of Documents in Question Answering

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/993,171
Abstract

A question-answering system that receive a natural-language question includes a database to provide a basis for that answer and a structured-query generator that constructs a structured query from the question and uses it to obtain an answer to the question from the database.

Claims (77)

1 . A method comprising:

constructing a knowledge base for a natural-language question-answering system,

wherein constructing said knowledge base for said natural-language question-and-answering system comprises:

ingesting a document that comprises first, second, and third visual segments, each of which comprises semantic content, and

pre-processing said document to enable said question-answering system to access said semantic content in response to a natural-language question that has been posed by a user of said natural-language question-answering system,

wherein pre-processing said document comprises:

using a visual structure of said document to extract structural information from said document, including

classifying said visual segments into different classes based on differences in visual appearances of said visual segments,

determining, based at least in part on classes of said first and second visual segments, that a context relationship is to exist between said first and second sematic segments,

determining, based at least in part on classes of said first and third visual segments, that no context relationship is to exist between said first and second sematic segments, and

establishing said context relationship between said first and second visual segments; and

incorporating said structural information in said knowledge base,

wherein said visual structure of said document comprises a spatial distribution of said visual segments in said document when said document is in a form that renders said document visible.

2 . (canceled)

3 . The method of claim 1 , wherein using said visual structure to extract structural information from said document comprises:

based at least in part on a distance between said first and second visual segments, determining that a context relationship is to exist between said first and second visual segments,

based at least in part on a distance between said first and third visual segments, determining that no context relationship is to exist between said first and third visual segments, and

establishing said context relationship between said first and second visual segments.

4 . The method of claim 1 , further comprising using semantic information in addition to structural information to determine whether a semantic relationship should exist between a pair of visual segments.

5 . The method of claim 1 , further comprising, based at least in part on a reference in said first visual segment to said second visual segment, determining that a context relationship is to exist between said first visual segment and said second visual segment,

based at least in part on an absence of a reference in said first visual segment to said third visual segment, determining that no context relationship is to exist between said first visual segment and said third visual segment, and

establishing said context relationship between said first visual segment and said second visual segment.

6 . (canceled)

7 . The method of claim 1 , wherein using said visual structure to extract structural information from said document comprises:

based at least in part on locations of said first and second visual segments relative to each other within said document and on semantic content of said first and second visual segments, determining that a context relationship is to exist between said first and second visual segments,

based at least in part on locations of said first and third visual segments relative to each other within said document and on semantic content of said first and second visual segments, determining that no context relationship is to exist between said first and third visual segments, and

establishing said context relationship between said first and second visual segments.

8 . The method of claim 1 wherein using said visual structure to extract structural information from said document comprises, based on having determined that a first visual segment is a figure and that a second visual segment is a caption for said figure, establishing a context relationship between said first and second visual segments.

9 . The method of claim 1 , wherein using said visual structure to extract structural information from said document comprises, based on having determined that said first and second visual segments have a common location, establishing a context relationship between said first and second visual segments.

10 . The method of claim 1 , wherein using said visual structure to extract structural information from said document comprises, based on having determined that a first visual segment is a figure and that a second visual segment is text that is superimposed on said figure, establishing a context relationship between said first and second visual segments.

11 . The method of claim 1 , wherein said document comprises instructions for causing said document to be rendered visible and wherein said method further comprises determining said visual structure based on said instructions.

12 . The method of claim 1 , wherein said document is in a portable document format and wherein said method further comprises determining said visual structure of said document based on rendering instructions that are expressed in said portable document format.

13 . The method of claim 1 , wherein said document is an HTML file and wherein said method further comprises determining said visual structure of said document based on tags in said HTML file.

14 . The method of claim 1 , further comprising determining that a first of said visual segments comprises text arranged in rows and columns, determining that a second of said visual segments includes text at an intersection of one of said rows and one of said columns, and establishing a context relationship between said first and second visual segments based at least in part on said determinations.

15 . The method of claim 1 , wherein using said visual structure to extract structural information from said document comprises tagging at least one of said visual segments as belonging to each class in a set of classes, said set comprising a list, a paragraph, a heading, and an image.

16 . The method of claim 1 , wherein said first visual segment comprises audio, wherein said second visual segment comprises a transcript of said audio, and wherein said method further comprises establishing a context relationship between said first and second visual segments.

17 . The method of claim 1 , wherein said first visual segment comprises audio, wherein said second visual segment comprises video, and wherein said method further comprises establishing a context relationship between said first and second visual segments to synchronize said audio and said video.

18 . A natural-language question-and-answering system for providing an answer to a question from a user, said natural-language question-and-answering system comprising:

an ingestor for ingesting documents, each of which, when rendered visible, has a visual structure in which visual segments, each of which comprises semantic content, are spatially distributed throughout said document,

a visual parser for extracting structural information from said documents based at least in part on said visual structures thereof, and

a knowledge base that incorporates semantic content from said documents and said structural information.

19 . The natural-language question-and-answering system of claim 18 , wherein said visual parser comprises:

a segmenting circuit configured to identify said visual segments and to classify said visual segments into classes,

an interpretation circuit to extract semantic information from said visual segments, and

a linkage circuit to establish context relationships between selected pairs of visual segments based on input from said segmenting circuit and said interpretation circuit.

20 . A non-transitory machine-readable comprising instructions stored thereon, said instructions when executed by a data processing system cause said system to perform:

constructing a knowledge base for a natural-language question-answering system,

wherein constructing said knowledge base for said natural-language question-and-answering system comprises:

ingesting a document that comprises first, second, and third visual segments, each of which comprises semantic content, and

pre-processing said document to enable said question-answering system to access said semantic content in response to a natural-language question that has been posed by a user of said natural-language question-answering system,

wherein pre-processing said document comprises:

using a visual structure of said document to extract structural information from said document, including

classifying said visual segments into different classes based on differences in visual appearances of said visual segments,

determining, based at least in part on classes of said first and second visual segments, that a context relationship is to exist between said first and second sematic segments,

determining, based at least in part on classes of said first and third visual segments, that no context relationship is to exist between said first and second sematic segments, and

establishing said context relationship between said first and second visual segments; and

incorporating said structural information in said knowledge base,

wherein said visual structure of said document comprises a spatial distribution of said visual segments in said document when said document is in a form that renders said document visible.

21 . The method of claim 1 , further comprising:

receiving a natural language question from a user;

processing the question to determine an answer by accessing the knowledge base; and

providing the answer to the user;

wherein the processing include determining the answer based a combination of semantic information in the first visual segment and semantic information in the second visual segment determined to have a semantic relationship with the first visual segment.

22 . The method of claim 21 ,

wherein storing the documents and the context relationships comprises for forming composite content of segments and content of segments with contextual relationships, including forming a first composite content comprising content of the first visual segment and content of the second visual segment; and

wherein processing the question comprises determining the answer by processing the question with composite content formed for a plurality of the visual segments, including processing the question and the first composite content to determine the answer.

23 . A method comprising:

constructing a knowledge base for a natural-language question-answering system,

wherein constructing said knowledge base for said natural-language question-and-answering system comprises:

ingesting a document that comprises first, second, and third visual segments, each of which comprises semantic content, and

pre-processing said document to enable said question-answering system to access said semantic content in response to a natural-language question that has been posed by a user of said natural-language question-answering system,

wherein pre-processing said document comprises:

using a visual structure of said document to extract structural information from said document, wherein said visual structure of said document comprises a spatial distribution of said visual segments in said document when said document is in a form that renders said document visible;

incorporating said structural information in said knowledge base;

classifying the visual segments into corresponding classes of said segments based on differences in visual appearances of said visual segments when the document is rendered;

determining context relationships between the visual segments of said document according to the classes of said segments and locations of said segments in the visual structure of the document, including determined a context relationship between the first visual segment and the second visual segment based on locations of the first and second visual segments and the classes of said segments; and

storing the documents and the context relationships of the visual segments of said documents in the knowledge base for use natural-language question-answering.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Oct 31, 2025
From: PRYON INCORPORATED
To: FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 073438/0899 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2025
From: MARCHERET, ETIENNE; NAHAMOO, DAVID; O'DONNELL, DOMINIQUE; GOEL, VAIBHAVA; SUNG, CHUL; RENNIE, STEVEN JOHN; JABLOKOV, IGOR RODITIS; ATIVANICHAYAPHONG, SOONTHORN; ZADBUKE, AJINKYA JITENDRA; ROTHBERG, CARMI; METEER, MARIE WENZEL; KISLAL, ELLEN EIDE; PRUITT, JOHN MICHAEL
To: PRYON INCORPORATED
Reel/Frame 070892/0696 →