IP Library Granted Patent US 8,744,835
Granted Patent B2
US 8,744,835 · App. 10/157,894 · Granted Jun 3, 2014

Content conversion method and apparatus

Inventor: Eli Abir (Cross River, NY)
Assignee: Meaningful Machines LLC
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,744,835
App. No.
10/157,894
Granted
Jun 3, 2014
Kind
B2
Abstract

A method and apparatus for analyzing documents and thereby determining the association between words in a language. The method includes providing a collection of documents, selecting a first word or word string, and a second word or word string occurring in the documents. The method further involves associating first word or word strings and second word or word strings with common word or word strings based on frequency of occurrence of the common word or word strings within the ranges.

Claims (28)

1. A non-transitory, computer readable storage medium storing computer instructions which when executed cause a processor to:

search at least one document for a query word or query word string, said at least one document being comprised of unstructured text;

identify occurrences of said query word or said query word string in said at least one document;

define, as user-defined first query ranges, words or word strings immediately preceding the occurrences of said query word or said query word string;

define, as user-defined second query ranges, words or word strings immediately following the occurrences of said query word or said query word string;

select one or more recurring words or word strings in the defined first query ranges as identified left contexts for said query word or said query word string;

select one or more recurring words or word strings in the defined second query ranges as identified right contexts for said query word or said query word string;

search said at least one document for one or more occurrences of said identified left contexts;

search said at least one document for one or more occurrences of said identified right contexts;

define, as first context ranges, words or word strings immediately following the occurrences of the identified left contexts;

define, as second context ranges, words or word strings immediately preceding the occurrences of the identified right contexts;

identify words and word strings that occur in both of at least one of said first context ranges and one of said second context ranges, as associated second words or word strings; and

determine an association between the associated second words or word strings and the query word or the query word string.

2. The non-transitory, computer readable storage medium according to claim 1 , wherein the query word or the query word string and the one or more recurring words or word strings are associated as being semantically similar or equivalent or of a common class or category.

3. A method for associating words and word strings in a language, the method comprising:

searching, using a processor, at least one document for a query word or query word string, said at least one document being comprised of unstructured text;

identifying occurrences of said query word or said query word string in said at least one document;

defining, as user-defined first query ranges, words or word strings immediately preceding the occurrences of said query word or said query word string;

defining, as user-defined second query ranges, words or word strings immediately following the occurrences of said query word or said query word string;

selecting one or more recurring words or word strings in the defined first query ranges as identified left contexts for said query word or said query word string;

selecting one or more recurring words or word strings in the defined second query ranges as identified right contexts for said query word or said query word string;

searching said at least one document for one or more occurrences of said identified left contexts;

searching said at least one document for one or more occurrences of said identified right contexts;

defining, as first context ranges, words or word strings immediately following the occurrences of the identified left contexts;

defining, as second context ranges, words or word strings immediately preceding the occurrences of the identified right contexts;

identifying words and word strings that occur in both of at least one of said first context ranges and one of said second context ranges, as associated second words or word strings; and

determining an association between the associated second words or word strings and the query word or the query word string.

4. The method according to claim 3 , wherein the query word or the query word string and the one or more recurring words or word strings are associated as being semantically similar or equivalent or of a common class or category.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2003
From: ABIR, ELI
To: MEANINGFUL MACHINES LLC
Reel/Frame 014452/0711 →
Continuity (4)
Continuation In Part 10024473 · Dec 21, 2001
Provisional Application 60276107 · Mar 16, 2001
Provisional Application 60299472 · Jun 21, 2001
Related Publication 20030061025A1 · Mar 27, 2003