IP Library Granted Patent US 8,086,592
Granted Patent B2
US 8,086,592 · App. 11/948,718 · Granted Dec 27, 2011

Apparatus and method for associating unstructured text with structured data

Assignee: SAP France S.A.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,086,592
App. No.
11/948,718
Granted
Dec 27, 2011
Kind
B2
Abstract

A computer readable storage medium includes executable instructions to receive a semantic abstraction describing at least one underlying data source. The semantic abstraction includes at least one dimension with at least one dimension value. Unstructured text is parsed into parsed text units. A dimension value is matched to a parsed text unit to form matched content. An indication of the matched content is stored.

Claims (40)

1. A non-transitory computer readable storage medium, comprising executable instructions to:

receive a semantic abstraction describing at least one underlying data source, wherein the semantic abstraction includes data model objects including source dimensions and corresponding source dimension values;

form an index with the source dimensions and the corresponding source dimension values derived from the data model objects of the semantic abstraction;

parse secondary source unstructured text into a plurality of parsed text units, wherein the secondary source unstructured text is from a data source separate from the underlying data source;

match a source dimension value from the index to a parsed text unit to form matched content;

store the matched content in a table specifying a source dimension of the underlying data source, corresponding source dimension values and a count for a parsed text unit from the external source unstructured text that matches a dimension value;

create an annotation in the table indicative of the number of matches between a selected dimension value and parsed text units; and

process the table to rank the relevance of the external source unstructured text to the semantic abstraction of the underlying data source.

2. The computer readable storage medium of claim 1 wherein the secondary source unstructured text is supplied from one of a local repository and a remote repository.

3. The computer readable storage medium of claim 1 wherein the secondary source unstructured text is selected from customer complaints, screen scrapings, news feeds, text files, Really Simple Syndication (RSS) feeds, data streams and emails.

4. The computer readable storage medium of claim 1 wherein the secondary source unstructured text is provided in response to a query.

5. The computer readable storage medium of claim 1 wherein the executable instructions to parse include executable instructions to tokenize the unstructured text to form tokenized text, stem the tokenized text, and remove stop words to produce the parsed text units.

6. The computer readable storage medium of claim 1 further comprising executable instructions to create a coordinate in the table indicative of a match between a plurality of source dimension values and the parsed text unit.

7. The computer readable storage medium of claim 6 further comprising executable instructions to rank the coordinate.

8. The computer readable storage medium of claim 1 further comprising executable instructions to simultaneously display a dimension value and associated unstructured text.

9. The computer readable storage medium of claim 1 further comprising executable instructions to apply a business intelligence tool to the matched content to generate a report.

10. A method for implementation by one or more data processors, the method comprising:

receiving, by at least one processor, a semantic abstraction describing at least one data source, the semantic abstraction comprising data model objects including source dimensions and corresponding source dimension values;

forming, by at least one processor, an index with the source dimensions and the corresponding source dimension values derived from the data model objects of the semantic abstraction;

parsing, by at least one processor, unstructured text into a plurality of parsed text units;

matching, by at least one processor, a source dimension value to a parsed text unit to form matched content;

storing, by at least one processor, the matched content in a table; and

creating, by at least one processor, an annotation in the table indicative of number of matches between a selected dimension value and parsed text units.

11. The method in accordance with claim 10 , wherein the table specifies a source dimension of the data source, corresponding source dimension values and a count for a parsed text unit from the unstructured text that matches a dimension value.

12. The method in accordance with claim 11 further comprising:

processing, by at least one processor, the table to rank the relevance of the unstructured text to the semantic abstraction of the data source.

13. The method in accordance with claim 10 , wherein the unstructured text is obtained from a secondary source separate from the at least one data source.

14. The method in accordance with claim 10 further comprising:

creating, by at least one processor, a coordinate in the table indicative of a match between a plurality of source dimension values and the parsed text unit; and

ranking, by at least one processor, the coordinate.

15. A method to associate unstructured text with structured data, the method being implemented by one or more data processors and comprising:

receiving, by at least one processor, structured data from at least one first data source, the structured data comprising source dimensions and corresponding source dimension values;

forming, by at least one processor, an index with the source dimensions and the corresponding source dimension values derived from the structured data;

receiving, by at least one processor, the unstructured text from a second data source;

parsing, by at least one processor, the unstructured text into a plurality of parsed text units;

matching, by at least one processor, a source dimension value of the source dimension values to a parsed text unit of the plurality of parsed text units to form matched content;

creating, by at least one processor, an annotation in a table indicative of number of matches between a selected dimension value and parsed text units;

processing, by at least one processor the table, to rank the matched content, the rank indicating relevance of the associating of the parsed text unit with the corresponding source dimension value.

16. The method in accordance with claim 15 , wherein the table stores the matched content and specifies a count for a parsed text unit that matches a dimension value.

17. The method in accordance with claim 15 , wherein the parsing comprises tokenizing the unstructured text to form tokenized text, stemming the tokenized text, and removing stop words to produce the parsed text units.

Assignments (2)
CHANGE OF NAME Recorded Jul 12, 2011
From: BUSINESS OBJECTS, S.A.
To: SAP FRANCE S.A.
Reel/Frame 026581/0190 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2008
From: MION, GILLES VERGNORY; CRAS, JEAN-YVES
To: BUSINESS OBJECTS, S.A.
Reel/Frame 020504/0907 →
Continuity (1)
Related Publication 20090144295A1 · Jun 4, 2009