IP Library Granted Patent US 11,036,939
Granted Patent B2
US 11,036,939 · App. 16/381,948 · Granted Jun 15, 2021

Data driven approach for automatically generating a natural language processing cartridge

Inventors: Mario J. Lorenzo (Miami, FL); Jennifer Lynn La Rocca (Cary, NC); Rebecca Lynn Dahlman (Rochester, MN); Kristin E. McNeil (Charlotte, NC)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F40/30G06F40/242G06F40/295G06N5/02G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,036,939
App. No.
16/381,948
Granted
Jun 15, 2021
Kind
B2
Abstract

An artifact identification engine identifies artifacts from structured and unstructured data in one or more documents based on pre-defined artifacts, by using cognitive annotations. The identified artifacts are analyzed based at least on received inputs. A cartridge that includes artifacts that are relevant to the structured and unstructured data is generated, based on the analyzing.

Claims (40)

1. A method, comprising:

identifying, via an artifact identification engine, artifacts from structured and unstructured data in one or more documents based on pre-defined artifacts, by using cognitive annotations; and

analyzing the identified artifacts, based at least on received inputs; and

generating a cartridge that includes artifacts that are targeted for processing the structured and unstructured data, based on the analyzing, wherein the identified artifacts exceed a frequency threshold of occurrence in the one or more documents, wherein the received inputs include a predetermined threshold in matching for entities within an artifact, wherein if entities within an identified artifact exceed the predetermined threshold in matching, then the identified artifact is added to the cartridge, and wherein the cartridge generates a cognitive model for the one or more documents.

2. The method of claim 1 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs.

3. The method of claim 1 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs, and wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

4. A method, comprising:

identifying, via an artifact identification engine, artifacts from structured and unstructured data in one or more documents based on pre-defined artifacts, by using cognitive annotations; and

analyzing the identified artifacts, based at least on received inputs; and

generating a cartridge that includes artifacts that are targeted for processing the structured and unstructured data, based on the analyzing, wherein the identified artifacts exceed a frequency threshold of occurrence in the one or more documents, wherein the received inputs include a predetermined threshold in matching for entities within an artifact, wherein if entities within an identified artifact do not exceed the predetermined threshold in matching, then a subset of the entities of the identified artifact is added to the cartridge, and wherein the cartridge generates a cognitive model for the one or more documents.

5. The method of claim 4 , wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

6. The method of claim 4 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs, and wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

7. A system, comprising:

a memory; and

a processor coupled to the memory, wherein the processor performs operations, the operations comprising:

identifying, via an artifact identification engine, artifacts from structured and unstructured data in one or more documents based on pre-defined artifacts, by using cognitive annotations; and

analyzing the identified artifacts, based at least on received inputs; and

generate a cartridge that includes artifacts that are targeted for processing the structured and unstructured data, based on the analyzing, wherein the identified artifacts exceed a frequency threshold of occurrence in the one or more documents, wherein the received inputs include a predetermined threshold in matching for entities within an artifact, wherein if entities within an identified artifact exceed the predetermined threshold in matching, then the identified artifact is added to the cartridge, and wherein the cartridge generates a cognitive model for the one or more documents.

8. The system of claim 7 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs.

9. The system of claim 8 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs, and wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

10. The system of claim 7 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs, and wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

11. A system, comprising:

a memory; and

a processor coupled to the memory, wherein the processor performs operations, the operations comprising:

identifying, via an artifact identification engine, artifacts from structured and unstructured data in one or more documents based on pre-defined artifacts, by using cognitive annotations; and

analyzing the identified artifacts, based at least on received inputs; and

generating a cartridge that includes artifacts that are targeted for processing the structured and unstructured data, based on the analyzing, wherein the identified artifacts exceed a frequency threshold of occurrence in the one or more documents, wherein the received inputs include a predetermined threshold in matching for entities within an artifact, wherein if entities within an identified artifact do not exceed the predetermined threshold in matching, then a subset of the entities of the identified artifact is added to the cartridge, and wherein the cartridge generates a cognitive model for the one or more documents.

12. The system of claim 11 , wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

13. A computer program product, the computer program product comprising a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code configured to perform operations, the operations comprising:

identifying, via an artifact identification engine, artifacts from structured and unstructured data in one or more documents based on pre-defined artifacts, by using cognitive annotations; and

analyzing the identified artifacts, based at least on received inputs; and

generate a cartridge that includes artifacts that are targeted for processing the structured and unstructured data, based on the analyzing, wherein the identified artifacts exceed a frequency threshold of occurrence in the one or more documents, wherein the received inputs include a predetermined threshold in matching for entities within an artifact, wherein if entities within an identified artifact exceed the predetermined threshold in matching, then the identified artifact is added to the cartridge, and wherein the cartridge generates a cognitive model for the one or more documents.

14. The computer program product of claim 13 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs, and wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

15. The computer program product of claim 14 , wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

16. The computer program product of claim 14 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs, and wherein the cognitive annotations are generated via a natural language processing software that process the structured and unstructured data in the one or more documents based on the pre-defined artifacts.

17. The computer program product of claim 13 , wherein the received inputs are used to generate filter artifacts based on concepts that are identified as not interesting in the received inputs.

18. A computer program product, the computer program product comprising a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code configured to perform operations, the operations comprising:

identifying, via an artifact identification engine, artifacts from structured and unstructured data in one or more documents based on pre-defined artifacts, by using cognitive annotations; and

analyzing the identified artifacts, based at least on received inputs; and

generate a cartridge that includes artifacts that are targeted for processing the structured and unstructured data, based on the analyzing, wherein the identified artifacts exceed a frequency threshold of occurrence in the one or more documents, wherein the received inputs include a predetermined threshold in matching for entities within an artifact, wherein if entities within an identified artifact do not exceed the predetermined threshold in matching, then a subset of the entities of the identified artifact is added to the cartridge, and wherein the cartridge generates a cognitive model for the one or more documents.

Assignments (3)
SECURITY INTEREST Recorded Oct 1, 2025
From: MERATIVE US L.P.; MERGE HEALTHCARE INCORPORATED
To: TCG SENIOR FUNDING L.L.C., AS COLLATERAL AGENT
Reel/Frame 072808/0442 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MERATIVE US L.P.
Reel/Frame 061496/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2019
From: LORENZO, MARIO J.; LA ROCCA, JENNIFER LYNN; DAHLMAN, REBECCA LYNN; MCNEIL, KRISTIN E.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048879/0083 →
Continuity (1)
Related Publication 20200327195A1 · Oct 15, 2020