IP Library Granted Patent US 11,893,136
Granted Patent B2
US 11,893,136 · App. 17/460,092 · Granted Feb 6, 2024

Token-based data security systems and methods with cross-referencing tokens in freeform text within structured document

Inventor: Walter Hughes Lindsay (Phoenix, AZ)
Assignee: OPEN TEXT HOLDINGS, INC.
G06F21/6254G06F16/93G06F21/6218G06F21/6227G06F40/103G06F40/166G06F40/284G06F2221/2141
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,893,136
App. No.
17/460,092
Granted
Feb 6, 2024
Kind
B2
Abstract

Multiple types of tokens can be generated and utilized in a highly structured document with freeform text. For example, a tokenization system may receive a request for tokenizing a document with a first portion having structured content and a second portion having unstructured or semi-structured content. In response, the tokenization system identifies sensitive information in the first portion of the document, generates format-preserving tokens for the sensitive information in the first portion of the document, identifies sensitive information in the second portion of the document, and generates self-describing tokens for the sensitive information in the second portion of the document. The self-describing tokens reference the sensitive information in the first portion of the document. The tokenization system may then communicate the format-preserving tokens and the self-describing tokens to the first client computing system or to a second client computing system.

Claims (43)

1. A method for securing data, the method comprising:

receiving, by a tokenization system from a first client computing system, a request for tokenizing a document with a first portion having structured content and a second portion having unstructured or semi-structured content;

identifying, by the tokenization system, sensitive information in the first portion of the document;

generating, by the tokenization system, format-preserving tokens for the sensitive information in the first portion of the document;

identifying, by the tokenization system, the sensitive information in the second portion of the document;

generating, by the tokenization system, self-describing tokens for the sensitive information in the second portion of the document, the self-describing tokens referencing the sensitive information in the first portion of the document such that the sensitive information in the first portion of the document is secured utilizing the format-preserving tokens and the sensitive information in the second portion of the document is secured utilizing the self-describing tokens; and

communicating, by the tokenization system, the format-preserving tokens and the self-describing tokens to the first client computing system or to a second client computing system.

2. The method according to claim 1 , wherein a format-preserving token has a one-to-one connection to the sensitive information in the structured content.

3. The method according to claim 1 , wherein a self-describing token contains a protection strategy that specifies a technique for generating or formatting a surrogate for an actual value and for mapping between the surrogate and the actual value.

4. The method according to claim 1 , wherein each of the self-describing tokens has a preconfigured pattern and a token value.

5. The method according to claim 1 , wherein each of the self-describing tokens has a preconfigured pattern, an indication of a protection strategy, and a token value.

6. The method according to claim 1 , further comprising:

marking a self-describing token in the second portion of the document with at least one visual marker in a human-readable form.

7. The method according to claim 1 , wherein the first portion of the document comprises a data structure having data fields and wherein the second portion of the document comprises freeform text in one of the data fields.

8. A tokenization system for securing data, the tokenization system comprising:

a processor;

a non-transitory computer-readable medium; and

stored instructions translatable by the processor for:

receiving, from a first client computing system, a request for tokenizing a document with a first portion having structured content and a second portion having unstructured or semi-structured content;

identifying sensitive information in the first portion of the document;

generating format-preserving tokens for the sensitive information in the first portion of the document;

identifying the sensitive information in the second portion of the document;

generating self-describing tokens for the sensitive information in the second portion of the document, the self-describing tokens referencing the sensitive information in the first portion of the document such that the sensitive information in the first portion of the document is secured utilizing the format-preserving tokens and the sensitive information in the second portion of the document is secured utilizing the self-describing tokens; and

communicating the format-preserving tokens and the self-describing tokens to the first client computing system or to a second client computing system.

9. The tokenization system of claim 8 , wherein a format-preserving token has a one-to-one connection to the sensitive information in the structured content.

10. The tokenization system of claim 8 , wherein a self-describing token contains a protection strategy that specifies a technique for generating or formatting a surrogate for an actual value and for mapping between the surrogate and the actual value.

11. The tokenization system of claim 8 , wherein each of the self-describing tokens has a preconfigured pattern and a token value.

12. The tokenization system of claim 8 , wherein each of the self-describing tokens has a preconfigured pattern, an indication of a protection strategy, and a token value.

13. The tokenization system of claim 8 , wherein the stored instructions are further translatable by the processor for:

marking a self-describing token in the second portion of the document with at least one visual marker in a human-readable form.

14. The tokenization system of claim 8 , wherein the first portion of the document comprises a data structure having data fields and wherein the second portion of the document comprises freeform text in one of the data fields.

15. A computer program product for tokenization, the computer program product comprising a non-transitory computer-readable medium storing instructions translatable by a processor of a tokenization system for:

receiving, from a first client computing system, a request for tokenizing a document with a first portion having structured content and a second portion having unstructured or semi-structured content;

identifying sensitive information in the first portion of the document;

generating format-preserving tokens for the sensitive information in the first portion of the document;

identifying the sensitive information in the second portion of the document;

generating self-describing tokens for the sensitive information in the second portion of the document, the self-describing tokens referencing the sensitive information in the first portion of the document such that the sensitive information in the first portion of the document is secured utilizing the format-preserving tokens and the sensitive information in the second portion of the document is secured utilizing the self-describing tokens; and

communicating the format-preserving tokens and the self-describing tokens to the first client computing system or to a second client computing system.

16. The computer program product of claim 15 , wherein a format-preserving token has a one-to-one connection to the sensitive information in the structured content.

17. The computer program product of claim 15 , wherein a self-describing token contains a protection strategy that specifies a technique for generating or formatting a surrogate for an actual value and for mapping between the surrogate and the actual value.

18. The computer program product of claim 15 , wherein each of the self-describing tokens has a preconfigured pattern and a token value.

19. The computer program product of claim 15 , wherein each of the self-describing tokens has a preconfigured pattern, an indication of a protection strategy, and a token value.

20. The computer program product of claim 15 , wherein the first portion of the document comprises a data structure having data fields and wherein the second portion of the document comprises freeform text in one of the data fields.

Assignments (2)
MERGER Recorded Jun 23, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 075054/0783 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2021
From: LINDSAY, WALTER HUGHES
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 057756/0198 →
Continuity (2)
Provisional Application 63071618 · Aug 28, 2020
Related Publication 20220067207A1 · Mar 3, 2022
Cited By (2)
US 12,292,999 US 12,561,476