IP Library Granted Patent US 12,548,365
Granted Patent B2
US 12,548,365 · App. 17/992,434 · Granted Feb 10, 2026

Detecting object burn-in on documents in a document management system

Inventors: Taiwo Raphael Alabi (Berkeley, CA); Ashwath Saran Mohan (San Ramon, CA); Nipun Dureja (Seattle, WA); Jerome Levadoux (San Mateo, CA)
Assignee: Docusign, Inc.
G06V30/418G06V30/153G06V30/19093
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,365
App. No.
17/992,434
Granted
Feb 10, 2026
Kind
B2
Abstract

A document management system surfaces changes to a portable document format (PDF) document to a user. The document management system converts each page of the PDF document into images, segments those images, and processes each segment of those images using computer vision and/or natural language processing. The document management system compares segments from an original copy of the PDF document with segments from a modified copy of the PDF document to identify significant changes to the PDF document.

Claims (75)

1 . A method for surfacing changes in secure electronic documents comprising:

receiving, using at least one processor, an original copy of a portable document format (PDF) document;

receiving, using the at least one processor, a modified copy of the PDF document;

rasterizing, using the at least one processor, both the original copy and the modified copy of the PDF document into a plurality of images, each image of the plurality of images representing a page, wherein at least one image in the plurality of images includes at least one burned-in object; and

segmenting, using the at least one processor, each of the plurality of images into a plurality of segments, the segmenting includes

determining a height and a buffer of a predetermined number of pixels for an image of each segment in the plurality of segments for expanding each segment;

expanding two or more adjacent segments using the height and the buffer of predetermined number of pixels of each of the two or more adjacent segments to generate at least two overlapping segments that overlap by the predetermined number of pixels;

determining at least one repeated portion in the original and modified copies of the PDF document in the at least two overlapping segments; and

selecting a segment in the at least two overlapping segments in the original and modified copies of the PDF document representative of the at least one repeated portion;

for corresponding selected segments derived from the original copy and the modified copy of the PDF document:

generating, using the at least one processor, a representation of text in the corresponding selected segments;

determining, using the at least one processor, similarity characteristics between the corresponding selected segments;

determining, using the at least one processor, based on the similarity characteristics, whether the PDF document has been changed; and

responsive to determining that the PDF document has been changed, generating, using the at least one processor, a graphical user interface including a graphical notification surfacing one or more changes between the original copy and the modified copy of the PDF document to a user, the graphical user interface including an indication of one or more significant changes altering a meaning of the selected segment.

2 . The method of claim 1 , further comprising receiving the original copy of the PDF document from a document originator.

3 . The method of claim 1 , further comprising receiving the modified copy of the PDF document from a signer.

4 . The method of claim 1 , wherein segmenting each of the plurality of images into the plurality of segments comprises:

determining at least a portion of the PDF present in each overlapping segment;

responsive to determining that a threshold amount of the portion of the PDF is in one of the overlapping segments, re-segmenting each of the plurality of images such that the portion of the PDF is present in a single segment.

5 . The method of claim 1 , wherein segmenting each of the plurality of images into the plurality of segments comprises:

determining at least a portion of the PDF presenting in each overlapping segment;

determining a confidence score for each overlapping segment, the confidence score corresponding to an optical character recognition (OCR) operation on the portion of the PDF present in each overlapping segment; and

responsive to confidence score above a threshold for one of the overlapping segments, re-segmenting each of the plurality of images such that the portion of the PDF is present in a single segment.

6 . The method of claim 1 , wherein generating the representation of text in the corresponding selected segments comprises:

performing optical character recognition (OCR) on each corresponding selected segment; and

generating a text vector based on the OCR of each corresponding selected segment.

7 . The method of claim 6 , wherein generating the text vector further comprises determining an average text vector for each sentence in each corresponding selected segment.

8 . The method of claim 7 , wherein determining similarity characteristics between the corresponding selected segments comprises performing a comparison between the average text vector for each sentence in each corresponding selected segment, the comparison based on a comparison metric.

9 . The method of claim 8 , wherein the comparison metric is at least one of: cosine similarity, dot product, Manhattan distance, Euclidean distance, Chebyshev distance, or any combination thereof.

10 . The method of claim 8 , wherein determining, based on the similarity characteristics whether the PDF document has been changed comprises determining an above threshold difference between the average text vector for each sentence in each corresponding selected segment.

11 . The method of claim 1 , further comprising generating a representation of images in the corresponding selected segments by performing pattern matching on each corresponding selected segment.

12 . The method of claim 11 , wherein the images in the corresponding selected segments include one or more signatures in the PDF document.

13 . The method of claim 1 , wherein surfacing one or more changes between the original copy and the modified copy of the PDF document to a user comprises notifying the document originator.

14 . A non-transitory computer-readable storage medium storing executable instructions that, when executed by one or more processors, cause the one or more processors to:

receive an original copy of a PDF document;

receive a modified copy of the PDF document;

rasterize both the original copy and the modified copy of the PDF document into a plurality of images, each image of the plurality of images representing a page, wherein at least one image in the plurality of images includes at least one burned-in object; and

segment each of the plurality of images into a plurality of segments, segmenting of each of the plurality of images includes

determining a height and a buffer of a predetermined number of pixels for an image of each segment in the plurality of segments for expanding each segment;

expanding two or more adjacent segments using the height and the buffer of predetermined number of pixels of each of the two or more adjacent segments to generate at least two overlapping segments that overlap by the predetermined number of pixels;

determining at least one repeated portion in the original and modified copies of the PDF document in the at least two overlapping segments; and

selecting a segment in the at least two overlapping segments in the original and modified copies of the PDF document representative of the at least one repeated portion;

for corresponding selected segments derived from the original copy and the modified copy of the PDF document:

generate a representation of text in the corresponding selected segments;

determine similarity characteristics between the corresponding selected segments;

determine, based on the similarity characteristics, whether the PDF document has been changed; and

responsive to determining that the PDF document has been changed, generate a graphical user interface including a graphical notification surfacing one or more changes between the original copy and the modified copy of the PDF document to a user, the graphical user interface including an indication of one or more significant changes altering a meaning of the selected segment.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more processors are configured to receive the original copy of the PDF document from a document originator.

16 . The non-transitory computer-readable storage medium of claim 14 , wherein the one or more processors are configured to receive the modified copy of the PDF document from a signer.

17 . The non-transitory computer-readable storage medium of claim 14 , wherein segmenting each of the plurality of images into the plurality of segments includes:

determining at least a portion of the PDF present in each overlapping segment;

responsive to determining that a threshold amount of the portion of the PDF is in one of the overlapping segments, re-segmenting each of the plurality of images such that the portion of the PDF is present in a single segment.

18 . The non-transitory computer-readable storage medium of claim 14 , wherein segmenting each of the plurality of images into the plurality of segments includes:

determining at least a portion of the PDF presenting in each overlapping segment;

determining a confidence score for each overlapping segment, the confidence score corresponding to an optical character recognition (OCR) operation on the portion of the PDF present in each overlapping segment; and

responsive to confidence score above a threshold for one of the overlapping segments, re-segmenting each of the plurality of images such that the portion of the PDF is present in a single segment.

19 . The non-transitory computer-readable storage medium of claim 14 , wherein the instructions that cause the one or more processors to generation of the representation of text in the corresponding selected segments includes:

performing optical character recognition (OCR) on each corresponding selected segment; and

generating a text vector based on the OCR of each corresponding selected segment.

20 . A document management system comprising:

a hardware processor; and

a non-transitory computer-readable storage medium storing executable instructions that, when executed, cause the hardware processor to:

receive an original copy of a PDF document;

receive a modified copy of the PDF document;

rasterize both the original copy and the modified copy of the PDF document into a plurality of images, each image of the plurality of images representing a page, wherein at least one image in the plurality of images includes at least one burned-in object; and

segment each of the plurality of images into a plurality of segments, segmenting of each of the plurality of images includes

determining a height and a buffer of a predetermined number of pixels for an image of each segment in the plurality of segments for expanding each segment;

expanding two or more adjacent segments using the height and the buffer of predetermined number of pixels of each of the two or more adjacent segments to generate at least two overlapping segments that overlap by the predetermined number of pixels;

determining at least one repeated portion in the original and modified copies of the PDF document in the at least two overlapping segments; and

selecting a segment in the at least two overlapping segments in the original and modified copies of the PDF document representative of the at least one repeated portion;

for corresponding selected segments derived from the original copy and the modified copy of the PDF document:

generate a representation of text in the corresponding selected segments;

determine similarity characteristics between the corresponding selected segments;

determine, based on the similarity characteristics, whether the PDF document has been changed; and

responsive to determining that the PDF document has been changed, generate a graphical user interface including a graphical notification surfacing one or more changes between the original copy and the modified copy of the PDF document to a user, the graphical user interface including an indication of one or more significant changes altering a meaning of the selected segment.

Assignments (2)
PATENT SECURITY AGREEMENT Recorded May 23, 2025
From: DOCUSIGN, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 071337/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2022
From: ALABI, TAIWO RAPHAEL; MOHAN, ASHWATH SARAN; DUREJA, NIPUN; LEVADOUX, JEROME
To: DOCUSIGN, INC.
Reel/Frame 061895/0911 →
Continuity (1)
Related Publication 20240169755A1 · May 23, 2024
References Cited (20)
US 9634875B2 · Porat · 2017 [cited by applicant]
US 10430570B2 · Gonser · 2019 [cited by applicant]
US 20020029232A1 · Bobrow · 2002 [cited by examiner]
US 20140280254A1 · Feichtner · 2014 [cited by examiner]
US 20150150090A1 · Carroll · 2015 [cited by applicant]
US 20170053164A1 · Eagleton · 2017 [cited by examiner]
US 20170357852A1 · Cai · 2017 [cited by examiner]
US 20190026550A1 · Yang · 2019 [cited by examiner]
US 20200372217A1 · Abuammar · 2020 [cited by examiner]
US 20220035993A1 · Bhandarkar · 2022 [cited by examiner]
US 20220108556A1 · Peng · 2022 [cited by examiner]
US 20230027526A1 · Kim · 2023 [cited by examiner]
US 20230162521A1 · Gao · 2023 [cited by examiner]
US 20240028651A1 · Coquard · 2024 [cited by examiner]
US 20240169755A1 · Alabi · 2024 [cited by examiner]
CN 112632952A · 2021 [cited by examiner]
CN 114896979A · 2022 [cited by examiner]
Anonymous, “Index of /courses/cse573/12sp/lectures,” (Jan. 1, 2012), Retrieved on Feb. 7, 2024; Retrieved from URL: https://courses.cs.washington.edu/courses/cse573/12sp/lectures/. [cited by applicant]
International Search Report and the Written Opinion of the International Searching Authority for International Application No. PCT/US2023/080303, mailed Feb. 23, 2024, 33 pages. [cited by applicant]
Mausam, “Document Similarity in Information Retrieval,” Based on slides of W. Arms, Thomas Hofmann, Ata Kaban, Melanie Martin, (May 2, 2012), Retrieved on Feb. 7, 2024; Retrieved from URL: https://courses.cs.washington.… [cited by applicant]