Text watermarking using bitstream encoding
Described are methods and systems that watermark text files in document or image formats using efficient encoding schemes. A unique identifier is encoded into a document by perturbing typographical properties of document elements, such as the lengths and widths of words, lines, or spaces, to encode multiple bits per element. Perturbations to the rendered dimensions of elements create patterns, digital watermarks, that can be decoded to recover the unique identifier, which can in turn be used to identify a user who disclosed or was otherwise responsible for a leaked document.
1 . A method for encoding an identifier in a digital document, the method comprising:
formatting a first text element of the digital document responsive to a first multi-bit pattern from the identifier to encode the first multi-bit pattern as a first areal perturbation of the first text element by adjusting at least one of a height and a width of the first text element in accordance with the first multi-bit pattern;
formatting a second text element of the digital document responsive to a second multi-bit pattern from the identifier to encode the second multi-bit pattern as a second areal perturbation of the second text element by adjusting at least one of a height and a width of the second text element in accordance with the second multi-bit pattern; and
formatting a third text element of the digital document responsive to a third multi-bit pattern from the identifier to encode the third multi-bit pattern as a third areal perturbation of the third text element by adjusting at least one of a height and a width of the third text element in accordance with the third multi-bit pattern.
2 . The method of claim 1 , further comprising correlating a fourth text element with a fourth multi-bit pattern from the identifier and foregoing formatting of the fourth text element responsive to the fourth multi-bit pattern.
3 . The method of claim 2 , wherein each of the first, second, third, and fourth multi-bit patterns is two bits.
4 . The method of claim 1 , wherein the first areal perturbation is a change in width and the second areal perturbation is a change in height.
5 . The method of claim 1 , wherein the text elements comprise words, the method further comprising, for each of the words:
delineating a bounding box around the word, the bounding box encompassing an area that includes the word and background;
scaling the area; and
replacing or covering the word with the scaled area.
6 . The method of claim 5 , wherein the background includes a margin around the text elements, the margin sufficient to cover the word when the area is scaled.
7 . The method of claim 5 , further comprising recognizing the text elements.
8 . The method of claim 1 , wherein the identifier has a number of bits, the method further comprising deriving a code delimiter from the number of bits and formatting at least one fourth text element to encode the code delimiter.
9 . The method of claim 1 , wherein the identifier combines a user identifier and a document identifier.
10 . The method of claim 1 , wherein at least two bits of the identifier are encoded in a single change to a renderable dimension of a corresponding text element.
11 . The method of claim 10 , wherein the single change comprises a dimensional variation of the corresponding text element selected from an increase in height, a decrease in height, an increase in width, and a decrease in width.
12 . The method of claim 1 , wherein more than two bits of the identifier are encoded in a fewer number of dimensional changes than the number of bits being encoded, such that the number of dimensional changes to the text elements is fewer than the number of bits encoded.
13 . A computer-implemented method of extracting an identifier encoded into a scanned document using an encoding scheme that encodes patterns in areal properties of text elements, the areal properties including at least one of width and height, the method comprising:
recognizing first text elements in the scanned document using at least one processor, each of the first text elements having respective first areal properties, wherein the areal properties of the first text elements differ among the first text elements to form a first encoded pattern;
detecting, with the at least one processor, at least one second text element in the scanned document, the second text element having second areal properties that differ from the first areal properties of the first text elements and form a second encoded pattern;
decoding, with the at least one processor, the second encoded pattern to select an inverse of the encoding scheme; and
decoding, with the at least one processor, the first encoded pattern using the inverse of the encoding scheme to extract the identifier.
14 . The method of claim 13 , wherein the first areal properties of the first text elements are encoded with variations in the height among the first text elements.
15 . The method of claim 14 , wherein the second areal properties are encoded with variations in the height of the at least one second text element relative to the heights of the first text elements.
16 . The method of claim 15 , the first and second text elements having respective first and second average heights, and wherein the first average height differs from the second average height.
17 . A method for encoding an identifier in a digital document, the method comprising:
formatting a first text element of the digital document responsive to a first multi-bit pattern from the identifier to encode the first multi-bit pattern by perturbing a first typographical property of the first text element;
formatting a second text element of the digital document responsive to a second multi-bit pattern from the identifier to encode the second multi-bit pattern by perturbing a second typographical property of the second text element; and
formatting a third text element of the digital document responsive to a third multi-bit pattern from the identifier to encode the third multi-bit pattern by perturbing a third typographical property of the third text element.
18 . The method of claim 17 , further comprising correlating a fourth text element with a fourth multi-bit pattern from the identifier and foregoing formatting of the fourth text element responsive to the fourth multi-bit pattern.
19 . The method of claim 18 , wherein each of the first, second, third, and fourth multi-bit patterns is two bits.
20 . The method of claim 17 , wherein the first typographical property is width and the second typographical property is height.
21 . The method of claim 17 , wherein the text elements comprise words, the method further comprising, for each of the words:
delineating a bounding box around the word, the bounding box encompassing an area that includes the word and background;
scaling the area; and
replacing or covering the word with the scaled area.
22 . The method of claim 21 , wherein the background includes a margin around the text elements, the margin sufficient to cover the word when the area is scaled.
23 . The method of claim 17 , wherein the identifier has a number of bits, the method further comprising deriving a code delimiter from the number of bits and formatting at least one fourth text element to encode the code delimiter.
24 . The method of claim 17 , wherein at least two bits of the identifier are encoded in a single change to a typographical property of a corresponding text element, thereby reducing a total number of property changes relative to the number of bits encoded.