IP Library Granted Patent US 12699839
Granted Patent B2
US 12699839 · App. 18/427,419 · Granted Aug 4, 2026

System and method for extracting information from partial images based on text stitching

Inventors: Rahul Kumar Gupta (Ballia, IN); Shilka Roy (Noida, IN); Sujit Jos (Ernakulam, IN)
Assignee: Walmart Apollo, LLC
G06F40/279G06F40/232G06V30/141G06V30/1801G06V30/245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12699839
App. No.
18/427,419
Filed
Jan 30, 2024
Granted
Aug 4, 2026
Kind
B2
Art Unit
2663
USPC
382/181
Abstract

A computer-implemented method including detecting respective one or more text boxes in each of multiple partial images of a text-bearing area. The method also can include determining respective one or more edge text boxes of the respective one or more text boxes in each of overlapping partial images of the multiple partial images, wherein each of the respective one or more edge text boxes comprise a respective incomplete text. The method additionally can include matching one or more pairs of corresponding edge text boxes from the respective one or more edge text boxes of two adjacent images of the overlapping partial images of the multiple partial images. The method also can include determining cross-image texts in the one or more pairs of the corresponding edge text boxes. The method further can include determining one or more entities in the text-bearing area based on entity texts of the cross-image texts and non-edge texts in respective one or more non-edge text boxes of the respective one or more text boxes in the multiple partial images. Other embodiments are described.

Claims (72)

1 . A system comprising one or more processors and one or more non-transitory computer-readable media storing computing instructions that, when run on the one or more processors, cause the one or more processors to perform operations comprising:

detecting respective one or more text boxes in each of multiple partial images of a text-bearing area;

determining respective one or more edge text boxes of the respective one or more text boxes in each of overlapping partial images of the multiple partial images, wherein each of the respective one or more edge text boxes comprise a respective incomplete text;

matching one or more pairs of corresponding edge text boxes from the respective one or more edge text boxes of two adjacent images of the overlapping partial images of the multiple partial images;

determining cross-image texts in the one or more pairs of the corresponding edge text boxes; and

determining one or more entities in the text-bearing area based on entity texts of the cross-image texts and non-edge texts in respective one or more non-edge text boxes of the respective one or more text boxes in the multiple partial images.

2 . The system in claim 1 , wherein detecting the respective one or more text boxes in each of the multiple partial images further comprises:

detecting respective characters in each of the multiple partial images;

determining small characters of the respective characters in each of the multiple partial images, wherein each of the small characters comprises a respective font size less than a respective font-size threshold for each of the multiple partial images;

masking the small characters from each of the multiple partial images; and

after masking, determining the respective one or more text boxes for remaining characters of the respective characters that are not masked in each of the multiple partial images.

3 . The system in claim 2 , wherein detecting the respective characters in each of the multiple partial images further comprises:

converting each of the multiple partial images into a respective binary image; and

detecting a respective character contour in the respective binary image for each of the respective characters.

4 . The system in claim 3 , wherein determining the small characters of the respective characters in each of the multiple partial images further comprises:

determining a respective font size of each of the respective characters based on the respective character contour for each of the respective characters; and

determining the respective font-size threshold for each of the multiple partial images based on the respective font size of each of the respective characters in each of the multiple partial images.

5 . The system in claim 1 , wherein determining the respective one or more edge text boxes in each of the overlapping partial images of the multiple partial images further comprises:

detecting a respective incomplete text for each of the respective one or more edge text boxes based on a respective distance between the respective incomplete text and a side edge of each of the overlapping partial images of the multiple partial images.

6 . The system in claim 1 , wherein matching the one or more pairs of the corresponding edge text boxes from the respective one or more edge text boxes of the two adjacent images of the overlapping partial images of the multiple partial images further comprises:

extracting a respective key descriptor for each of the respective one or more edge text boxes of the two adjacent images; and

determining whether a first edge text box of the respective one or more edge text boxes of a first partial image of the two adjacent images matches a second edge text box of the respective one or more edge text boxes of a second partial image of the two adjacent images based on a feature distance between the respective key descriptor for the first edge text box and the respective key descriptor for the second edge text box, wherein:

when the first edge text box, as determined, matches the second edge text box, one of the one or more pairs of the corresponding edge text boxes comprises the first edge text box and the second edge text box.

7 . The system in claim 1 , wherein determining the cross-image texts in the one or more pairs of the corresponding edge text boxes further comprises:

removing overlapping characters from the one or more pairs of the corresponding edge text boxes based on a respective matching prefix-suffix pair of each of the one or more pairs of the corresponding edge text boxes.

8 . The system in claim 1 , wherein:

determining the one or more entities further comprises determining one or more word groups based on the cross-image texts and the non-edge texts; and

each of the entity texts comprises a corresponding word or one of the one or more word groups, as determined.

9 . The system in claim 1 , wherein determining the one or more entities further comprises one or more of:

(a) correcting one or more spelling errors in the entity texts; or

(b) detecting one or more entity signals for the one or more entities based on one or more of:

a keyword or a pattern in each of the entity texts; or

a machine learning module pre-trained to determine whether a text of the entity texts comprises an entity signal of the one or more entity signals; and

determining one or more entity values for the one or more entities based on the one or more entity signals.

10 . The system in claim 9 , wherein determining the one or more entity values for the one or more entities further comprises:

performing a Regex-based extraction on the entity texts based on the one or more entity signals.

11 . A computer-implemented method comprising:

detecting respective one or more text boxes in each of multiple partial images of a text-bearing area;

determining respective one or more edge text boxes of the respective one or more text boxes in each of overlapping partial images of the multiple partial images, wherein each of the respective one or more edge text boxes comprise a respective incomplete text;

matching one or more pairs of corresponding edge text boxes from the respective one or more edge text boxes of two adjacent images of the overlapping partial images of the multiple partial images;

determining cross-image texts in the one or more pairs of the corresponding edge text boxes; and

determining one or more entities in the text-bearing area based on entity texts of the cross-image texts and non-edge texts in respective one or more non-edge text boxes of the respective one or more text boxes in the multiple partial images.

12 . The computer-implemented method in claim 11 , wherein detecting the respective one or more text boxes in each of the multiple partial images further comprises:

detecting respective characters in each of the multiple partial images;

determining small characters of the respective characters in each of the multiple partial images, wherein each of the small characters comprises a respective font size less than a respective font-size threshold for each of the multiple partial images;

masking the small characters from each of the multiple partial images; and

after masking, determining the respective one or more text boxes for remaining characters of the respective characters that are not masked in each of the multiple partial images.

13 . The computer-implemented method in claim 12 , wherein detecting the respective characters in each of the multiple partial images further comprises:

converting each of the multiple partial images into a respective binary image; and

detecting a respective character contour in the respective binary image for each of the respective characters.

14 . The computer-implemented method in claim 13 , wherein determining the small characters of the respective characters in each of the multiple partial images further comprises:

determining a respective font size of each of the respective characters based on the respective character contour for each of the respective characters; and

determining the respective font-size threshold for each of the multiple partial images based on the respective font size of each of the respective characters in each of the multiple partial images.

15 . The computer-implemented method in claim 11 , wherein determining the respective one or more edge text boxes in each of the overlapping partial images of the multiple partial images further comprises:

detecting a respective incomplete text for each of the respective one or more edge text boxes based on a respective distance between the respective incomplete text and a side edge of each of the overlapping partial images of the multiple partial images.

16 . The computer-implemented method in claim 11 , wherein matching the one or more pairs of the corresponding edge text boxes from the respective one or more edge text boxes of the two adjacent images of the overlapping partial images of the multiple partial images further comprises:

extracting a respective key descriptor for each of the respective one or more edge text boxes of the two adjacent images; and

determining whether a first edge text box of the respective one or more edge text boxes of a first partial image of the two adjacent images matches a second edge text box of the respective one or more edge text boxes of a second partial image of the two adjacent images based on a feature distance between the respective key descriptor for the first edge text box and the respective key descriptor for the second edge text box, wherein:

when the first edge text box, as determined, matches the second edge text box, one of the one or more pairs of the corresponding edge text boxes comprises the first edge text box and the second edge text box.

17 . The computer-implemented method in claim 11 , wherein determining the cross-image texts in the one or more pairs of the corresponding edge text boxes further comprises:

removing overlapping characters from the one or more pairs of the corresponding edge text boxes based on a respective matching prefix-suffix pair of each of the one or more pairs of the corresponding edge text boxes.

18 . The computer-implemented method in claim 11 , wherein:

determining the one or more entities further comprises determining one or more word groups based on the cross-image texts and the non-edge texts; and

each of the entity texts comprises a corresponding word or one of the one or more word groups, as determined.

19 . The computer-implemented method in claim 11 , wherein determining the one or more entities further comprises one or more of:

(a) correcting one or more spelling errors in the entity texts; or

(b) detecting one or more entity signals for the one or more entities based on one or more of:

a keyword or a pattern in each of the entity texts; or

a machine learning module pre-trained to determine whether a text of the entity texts comprises an entity signal of the one or more entity signals; and

determining one or more entity values for the one or more entities based on the one or more entity signals.

20 . The computer-implemented method in claim 19 , wherein determining the one or more entity values for the one or more entities further comprises:

performing a Regex-based extraction on the entity texts based on the one or more entity signals.