Method and apparatus for processing image, device and storage medium
A method and apparatus for processing an image, a device and a storage medium are provided. An implementation of the method includes: acquiring a template image, the template image including at least one region of interest; determining a first feature map corresponding to each region of interest in the template image; acquiring a target image; determining a second feature map of the target image; and determining at least one region of interest in the target image according to the first feature map and the second feature map.
1. A method for processing an image, comprising:
acquiring a template image, the template image including at least one region of interest;
determining a first feature map corresponding to each region of interest in the template image, wherein determining the first feature map corresponding to each region of interest comprises: extracting a feature of the template image to obtain a third feature map, and determining the first feature map corresponding to the to the each region of interest based on the third feature map and the at feast one region of interest;
acquiring a target image;
determining a second feature map of the target image; and
determining at least one region of interest in the target image according to the first feature map and the second feature map.
2. The method according to claim 1 , wherein the determining the first feature map corresponding to the each region of interest based on the third feature map and the at least one region of interest comprises:
performing an interception on the third feature map by using the each region of interest, to obtain an interception feature map corresponding to the each region of interest; and
performing at least one of a convolution operation or a scaling operation on the interception feature map, to obtain the first feature map corresponding to the each region of interest.
3. The method according to claim 1 , wherein the determining the at least one region of interest in the target image based on the first feature map and the second feature map comprises:
for each first feature map, performing a convolution calculation on the second feature map with the each first feature map as a convolution kernel, to obtain a response map corresponding to the each first feature map; and
determining the at least one region of interest in the target image according to the response map corresponding to the each first feature map.
4. The method according to claim 3 , wherein the determining the at least one region of interest in the target image according to the response map corresponding to the each first feature map comprises:
performing a threshold-based segmentation and a connected component analysis on the response map corresponding to the each first feature map, to determine the at least one region of interest in the target image.
5. The method according to claim 1 , further comprising:
determining an envelope box for the each region of interest; and
highlighting the at least one region of interest in the target image according to the envelope box for the each region of interest.
6. The method according to claim 1 , further comprising:
performing a text recognition on each of the determined at least one region of interest in the target image.
7. An electronic device for processing an image, comprising:
at least one computation unit; and
a storage unit, communicated with the at least one computation unit,
wherein the storage unit stores an instruction executable by the at least one computation unit, and the instruction is executed by the at least one computation unit, to enable the at least one computation unit to perform operations, the operations comprising:
acquiring a template image, the template image including at least one region of interest;
determining a first feature map corresponding to each region of interest in the template image, determining the first feature map corresponding to each region of interest comprises: extracting a feature of the template image to obtain a third feature map, and determining the first feature map corresponding to the each region of interest based on the third feature map and the at least one region of interest;
acquiring a target image;
determining a second feature map of the target image; and
determining at least one region of interest in the target image according to the first feature map and the second feature map.
8. The electronic device according to claim 7 , wherein the determining the first feature map corresponding to the each region of interest based on the third feature map and the at least one region of interest comprises:
performing an interception on the third feature map by using the each region of interest, to obtain an interception feature map corresponding to the each region of interest; and
performing at least one of a convolution operation or a scaling operation on the interception feature map, to obtain the first feature map corresponding to the each region of interest.
9. The electronic device according to claim 7 , wherein the determining the at least one region of interest in the target image based on the first feature map and the second feature map comprises:
for each first feature map, performing a convolution calculation on the second feature map with the each first feature map as a convolution kernel, to obtain a response map corresponding to the each first feature map; and
determining the at least one region of interest in the target image according to the response map corresponding to the each first feature map.
10. The electronic device according to claim 9 , wherein the determining the at least one region of interest in the target image according to the response map corresponding to the each first feature map comprises:
performing a threshold-based segmentation and a connected component analysis on the response map corresponding to the each first feature map, to determine the at least one region of interest in the target image.
11. The electronic device according to claim 7 , wherein the operations further comprise:
determining an envelope box for the each region of interest; and
highlighting the at least one region of interest in the target image according to the envelope box for the each region of interest.
12. The electronic device according to claim 7 , wherein the operations further comprise:
performing a text recognition on each of the determined at least one region of interest in the target image.
13. A non-transitory computer readable storage medium, storing a computer instruction, wherein the computer instruction, when executed by a processor, cause the processor to perform operations, the operations comprising:
acquiring a template image, the template image including at least one region of interest;
determining a first feature map corresponding to each region of interest in the template image, wherein determining the first feature map comes corresponding to each region of interest comprises: extracting a feature of the template image to obtain a third feature map, and determining the first feature map corresponding to the each region of interest based on the third feature map and the at least one region of interest;
acquiring a target image;
determining a second feature map of the target image; and
determining at least one region of interest in the target image according to the first feature map and the second feature map.
14. The medium according to claim 13 , wherein the determining the first feature map corresponding to the each region of interest based on the third feature map and the at least one region of interest comprises:
performing an interception on the third feature map by using the each region of interest, to obtain an interception feature map corresponding to the each region of interest; and
performing at least one of a convolution operation or a scaling operation on the interception feature map, to obtain the first feature map corresponding to the each region of interest.
15. The medium according to claim 13 , wherein the determining the at least one region of interest in the target image based on the first feature map and the second feature map comprises:
for each first feature map, performing a convolution calculation on the second feature map with the each first feature map as a convolution kernel, to obtain a response map corresponding to the each first feature map; and
determining the at least one region of interest in the target image according to the response map corresponding to the each first feature map.
16. The medium according to claim 15 , wherein the determining the at least one region of interest in the target image according to the response map corresponding to the each first feature map comprises:
performing a threshold-based segmentation and a connected component analysis on the response map corresponding to the each first feature map, to determine the at least one region of interest in the target image.
17. The medium according to claim 13 , wherein the operations further comprise:
determining an envelope box for the each region of interest; and
highlighting the at least one region of interest in the target image according to the envelope box for the each region of interest.