Information processing method and apparatus
View Patent ↗This application describes examples of an information processing method and apparatus. In one example, the method includes: obtaining an image; inputting the image into a feature extraction model; obtaining, from the feature extraction model, a first feature map; inputting the first feature map into a first semantic recognition model; and obtaining, from the first semantic recognition model, first target semantic information.
1 . An information processing method, comprising:
obtaining an image;
inputting the image into a feature extraction model, wherein the feature extraction model is used to extract a feature map of a target object based on an input image;
obtaining, from the feature extraction model, a first feature map corresponding to the image, wherein the first feature map describes a first form of the target object, wherein obtaining the first feature map comprises performing a three-dimensional model reconstruction process on the target object, and wherein the three-dimensional model reconstruction process comprises generating a primal sketch based on the image, deriving an intrinsic image based on the primal sketch, and forming a three-dimensional model of the target object based on the intrinsic image;
inputting the first feature map into a first semantic recognition model, wherein the first semantic recognition model is used to determine first semantic information based on an input feature map; and
obtaining, from the first semantic recognition model, first target semantic information corresponding to the first feature map, wherein the first target semantic information describes a meaning expressed by the first form.
2 . The information processing method according to claim 1 , wherein a training process of the feature extraction model and a training process of the first semantic recognition model are independent of each other.
3 . The information processing method according to claim 2 , wherein the first feature map comprises information about a three-dimensional model, and the three-dimensional model is determined by the feature extraction model by fitting the target object based on the image by using a parametric model.
4 . The information processing method according to claim 3 , wherein the information about the three-dimensional model comprises at least one of information about a grid vertex in the three-dimensional model or information about a fitting parameter, and the information about the fitting parameter is used to determine the three-dimensional model based on the parametric model.
5 . The information processing method according to claim 1 , wherein after the inputting the image into a feature extraction model, the information processing method further comprises:
obtaining, from the feature extraction model, a second feature map corresponding to the image, wherein the second feature map describes a second form of the target object;
inputting the second feature map into a second semantic recognition model, wherein the second semantic recognition model is used to determine second semantic information based on an input feature map; and
obtaining, from the second semantic recognition model, second target semantic information corresponding to the image, wherein the second target semantic information describes a meaning expressed by the second form.
6 . The information processing method according to claim 5 , wherein a training process of the first semantic recognition model and a training process of the second semantic recognition model are independent of each other.
7 . The information processing method according to claim 6 , wherein the information processing method further comprises:
executing a first visual task based on the first target semantic information; and
executing a second visual task based on the second target semantic information.
8 . The information processing method according to claim 1 , wherein the image is from sensing information of a vehicle-mounted sensor.
9 . The information processing method according to claim 8 , wherein the vehicle-mounted sensor comprises at least one of the following sensors:
a radar, an infrared detector, a depth camera, a full-color camera, or a fisheye camera.
10 . The information processing method according to claim 1 , wherein the target object comprises a person, a vehicle, or a road scenario.
11 . An information processing apparatus, comprising at least one processor and one or more memories coupled to the at least one processor, wherein the one or more memories store instructions for execution by the at least one processor to:
obtain an image;
input the image into a feature extraction model, wherein the feature extraction model is used to extract a feature map of a target object based on an input image;
obtain, from the feature extraction model, a first feature map corresponding to the image, wherein the first feature map describes a first form of the target object, wherein obtaining the first feature map comprises performing a three-dimensional model reconstruction process on the target object, and wherein the three-dimensional model reconstruction process comprises generating a primal sketch based on the image, deriving an intrinsic image based on the primal sketch, and forming a three-dimensional model of the target object based on the intrinsic image;
input the first feature map into a first semantic recognition model, wherein the first semantic recognition model is used to determine first semantic information based on an input feature map; and
obtain, from the first semantic recognition model, first target semantic information corresponding to the first feature map, wherein the first target semantic information describes a meaning expressed by the first form.
12 . The information processing apparatus according to claim 11 , wherein a training process of the feature extraction model and a training process of the first semantic recognition model are independent of each other.
13 . The information processing apparatus according to claim 12 , wherein the first feature map comprises information about a three-dimensional model, and the three-dimensional model is determined by the feature extraction model by fitting the target object based on the image by using a parametric model.
14 . The information processing apparatus according to claim 13 , wherein the information about the three-dimensional model comprises at least one of information about a grid vertex in the three-dimensional model or information about a fitting parameter, and the information about the fitting parameter is used to determine the three-dimensional model based on the parametric model.
15 . The information processing apparatus according to claim 11 , wherein the one or more memories store instructions for execution by the at least one processor to:
obtain, from the feature extraction model, a second feature map corresponding to the image, wherein the second feature map describes a second form of the target object;
input the second feature map into a second semantic recognition model, wherein the second semantic recognition model is used to determine second semantic information based on an input feature map; and
obtain, from the second semantic recognition model, second target semantic information corresponding to the image, wherein the second target semantic information describes a meaning expressed by the second form.
16 . The information processing apparatus according to claim 15 , wherein a training process of the first semantic recognition model and a training process of the second semantic recognition model are independent of each other.
17 . The information processing apparatus according to claim 16 , wherein the one or more memories store instructions for execution by the at least one processor to:
execute a first visual task based on the first target semantic information; and
execute a second visual task based on the second target semantic information.
18 . The information processing apparatus according to claim 17 , wherein the image is from sensing information of a vehicle-mounted sensor.
19 . The information processing apparatus according to claim 18 , wherein the vehicle-mounted sensor comprises at least one of the following sensors: a radar, an infrared detector, a depth camera, a full-color camera, or a fisheye camera.
20 . A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores programming instructions for execution by at least one processor to:
obtain an image;
input the image into a feature extraction model, wherein the feature extraction model is used to extract a feature map of a target object based on an input image;
obtain, from the feature extraction model, a first feature map corresponding to the image, wherein the first feature map describes a first form of the target object, wherein obtaining the first feature map comprises performing a three-dimensional model reconstruction process on the target object, and wherein the three-dimensional model reconstruction process comprises generating a primal sketch based on the image, deriving an intrinsic image based on the primal sketch, and forming a three-dimensional model of the target object based on the intrinsic image;
input the first feature map into a first semantic recognition model, wherein the first semantic recognition model is used to determine first semantic information based on an input feature map; and
obtain, from the first semantic recognition model, first target semantic information corresponding to the first feature map, wherein the first target semantic information describes a meaning expressed by the first form.