IP Library › Granted Patent US 11,928,800
Granted Patent B2
US 11,928,800 · App. 17/373,768 · Granted Mar 12, 2024

Image coordinate system transformation method and apparatus, device, and storage medium

Inventor: Xiangqi Huang (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06T5/50G06T7/11G06T2207/10016G06T2207/20221G06T2207/30204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,928,800
App. No.
17/373,768
Granted
Mar 12, 2024
Kind
B2
Abstract

An image coordinate system transformation method includes obtaining video images acquired by adjacent cameras, the adjacent cameras including a first camera and a second camera that have an overlapping photography region on a ground plane, recognizing N groups of key points of a target object on the ground plane from the video images acquired by the adjacent cameras, each group of key points including first key point extracted from a video image of the first camera and a second key point extracted from a video image of the second camera, the first key point and the second key point being the same feature point of the same target object appearing in the adjacent cameras at the same moment, and N being an integer greater than or equal to 3, and calculating a transformation relationship between image coordinate systems of the adjacent cameras according to the N groups of key points.

Claims (59)

1. An image coordinate system transformation method, applied to a computing device, the method comprising:

obtaining video images acquired by adjacent cameras, the adjacent cameras including a first camera and a second camera that have an overlapping photography region on a ground plane;

recognizing N groups of key points of a target object on the ground plane from the video images acquired by the adjacent cameras, each group of key points including first key point extracted from a video image of the first camera and a second key point extracted from a video image of the second camera, the first key point and the second key point being the same feature point of the same target object appearing in the adjacent cameras at the same moment, and N being an integer greater than or equal to 3; and

calculating a transformation relationship between image coordinate systems of the adjacent cameras according to the N groups of key points.

2. The method according to claim 1 , wherein recognizing N groups of key points of the target object comprises:

performing target detection and tracking on the video images acquired by the adjacent cameras respectively, to obtain a detection and tracking result corresponding to the first camera and a detection and tracking result corresponding to the second camera;

sifting out a standard target object according to the detection and tracking result corresponding to the first camera and the detection and tracking result corresponding to the second camera, the standard target object being the same target object appearing in the adjacent cameras at the same moment; and

performing key point detection on the standard target object to obtain the N groups of key points.

3. The method according to claim 2 , wherein sifting out the standard target object according to the detection and tracking result comprises:

obtaining, according to the detection and tracking result corresponding to the first camera, an appearance feature of a first target object obtained through detection and tracking from a first video image acquired by the first camera;

obtaining, according to the detection and tracking result corresponding to the second camera, an appearance feature of a second target object obtained through detection and tracking from a second video image acquired by the second camera, the first video image and the second video image being video images acquired by the adjacent cameras at the same moment;

calculating a similarity between the appearance feature of the first target object and the appearance feature of the second target object; and

determining that the first target object and the second target object are the standard target object in response to determining the similarity is greater than a similarity threshold.

4. The method according to claim 3 , further comprising:

sifting out the first video image and the second video image that meet a condition according to the detection and tracking result corresponding to the first camera and the detection and tracking result corresponding to the second camera,

the condition including that a quantity of target objects obtained through detection and tracking from the first video image is 1 and a quantity of target objects obtained through detection and tracking from the second video image is also 1.

5. The method according to claim 2 , wherein performing key point detection on the standard target object comprises:

extracting, in response to determining the standard target object is a pedestrian, midpoints of connecting lines between two feet of the standard target object, to obtain the N groups of key points.

6. The method according to claim 2 , further comprising:

obtaining, for each group of key points, a confidence level corresponding to the key point; and

removing the key point in response to determining the confidence level corresponding to the key point is less than a confidence level threshold.

7. The method according to claim 1 , wherein the N groups of key points come from video images of the same target object at N different moments.

8. The method according to claim 1 , wherein calculating the transformation relationship between image coordinate systems of the adjacent cameras comprises:

calculating an affine transformation matrix between the image coordinate systems of the adjacent cameras according to the N groups of key points.

9. The method according to claim 1 , further comprising:

calculating, for any object obtained through detection and tracking from the video image of the first camera, position coordinates of the object in an image coordinate system corresponding to the second camera according to position coordinates of the object in an image coordinate system corresponding to the first camera and the transformation relationship.

10. An image coordinate system transformation apparatus, the apparatus comprising: a memory storing computer program instructions; and a processor coupled to the memory and configured to execute the computer program instructions and perform:

obtaining video images acquired by adjacent cameras, the adjacent cameras including a first camera and a second camera that have an overlapping photography region on a ground plane;

recognizing N groups of key points of a target object on the ground plane from the video images acquired by the adjacent cameras, each group of key points including a first key point extracted from a video image of the first camera and a second key point extracted from a video image of the second camera, the first key point and the second key point being the same feature point of the same target object appearing in the adjacent cameras at the same moment, and N being an integer greater than or equal to 3; and

calculating a transformation relationship between image coordinate systems of the adjacent cameras according to the N groups of key points.

11. The apparatus according to claim 10 , wherein the processor is further configured to execute the computer program instructions and perform:

performing target detection and tracking on the video images acquired by the adjacent cameras respectively, to obtain a detection and tracking result corresponding to the first camera and a detection and tracking result corresponding to the second camera;

sifting out a standard target object according to the detection and tracking result corresponding to the first camera and the detection and tracking result corresponding to the second camera, the standard target object being the same target object appearing in the adjacent cameras at the same moment; and

performing key point detection on the standard target object to obtain the N groups of key points.

12. The apparatus according to claim 11 , wherein the processor is further configured to execute the computer program instructions and perform:

obtaining, according to the detection and tracking result corresponding to the first camera, an appearance feature of a first target object obtained through detection and tracking from a first video image acquired by the first camera; obtain, according to the detection and tracking result corresponding to the second camera, an appearance feature of a second target object obtained through detection and tracking from a second video image acquired by the second camera, the first video image and the second video image being video images acquired by the adjacent cameras at the same moment;

calculating a similarity between the appearance feature of the first target object and the appearance feature of the second target object; and

determining that the first target object and the second target object are the standard target object in response to determining the similarity is greater than a similarity threshold.

13. The apparatus according to claim 12 , wherein the processor is further configured to execute the computer program instructions and perform:

sifting out the first video image and the second video image that meet a condition according to the detection and tracking result corresponding to the first camera and the detection and tracking result corresponding to the second camera,

the condition including that a quantity of target objects obtained through detection and tracking from the first video image is 1 and a quantity of target objects obtained through detection and tracking from the second video image is also 1.

14. The apparatus according to claim 11 , wherein the processor is further configured to execute the computer program instructions and perform:

in response to determining the standard target object is a pedestrian, extracting midpoints of connecting lines between two feet of the standard target object, to obtain the N groups of key points.

15. The apparatus according to claim 11 , wherein the processor is further configured to execute the computer program instructions and perform:

obtaining, for each group of key points, a confidence level corresponding to the key point; and

removing the key point in response to determining the confidence level corresponding to the key point is less than a confidence level threshold.

16. The apparatus according to claim 10 , wherein the N groups of key points come from video images of the same target object at N different moments.

17. The apparatus according to claim 10 , wherein the processor is further configured to execute the computer program instructions and perform:

calculating an affine transformation matrix between the image coordinate systems of the adjacent cameras according to the N groups of key points.

18. The apparatus according to claim 10 , wherein the processor is further configured to execute the computer program instructions and perform:

calculating, for any object obtained through detection and tracking from the video image of the first camera, position coordinates of the object in an image coordinate system corresponding to the second camera according to position coordinates of the object in an image coordinate system corresponding to the first camera and the transformation relationship.

19. A non-transitory computer-readable storage medium storing computer program instructions executable by at least one processor to perform:

obtaining video images acquired by adjacent cameras, the adjacent cameras including a first camera and a second camera that have an overlapping photography region on a ground plane;

recognizing N groups of key points of a target object on the ground plane from the video images acquired by the adjacent cameras, each group of key points including first key point extracted from a video image of the first camera and a second key point extracted from a video image of the second camera, the first key point and the second key point being the same feature point of the same target object appearing in the adjacent cameras at the same moment, and N being an integer greater than or equal to 3; and

calculating a transformation relationship between image coordinate systems of the adjacent cameras according to the N groups of key points.

20. The non-transitory computer-readable storage medium according to claim 19 , wherein the computer program instructions are executable by the at least one processor to further perform:

performing target detection and tracking on the video images acquired by the adjacent cameras respectively, to obtain a detection and tracking result corresponding to the first camera and a detection and tracking result corresponding to the second camera;

sifting out a standard target object according to the detection and tracking result corresponding to the first camera and the detection and tracking result corresponding to the second camera, the standard target object being the same target object appearing in the adjacent cameras at the same moment; and

performing key point detection on the standard target object to obtain the N groups of key points.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2021
From: HUANG, XIANGQI
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 056830/0725 →
Priority Claims (1)
CN 201910704514.8 · Jul 31, 2019 · national
Continuity (2)
Continuation PCTCN2020102493 · Jul 16, 2020
Related Publication 20210342990A1 · Nov 4, 2021