IP Library Granted Patent US 12,380,666
Granted Patent B2
US 12,380,666 · App. 17/959,952 · Granted Aug 5, 2025

Compressed fixed-point SIMD macroblock rotation systems and methods

Inventors: Lars Petter Endresen (Nesoddtangen, NO); Øystein Hovind (Hvalstad, NO)
Assignee: FLIR UNMANNED AERIAL SYSTEMS AS
G06V10/242
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,666
App. No.
17/959,952
Granted
Aug 5, 2025
Kind
B2
Abstract

Various techniques are provided for efficient bilinear interpolation of rotated pixels. In one example, a method includes identifying a rotation angle for an image; performing a vector load of pixel positions for the image at the rotation angle; performing a vector load of rows of pixels associated with the pixel positions; performing a vector selection of a subset of pixels from the rows of pixels based on the identified pixel positions; performing a vector load of a set of coefficients at the rotation angle; and applying the set of coefficients to the subset of pixels to determine an updated value for the image. Additional methods and systems are also provided.

Claims (62)

1. A method comprising:

identifying a rotation angle for an input image comprising first pixels; and

determining second pixels of a rotated image obtained by rotating the input image by the rotation angle;

wherein for at least a block of pixels, determining the second pixels in the block comprises a bilinear interpolation of each second pixel in the block from corresponding first pixels, the bilinear interpolation comprising a bilinear interpolation process performed to determine a plurality of the second pixels, the bilinear interpolation process comprising:

performing a vector load of pixel positions from which the plurality of the second pixels will be interpolated, the pixel positions being at least partially defined by the rotation angle;

performing a vector load of rows of first pixels associated with the pixel positions;

performing a vector selection of a subset of pixels from the rows of first pixels based on the pixel positions;

performing a vector load of a set of interpolation coefficients at least partially defined by the rotation angle; and

applying the set of interpolation coefficients to the subset of pixels to determine the plurality of the second pixels.

2. The method of claim 1 , wherein:

the plurality of the second pixels is a first plurality of the second pixels;

the method further comprises determining a second plurality of the second pixels, wherein determining the second plurality comprises a rotation-interpolation process comprising:

performing the bilinear interpolation process on a copy of the block, the copy being rotated by a first angle of 90°, 180°, or 270°, wherein the bilinear interpolation process on the copy reuses the pixel positions and the interpolation coefficients used for the first plurality of the second pixels; and

back-rotating the second plurality of the second pixels by the first angle.

3. The method of claim 2 , wherein:

determining the second plurality is performed for the first angle of 90°; and

the method further comprises:

determining a third plurality of the second pixels by performing the rotation-interpolation process for the first angle of 180°; and

determining a fourth plurality of the second pixels by performing the rotation-interpolation process for the first angle of 270°;

wherein each of the first, second, third, and fourth pluralities of the second pixels is a respective quadrant of the block.

4. The method of claim 1 , wherein the subset of pixels from the rows of first pixels are selected based on adjacency with the pixel positions.

5. The method of claim 1 , wherein the subset of pixels from the rows of first pixels are identified utilizing a first set of instructions that are a set of table (TBL) instructions.

6. The method of claim 1 , wherein the vector selection of the subset of pixels from the rows of first pixels based on the identified pixel positions comprises:

utilizing a set of table (TBL) instructions that each return pixels within the subset of pixels as required by a UDOT instruction that calculates a dot product between two vectors with 4 groups of 4 bytes.

7. The method of claim 6 , wherein the dot product is between two vectors with four groups of four bytes.

8. The method of claim 6 , wherein dot products for the set of TBL instructions are determined simultaneously utilizing an Advanced RISC Machines (Arm) Advanced Single Instruction, Multiple Data (ASIMD) instruction that provides efficient dot product UDOT instructions.

9. The method of claim 1 , wherein the vector load of the pixel positions and the vector load of the set of interpolation coefficients are performed utilizing a zigzag pattern.

10. The method of claim 1 , further comprising:

performing a vector select of a result utilizing a reorganization matrix (RM) that ensures right-shift, reduction, reordering, and rotation thereby forming a transposed result; and

performing a vector store of the transposed result to memory.

11. A system comprising:

a memory component storing machine-executable instructions; and

a logic device configured to execute the instructions to cause the system to:

identify a rotation angle for an input image comprising first pixels; and

determine second pixels of a rotated image obtained by rotating the input image by the rotation angle;

wherein for at least a block of pixels, determining the second pixels in the block comprises a bilinear interpolation of each second pixel in the block from corresponding first pixels, the bilinear interpolation comprising a bilinear interpolation process performed to determine a plurality of the second pixels, the bilinear interpolation process comprising:

performing a vector load of pixel positions from which the plurality of the second pixels will be interpolated, the pixel positions being at least partially defined by the rotation angle;

performing a vector load of rows of first pixels associated with the pixel positions;

performing a vector selection of a subset of pixels from the rows of first pixels based on the pixel positions;

performing a vector load of a set of interpolation coefficients at least partially defined by the rotation angle; and

applying the set of interpolation coefficients to the subset of pixels to determine the plurality of the second pixels.

12. The system of claim 11 , wherein:

the plurality of the second pixels is a first plurality of the second pixels;

the logic device is further configured to determine a second plurality of the second pixels, wherein determining the second plurality comprises a rotation-interpolation process comprising:

performing the bilinear interpolation process on a copy of the block, the copy being rotated by a first angle of 90°, 180°, or 270°, wherein the bilinear interpolation process on the copy reuses the pixel positions and the interpolation coefficients used for the first plurality of the second pixels; and

back-rotating the second plurality of the second pixels by the first angle.

13. The system of claim 12 , wherein:

determining the second plurality is performed for the first angle of 90°; and

the logic device is further configured to:

determine a third plurality of the second pixels by performing the rotation-interpolation process for the first angle of 180°; and

determine a fourth plurality of the second pixels by performing the rotation-interpolation process for the first angle of 270°;

wherein each of the first, second, third, and fourth pluralities of the second pixels is a respective quadrant of the block.

14. The system of claim 11 , wherein the subset of pixels from the rows of first pixels are selected based on adjacency with the pixel positions.

15. The method of claim 11 , wherein the subset of pixels from the rows of first pixels are identified utilizing a first set of instructions that are a set of table (TBL) instructions.

16. The method of claim 11 , wherein the vector selection of the subset of pixels from the rows of first pixels based on the identified pixel positions comprises:

utilizing a set of table (TBL) instructions that each return pixels within the subset of pixels as required by a UDOT instruction that calculates a dot product between two vectors with 4 groups of 4 bytes.

17. The system of claim 16 , wherein the dot product is between two vectors with four groups of four bytes.

18. The system of claim 16 , wherein dot products for the set of TBL instructions are determined simultaneously utilizing an Advanced RISC Machines (Arm) Advanced Single Instruction, Multiple Data (ASIMD) instruction that provides efficient dot product UDOT instructions.

19. The system of claim 11 , wherein the vector load of the pixel positions and the vector load of the set of interpolation coefficients are performed utilizing a zigzag pattern.

20. The method of claim 11 , wherein the logic device is further configured to execute the instructions to cause the system to:

perform a vector select of a result utilizing a reorganization matrix (RM) that ensures right-shift, reduction, reordering, and rotation thereby forming a transposed result; and

perform a vector store of the transposed result to memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: ENDRESEN, LARS PETTER; HOVIND, OYSTEIN
To: FLIR UNMANNED AERIAL SYSTEMS AS
Reel/Frame 061455/0711 →
Continuity (2)
Provisional Application 63252568 · Oct 5, 2021
Related Publication 20230105192A1 · Apr 6, 2023
References Cited (19)
US 5111192A · Kadakia · 1992 [cited by applicant]
US 7376286B2 · Locker et al. · 2008 [cited by applicant]
US 8463837B2 · Chen · 2013 [cited by examiner]
US 8797359B2 · Tripathi et al. · 2014 [cited by applicant]
US 9143799B2 · Endresen · 2015 [cited by examiner]
US 9544451B2 · Silverbrook · 2017 [cited by examiner]
US 9756360B2 · Joshi · 2017 [cited by examiner]
US 10339664B2 · Rhoads · 2019 [cited by examiner]
US 10735659B2 · Messely et al. · 2020 [cited by applicant]
US 11184617B2 · Egilmez · 2021 [cited by examiner]
US 20050094899A1 · Kim · 2005 [cited by examiner]
US 20130300769A1 · Huang et al. · 2013 [cited by applicant]
WO WO2020247212A1 · 2020 [cited by applicant]
S. A. Fahmy, C.-s. Bouganis, P. Y. k. Cheung and W. Luk, “Efficient Realtime FPGA Implementation of the Trace Transform,” 2006 International Conference on Field Programmable Logic and Applications, Madrid, Spain, 2006, … [cited by examiner]
M. K. Rahman, M. H. Sujon and A. Azad, “FusedMM: A Unified SDDMM-SpMM Kernel for Graph Embedding and Graph Neural Networks,” 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS), Portland, OR, U… [cited by examiner]
Arabnia et al., “Arbitrary rotation of raster images with SIMD machine architectures”, Computer Graphics Forum, Mar. 1, 1987, p. 3-11, vol. 6, Blackwell Publishing Ltd, Oxford, United Kingdom. [cited by applicant]
Danielsson et al., “High-Accuracy Rotation of Images”, CVGIP: Graphical Models and Image Processing, Jul. 1992, p. 340-344, vol. 54, No. 4, Elsevier, Amsterdam, Netherlands. [cited by applicant]
Ramesh et al., “Optimization and Evaluation of Image-and Signal-Processing Kernels on the TI C6678 Multi-Core DSP” 2014 IEEE High Performance Extreme Computing Conference (HPEC), Sep. 2014, p. 1-6, IEEE, Waltham, MA, Un… [cited by applicant]
Zuo, Song, “Fast Data-Parallel Rendering of Digital Volume Images”, Thesis, Jun. 1995, 81 pages, The Chinese University of Hong Kong, Hong Kong, China. [cited by applicant]