IP Library Granted Patent US 12,425,636
Granted Patent B2
US 12,425,636 · App. 18/521,182 · Granted Sep 23, 2025

Segmentation-based parameterized motion models

Inventors: Debargha Mukherjee (Cupertino, CA); Yuxin Liu (Palo Alto, CA); Sarah Parker (San Francisco, CA)
Assignee: GOOGLE LLC
H04N19/517H04N19/17H04N19/20H04N19/521H04N19/54H04N19/543H04N19/547H04N19/557H04N19/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,425,636
App. No.
18/521,182
Granted
Sep 23, 2025
Kind
B2
Abstract

Multiple global motion models associated with respective segments of a current frame are decoded from a compressed bitstream. Each global motion model is based on a segmentation of the current frame and represents a respective underlying motion of blocks within a respective segment. Blocks of the current frame are decoded by: for each inter-predicted block of a segment, decoding, form the compressed bitstream, an indication of whether to decode the each inter-predicted block based on a global motion model of the multiple global motion models and associated with the segment, or whether to decode the each inter-predicted block based on a motion vector that is different from the global motion model; and decoding the each inter-predicted block based on the indication.

Claims (50)

1. A method, comprising:

decoding, from a compressed bitstream and with respect to each reference frame of two or more reference frames available for decoding a current frame, at least two global motion models, wherein decoding a global motion model comprises decoding parameters for the global motion model, wherein the global motion models constitute multiple global motion models and wherein each global motion model is based on a segmentation of the current frame and represents a respective underlying motion of blocks within a respective segment; and

decoding blocks of the current frame by:

for each inter-predicted block of a segment, decoding, from the compressed bitstream, a per block indication specific to the each inter-predicted block indicating either

to generate a prediction block for the each inter-predicted block based on a global motion model of the multiple global motion models and associated with one of the two or more reference frames, or

to generate the prediction block for the each inter-predicted block based on a motion vector that is different from the global motion model; and

decoding the each inter-predicted block based on the per block indication.

2. The method of claim 1 , wherein the multiple global motion models include transformations selected from a group comprising a translational motion model type, a similarity motion model type, an affine motion model type, and a homographic motion model type.

3. The method of claim 1 , wherein decoding the at least two global motion models comprises:

decoding respective global motion model types for at least some of the multiple global motion models.

4. The method of claim 3 , further comprising:

determining parameters of one of the at least some of the multiple global motion models based on a motion model type.

5. The method of claim 1 , wherein the segmentation of the current frame is based on motion analysis between the current frame and the two or more reference frames.

6. The method of claim 1 , further comprising:

decoding parameters of the global motion model of the multiple global motion models.

7. The method of claim 1 , wherein decoding the blocks of the current frame further comprises:

decoding at least one block of the segment using intra-prediction.

8. A device, comprising:

a processor, the processor configured to:

decode, from a compressed bitstream and with respect to each reference frame of two or more reference frames available for decoding a current frame, at least two global motion models, wherein decoding a global motion model comprises decoding parameters for the global motion model, wherein the global motion models constitute multiple global motion models, and wherein each global motion model is based on a segmentation of the current frame and represents a respective underlying motion of blocks within a respective segment; and

decode blocks of the current frame by configuration to:

for each inter-predicted block of a segment, decode, from the compressed bitstream, a per block indication specific to the each inter-predicted block indicating either

to generate a prediction block for the each inter-predicted block based on a global motion model of the multiple global motion models and associated with one of the two or more reference frames, or

to generate the prediction block for the each inter-predicted block based on a motion vector that is different from the global motion model; and

decode the each inter-predicted block based on the per block indication.

9. The device of claim 8 , wherein the multiple global motion models include transformations selected from a group comprising a translational motion model type, a similarity motion model type, an affine motion model type, and a homographic motion model type.

10. The device of claim 8 , wherein to decode the at least two global motion models comprises to:

decode respective global motion model types for at least some of the multiple global motion models.

11. The device of claim 10 , wherein the processor is further configured to:

determine parameters of one of the at least some of the multiple global motion models based on a motion model type.

12. The device of claim 8 , wherein the segmentation of the current frame is based on motion analysis between the current frame and the two or more reference frames.

13. The device of claim 8 , wherein the processor is further configured to:

decode parameters of the global motion model of the multiple global motion models.

14. The device of claim 8 , wherein to decode the blocks of the current frame further comprises to:

decode, using intra-prediction, at least one block of the segment.

15. A non-transitory computer-readable medium storing instructions operable to cause one or more processors to perform operations comprising:

decoding, from a compressed bitstream and with respect to each reference frame of two or more reference frames available for decoding a current frame, at least two global motion models, wherein decoding a global motion model comprises decoding parameters for the global motion model, wherein the global motion models constitute multiple global motion models, and wherein each global motion model is based on a segmentation of the current frame and represents a respective underlying motion of blocks within a respective segment; and

decoding blocks of the current frame by:

for each inter-predicted block of a segment, decoding, from the compressed bitstream, a per block indication specific to the each inter-predicted block indicating either

to generate a prediction block for the each inter-predicted block based on a global motion model of the multiple global motion models and associated with the one of the two or more reference frames, or

to generate a prediction block for the each inter-predicted block based on a motion vector that is different from the global motion model; and

decoding the each inter-predicted block based on the per block indication.

16. The non-transitory computer-readable medium of claim 15 , wherein the multiple global motion models include transformations selected from a group comprising a translational motion model type, a similarity motion model type, an affine motion model type, and a homographic motion model type.

17. The non-transitory computer-readable medium of claim 15 , wherein decoding the at least two global motion models comprises:

decoding respective global motion model types for at least some of the multiple global motion models.

18. The non-transitory computer-readable medium of claim 17 , wherein the operations further comprise:

determining parameters of one of the at least some of the multiple global motion models based on a motion model type.

19. The non-transitory computer readable medium of claim 15 , wherein the segmentation of the current frame is based on motion analysis between the current frame and the two or more reference frames.

20. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

decoding parameters of the global motion model of the multiple global motion models.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2023
From: MUKHERJEE, DEBARGHA; LIU, YUXIN; PARKER, SARAH
To: GOOGLE LLC
Reel/Frame 065765/0114 →
Continuity (4)
Continuation 16693425 · Nov 25, 2019
Continuation 15838748 · Dec 12, 2017
Provisional Application 62471659 · Mar 15, 2017
Related Publication 20240098298A1 · Mar 21, 2024
References Cited (57)
US 5510838A · Yomdin · 1996 [cited by examiner]
US 7200174B2 · Lainema · 2007 [cited by examiner]
US 9438910B1 · Han et al. · 2016 [cited by applicant]
US 9609343B1 · Chen · 2017 [cited by examiner]
US 10225573B1 · Mukherjee · 2019 [cited by examiner]
US 20010054989A1 · Zavracky · 2001 [cited by examiner]
US 20030081836A1 · Averbuch · 2003 [cited by examiner]
US 20030108099A1 · Nagumo · 2003 [cited by examiner]
US 20030174775A1 · Nagaya · 2003 [cited by examiner]
US 20030202596A1 · Lainema · 2003 [cited by examiner]
US 20060098886A1 · De Haan · 2006 [cited by examiner]
US 20060227865A1 · Sherigar · 2006 [cited by examiner]
US 20060285747A1 · Blake · 2006 [cited by examiner]
US 20070025444A1 · Okada · 2007 [cited by examiner]
US 20080144716A1 · De Haan · 2008 [cited by examiner]
US 20080240247A1 · Lee · 2008 [cited by examiner]
US 20090153730A1 · Knee · 2009 [cited by examiner]
US 20100085253A1 · Ferguson · 2010 [cited by examiner]
US 20110129016A1 · Sekiguchi · 2011 [cited by examiner]
US 20130028325A1 · Le Floch · 2013 [cited by examiner]
US 20130039422A1 · Kirchhoffer · 2013 [cited by examiner]
US 20130121416A1 · He · 2013 [cited by examiner]
US 20130294513A1 · Seregin · 2013 [cited by examiner]
US 20140029675A1 · Su · 2014 [cited by examiner]
US 20140133567A1 · Rusanovskyy · 2014 [cited by examiner]
US 20140192886A1 · Fran · 2014 [cited by examiner]
US 20140204228A1 · Yokokawa · 2014 [cited by examiner]
US 20150036737A1 · Puri · 2015 [cited by examiner]
US 20150325001A1 · Riemens · 2015 [cited by examiner]
US 20160330468A1 · Minezawa · 2016 [cited by examiner]
US 20160366435A1 · Chien · 2016 [cited by examiner]
US 20170070745A1 · Lee · 2017 [cited by examiner]
US 20170301096A1 · Weese · 2017 [cited by examiner]
US 20180199053A1 · Tsai · 2018 [cited by examiner]
US 20190082191A1 · Chuang · 2019 [cited by examiner]
US 20190158870A1 · Xu · 2019 [cited by examiner]
US 20200260111A1 · Liu · 2020 [cited by examiner]
EP 1404135A2 · 2004 [cited by applicant]
Bankoski et al., “VP8 Data Format and Decoding Guide”, Independent Submission RFC 6389, Nov. 2011, 305 pp. [cited by applicant]
Bankoski et al., “VP8 Data Format and Decoding Guide draft-bankoski-vp8-bitstream-02”, Network Working Group, Internet-Draft, May 18, 2011, 288 pp. [cited by applicant]
“Introduction to Video Coding Part 1: Transform Coding”, Mozilla, Mar. 2012, 171 pp. [cited by applicant]
“Overview VP7 Data Format and Decoder”, Version 1.5, On2 Technologies, Inc., Mar. 28, 2005, 65 pp. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Amendment 2: New profiles for professional applications, International Telecommunication Union, Apr. 2007, 75 … [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Amendment 1: Support of additional colour spaces and r… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Version 3, International Telecommunication Union, Mar.… [cited by applicant]
Bankoski, et al., “Technical Overview of VP8, An Open Source Video Codec for the Web”, Jul. 11, 2011, 6 pp. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Coding of moving video: Implementors Guide for H.264: Advanced video coding for generic audiovisual services, International Telecommunication Union, Jul. 30, 2010, 15 pp. [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, International Telecommunication Union, Version 11, Mar… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, International Telecommunication Union, Version 12, Mar… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Version 8, International Telecommunication Union, Nov.… [cited by applicant]
Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services, Version 1, International Telecommunication Union, May … [cited by applicant]
“VP8 Data Format and Decoding Guide, WebM Project”, Google On2, Dec. 1, 2010, 103 pp. [cited by applicant]
“VP6 Bitstream and Decoder Specification”, Version 1.02, On2 Technologies, Inc., Aug. 17, 2006, 88 pp. [cited by applicant]
“VP6 Bitstream and Decoder Specification”, Version 1.03, On2 Technologies, Inc., Oct. 29, 2007, 95 pp. [cited by applicant]
Dufaux F et al.; “Motion Estimation Techniques for Digital TV: A Review and A New Contribution”; Proceedings of the IEEE, New York; Jun. 1, 1995; pp. 858-876. [cited by applicant]
Moscheni et al; “A new two-stage global/ local motion estimation based on a background/foreground segmentation”, 1995 International Conference on Acoustics, Speech and Signal Processing; May 1995; pp. 2261-2264. [cited by applicant]
International Search Report and Written Opinion for Internation Application No. PCT/US2017/059306; dated Feb. 2, 2018. [cited by applicant]