IP Library Granted Patent US 12695846
Granted Patent B2
US 12695846 · App. 18/438,615 · Granted Jul 28, 2026

Video processing method and apparatus, computer, and readable storage medium

Inventor: Hongtao Zuo (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
H04N7/152G06V10/56H04N7/155
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12695846
App. No.
18/438,615
Granted
Jul 28, 2026
Kind
B2
Abstract

A video processing method includes: obtaining video stream data from N video acquisition devices, N being a positive integer, and each video acquisition device corresponding to one visual angle; stitching N pieces of video stream data into a video synthesis picture based on region division information for video picture synthesis, the region division information indicating locations of the N pieces of video stream data in the video synthesis picture; and transmitting the video synthesis picture to a participating client, to enable the participating client to split the video synthesis picture into the N pieces of video stream data based on the region division information, and perform synchronous rendering on the N pieces of video stream data based on video display region information corresponding to the participating client.

Claims (93)

1 . A video processing method, performed by a video processing device, and comprising:

obtaining video stream data from N video acquisition devices, N being a positive integer, and each video acquisition device corresponding to one visual angle;

stitching N pieces of video stream data into a video synthesis picture based on region division information for video picture synthesis, the region division information indicating locations of the N pieces of video stream data in the video synthesis picture, the region division information including N association relationships each associating a device identifier of a respective one of the N video acquisition devices with a corresponding coordinate-based video region location in the video synthesis picture; and

transmitting the video synthesis picture and the region division information to a participating client, to enable the participating client to split the video synthesis picture into the N pieces of video stream data based on the region division information, and perform synchronous rendering on the N pieces of video stream data based on video display region information corresponding to the participating client.

2 . The method according to claim 1 , further comprising:

obtaining a video playing frame rate, and determining a video frame switching periodicity based on the video playing frame rate; and

the obtaining the video stream data from the N video acquisition devices comprising:

deleting historical video stream data associated with the N video acquisition devices, and associating and storing the N pieces of video stream data and the N video acquisition devices within the video frame switching periodicity, in response to receiving the video stream data transmitted by the N video acquisition devices; and

obtaining, in response to that a video frame switching condition indicated by the video frame switching periodicity is satisfied, video stream data that is currently associated with the N video acquisition devices.

3 . The method according to claim 1 , further comprising:

obtaining device acquisition information and coverage regions corresponding to the N video acquisition devices, and determining a primary device in the N video acquisition devices based on the device acquisition information and the coverage regions corresponding to the N video acquisition devices; and

wherein the obtaining the video stream data from the N video acquisition devices comprises:

obtaining video stream data from a video acquisition device other than the primary device in the N video acquisition devices, in response to receiving video stream data transmitted by the primary device.

4 . The method according to claim 1 , further comprising:

establishing data connections to the N video acquisition devices, and determining, based on data transmission losses corresponding to the data connections and media streaming manners, a target media streaming manner from the media streaming manners; and

transmitting the target media streaming manner to the N video acquisition devices; and

wherein the obtaining the video stream data from the N video acquisition devices comprises:

obtaining video stream data transmitted by the N video acquisition devices based on the target media streaming manner.

5 . The method according to claim 1 , wherein the region division information comprises video region locations corresponding to the N video acquisition devices; and

the stitching the N pieces of video stream data into the video synthesis picture based on the region division information comprises:

obtaining image data and audio data that are comprised in each piece of video stream data, and stitching the image data comprised in the N pieces of video stream data into a synthetic image based on the video region locations corresponding to the N video acquisition devices;

associating the audio data corresponding to N pieces of image data with the N pieces of image data in the synthetic image, to obtain the video synthesis picture.

6 . The method according to claim 1 , further comprising:

obtaining image synthesis data and audio synthesis data that form the video synthesis picture;

obtaining d pixels comprised in the image synthesis data and pixel color value information corresponding to each pixel of the d pixels, and obtaining color value difference data between pixel color value information corresponding to every two adjacent pixels in the d pixels, d being a positive integer;

dividing the d pixels into k pixel sets based on the color value difference data between the pixel color value information corresponding to every two adjacent pixels, k being a positive integer less than or equal to d, and pixels comprised in each pixel set being consecutive in the image synthesis data;

forming, by using pixel color value information of a second pixel in each pixel set and color value difference data between a first pixel and a previous pixel of the first pixel, image coded data corresponding to the image synthesis data, the second pixel being a leading pixel in a corresponding pixel set, and the first pixel being a pixel other than the second pixel in each pixel set;

performing audio coding processing on the audio synthesis data, to obtain audio coded data; and

transmitting video synthesis coded data formed by the image coded data and the audio coded data to the participating client, to enable the participating client to perform a decoding operation on the video synthesis coded data, to obtain the video synthesis picture.

7 . The method according to claim 1 , further comprising:

obtaining device acquisition information corresponding to the N video acquisition devices, and determining, based on the device acquisition information corresponding to the N video acquisition devices, device priorities corresponding to the N video acquisition devices;

determining a video synthesis size based on acquisition resolution corresponding to the N video acquisition devices;

determining, based on the device priorities corresponding to the N video acquisition devices and the video synthesis size, video region locations corresponding to the N video acquisition devices;

determining the region division information based on device identifiers corresponding to the N video acquisition devices and the video region locations corresponding to the N video acquisition devices; and

transmitting the region division information to the participating client.

8 . A video processing method, performed by a participating client, and comprising:

obtaining a video synthesis picture and region division information from a video processing device;

splitting the video synthesis picture into N pieces of video stream data based on the region division information corresponding to video picture synthesis performed by the video processing device, N being a positive integer, the region division information indicating locations of the N pieces of video stream data in the video synthesis picture, the video synthesis picture being obtained by stitching the N pieces of video stream data by the video processing device based on the region division information, the N pieces of video stream data being obtained from N video acquisition devices, and each video acquisition device corresponding to one visual angle, the region division information including N association relationships each associating a device identifier of a respective one of the N video acquisition devices with a corresponding coordinate-based video region location in the video synthesis picture; and

performing synchronous rendering on the N pieces of video stream data based on video display region information.

9 . The method according to claim 8 , further comprising:

obtaining a currently displayed livestreaming page, and obtaining M livestreaming windows comprised in the livestreaming page and livestreaming visual angles corresponding to the M livestreaming windows, M being a positive integer; and

determining the video display region information based on video acquisition devices corresponding to the M livestreaming visual angles, the video display region information indicating video acquisition devices corresponding to the M livestreaming windows; and

the performing the synchronous rendering on the N pieces of video stream data based on the video display region information comprising:

determining video stream data corresponding to each livestreaming window from the N pieces of video stream data based on the video display region information; and

rendering the video stream data corresponding to the livestreaming window in the corresponding livestreaming window based on the video stream data corresponding to each livestreaming window.

10 . The method according to claim 9 , the M livestreaming windows comprising a primary window and at least one secondary window, each secondary window being a livestreaming window other than the primary window in the M livestreaming windows; and

the method further comprising:

switching from displaying of video stream data in the primary window to displaying of video stream data in a first livestreaming window in the at least one secondary window, in response to a primary picture switching request for the first livestreaming window, content displayed in the primary window after the switching being the same as content displayed in the first livestreaming window.

11 . The method according to claim 9 , further comprising:

obtaining a first video acquisition device corresponding to a second livestreaming window in the N video acquisition devices, in response to a playing request for the second livestreaming window, the second livestreaming window not belonging to the M livestreaming windows; and

obtaining video stream data corresponding to the first video acquisition device from the N pieces of video stream data, and outputting the video stream data corresponding to the first video acquisition device in the second livestreaming window.

12 . The method according to claim 8 , the obtaining the video synthesis picture from the video processing device comprising:

obtaining video synthesis coded data from the video processing device, and obtaining image coded data and audio coded data that form the video synthesis coded data;

obtaining k pixel sets of the image coded data, k being a positive integer;

determining, based on pixel color value information of corresponding second pixels in the k pixel sets and color value difference data between a first pixel and a previous pixel of the first pixel, pixel color value information corresponding to d pixels comprised in the k pixel sets, d being a positive integer greater than or equal to k, the second pixel being the leading pixel in a corresponding pixel set, and the first pixel being a pixel other than the second pixel in each pixel set:

forming image synthesis data by using the pixel color value information corresponding to the d pixels;

performing audio decoding processing on the audio coded data, to obtain audio synthesis data; and

forming the video synthesis picture by using the image synthesis data and the audio synthesis data.

13 . A non-transitory computer-readable storage medium, storing a computer program, the computer program being applied to be loaded and executed by at least one processor, to enable a computer device that comprises the at least one processor to perform:

obtaining video stream data from N video acquisition devices, N being a positive integer, and each video acquisition device corresponding to one visual angle;

stitching N pieces of video stream data into a video synthesis picture based on region division information for video picture synthesis, the region division information indicating locations of the N pieces of video stream data in the video synthesis picture, the region division information including N association relationships each associating a device identifier of a respective one of the N video acquisition devices with a corresponding coordinate-based video region location in the video synthesis picture; and

transmitting the video synthesis picture and the region division information to a participating client, to enable the participating client to split the video synthesis picture into the N pieces of video stream data based on the region division information, and perform synchronous rendering on the N pieces of video stream data based on video display region information corresponding to the participating client.

14 . The storage medium according to claim 13 , wherein the computer program further causes the at least one processor to perform:

obtaining a video playing frame rate, and determining a video frame switching periodicity based on the video playing frame rate; and

the obtaining the video stream data from the N video acquisition devices comprising:

deleting historical video stream data associated with the N video acquisition devices, and associating and storing the N pieces of video stream data and the N video acquisition devices within the video frame switching periodicity, in response to receiving the video stream data transmitted by the N video acquisition devices; and

obtaining, in response to that a video frame switching condition indicated by the video frame switching periodicity is satisfied, video stream data that is currently associated with the N video acquisition devices.

15 . The storage medium according to claim 13 , wherein the computer program further causes the at least one processor to perform:

obtaining device acquisition information and coverage regions corresponding to the N video acquisition devices, and determining a primary device in the N video acquisition devices based on the device acquisition information and the coverage regions corresponding to the N video acquisition devices; and

wherein the obtaining the video stream data from the N video acquisition devices comprises:

obtaining video stream data from a video acquisition device other than the primary device in the N video acquisition devices, in response to receiving video stream data transmitted by the primary device.

16 . The storage medium according to claim 13 , wherein the computer program further causes the at least one processor to perform:

establishing data connections to the N video acquisition devices, and determining, based on data transmission losses corresponding to the data connections and media streaming manners, a target media streaming manner from the media streaming manners; and

transmitting the target media streaming manner to the N video acquisition devices; and

wherein the obtaining the video stream data from the N video acquisition devices comprises:

obtaining video stream data transmitted by the N video acquisition devices based on the target media streaming manner.

17 . The storage medium according to claim 13 , wherein the region division information comprises video region locations corresponding to the N video acquisition devices.

18 . The storage medium according to claim 17 , wherein the stitching the N pieces of video stream data into the video synthesis picture based on the region division information comprises:

obtaining image data and audio data that are comprised in each piece of video stream data, and stitching the image data comprised in the N pieces of video stream data into a synthetic image based on the video region locations corresponding to the N video acquisition devices;

associating the audio data corresponding to N pieces of image data with the N pieces of image data in the synthetic image, to obtain the video synthesis picture.

19 . The storage medium according to claim 13 , wherein the computer program further causes the at least one processor to perform:

obtaining image synthesis data and audio synthesis data that form the video synthesis picture;

obtaining d pixels comprised in the image synthesis data and pixel color value information corresponding to each pixel of the d pixels, and obtaining color value difference data between pixel color value information corresponding to every two adjacent pixels in the d pixels, d being a positive integer;

dividing the d pixels into k pixel sets based on the color value difference data between the pixel color value information corresponding to every two adjacent pixels, k being a positive integer less than or equal to d, and pixels comprised in each pixel set being consecutive in the image synthesis data;

forming, by using pixel color value information of a second pixel in each pixel set and color value difference data between a first pixel and a previous pixel of the first pixel, image coded data corresponding to the image synthesis data, the second pixel being a leading pixel in a corresponding pixel set, and the first pixel being a pixel other than the second pixel in each pixel set;

performing audio coding processing on the audio synthesis data, to obtain audio coded data; and

transmitting video synthesis coded data formed by the image coded data and the audio coded data to the participating client, to enable the participating client to perform a decoding operation on the video synthesis coded data, to obtain the video synthesis picture.

20 . The storage medium according to claim 13 , wherein the computer program further causes the at least one processor to perform:

obtaining device acquisition information corresponding to the N video acquisition devices, and determining, based on the device acquisition information corresponding to the N video acquisition devices, device priorities corresponding to the N video acquisition devices;

determining a video synthesis size based on acquisition resolution corresponding to the N video acquisition devices;

determining, based on the device priorities corresponding to the N video acquisition devices and the video synthesis size, video region locations corresponding to the N video acquisition devices;

determining the region division information based on device identifiers corresponding to the N video acquisition devices and the video region locations corresponding to the N video acquisition devices; and

transmitting the region division information to the participating client.