IP Library Granted Patent US 12,563,246
Granted Patent B2
US 12,563,246 · App. 18/572,317 · Granted Feb 24, 2026

Method, apparatus, electronic device and storage medium for audio and video synchronization monitoring

Inventors: Wei Zhang (Beijing, CN); Xianhua Zeng (Beijing, CN)
Assignee: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
H04N21/2407H04N21/235
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,563,246
App. No.
18/572,317
Granted
Feb 24, 2026
Kind
B2
Abstract

The disclosure discloses a method, apparatus, electronic device, and storage medium for audio and video synchronization monitoring, wherein the method is applied to a data push streaming end includes: collecting audio data and video data to be pushed for streaming, and encoding the audio data and video data; selecting a video reference frame from the video data and adding supplemental enhancement information during an encoding process of the video reference frame, wherein the supplemental enhancement information comprises reference information for synchronized playback of the audio data and video data; and pushing the encoded audio data and video data into a target content distribution network for the data pull streaming end pulls the audio data and video data by a third-party server, and monitoring the audio data and the video data based on the additional enhancement information to achieve the synchronized playback.

Claims (48)

1 . A method for audio and video synchronization monitoring, wherein the method is applied to a data push streaming end and comprises:

collecting audio data and video data to be pushed for streaming, and encoding the audio data and video data;

selecting a video reference frame from the video data and adding supplemental enhancement information during an encoding process of the video reference frame, wherein the supplemental enhancement information comprises reference information for synchronized playback of the audio data and video data; and

pushing the encoded audio data and video data into a target content distribution network for the data pull streaming end pulls the audio data and video data by a third-party server, and monitoring the audio data and the video data based on the supplemental enhancement information to achieve the synchronized playback,

wherein the adding supplemental enhancement information during an encoding process of the video reference frame comprises:

determining an audio reference frame corresponding to the video reference frame; and

adding, as the supplemental enhancement information, a signature of the audio reference frame, an audio frame length, an audio data sampling rate and audio frame rendering time, a video data sampling rate of the video reference frame, and a video frame rendering time into encoded data of the reference video frame.

2 . A method for audio and video synchronization monitoring, wherein the method is applied to a data pull streaming end and comprises:

pulling audio data and video data to be played, and obtaining supplemental enhancement information of a video reference frame in the video data;

determining, based on the supplemental enhancement information, a rendering time of a video frame in the video data and a rendering time of an audio frame in the audio data; and

monitoring the video data and the audio data for synchronized playback based on the rendering time of the video frame in the video data and the rendering time of the audio frame in the audio data,

wherein the video data comprises a plurality of video frames, the audio data comprises a plurality of audio frames, and the determining a rendering time of a video frame in the video data and a rendering time of an audio frame in the audio data based on the supplemental enhancement information comprises:

determining a corresponding audio reference frame that matches the video reference frame based on an audio reference frame signature and an audio frame length in the supplemental enhancement information;

determining an audio frame rendering time of each audio frame in the audio data based on an audio data sampling rate and an audio frame rendering time of the audio reference frame in the supplemental enhancement information, and a sending timestamp of each audio frame in the audio data; and

determining a video frame rendering time of each video frame in the video data based on a video data sampling rate and a video frame rendering time of the video reference frame in the supplemental enhancement information, and a sending timestamp of each video frame in the video data.

3 . The method of claim 2 , wherein the determining a video frame rendering time of each video frame in the video data based on a video data sampling rate and a video frame rendering time of the video reference frame in the supplemental enhancement information, and a sending timestamp of each video frame in the video data comprises:

calculating, for each video frame in the video data, a first time difference between a sending timestamp of the each video frame and the video reference frame;

determining a first rendering time difference between the each video frame and the video reference frame based on the first time difference and the video data sampling rate; and

determining the video frame rendering time of each video frame by adding the first rendering time difference and the video frame rendering time of the video reference frame.

4 . The method of claim 3 , wherein the determining an audio frame rendering time of each audio frame in the audio data based on an audio data sampling rate and an audio frame rendering time of the audio reference frame in the supplemental enhancement information, and a sending timestamp of each audio frame in the audio data comprises:

calculating, for each audio frame in the audio data, a second time difference between a sending timestamp of the each audio frame and the audio reference frame;

determining a second rendering time difference between the each audio frame and the audio reference frame based on the second time difference and the audio data sampling rate; and

determining the video frame rendering time of the each audio frame by adding the second rendering time difference and the video frame rendering time of the audio reference frame.

5 . The method of claim 2 , wherein the monitoring the video data and the audio data for synchronized playback based on the rendering time of the video frame in the video data and the rendering time of the audio frame in the audio data comprises:

determining an arrival time difference of the video data relative to the audio data based on a video rendering time of a latest video frame in the video data and an arrival timestamp of the video data, and an audio frame rendering time of a latest audio frame in the audio data and an arrival timestamp of the audio data; and

monitoring the video data and the audio data for synchronized playback based on the arrival time difference.

6 . An electronic device comprising:

one or more processors; and

a storage apparatus configured to store one or more programs,

the one or more programs, when executed by the one or more processors, causing the one or more processors to implement a method of audio and video synchronization monitoring which is applied to a data pull streaming end, comprising:

pulling audio data and video data to be played, and obtaining supplemental enhancement information of a video reference frame in the video data;

determining, based on the supplemental enhancement information, a rendering time of a video frame in the video data and a rendering time of an audio frame in the audio data; and

monitoring the video data and the audio data for synchronized playback based on the rendering time of the video frame in the video data and the rendering time of the audio frame in the audio data,

wherein the video data comprises a plurality of video frames, the audio data comprises a plurality of audio frames, and the determining a rendering time of a video frame in the video data and a rendering time of an audio frame in the audio data based on the supplemental enhancement information comprises:

determining a corresponding audio reference frame that matches the video reference frame based on an audio reference frame signature and an audio frame length in the supplemental enhancement information;

determining an audio frame rendering time of each audio frame in the audio data based on an audio data sampling rate and an audio frame rendering time of the audio reference frame in the supplemental enhancement information, and a sending timestamp of each audio frame in the audio data; and

determining a video frame rendering time of each video frame in the video data based on a video data sampling rate and a video frame rendering time of the video reference frame in the supplemental enhancement information, and a sending timestamp of each video frame in the video data.

7 . The electronic device of claim 6 , wherein the determining a video frame rendering time of each video frame in the video data based on a video data sampling rate and a video frame rendering time of the video reference frame in the supplemental enhancement information, and a sending timestamp of each video frame in the video data comprises:

calculating, for each video frame in the video data, a first time difference between a sending timestamp of the each video frame and the video reference frame;

determining a first rendering time difference between the each video frame and the video reference frame based on the first time difference and the video data sampling rate; and

determining the video frame rendering time of each video frame by adding the first rendering time difference and the video frame rendering time of the video reference frame.

8 . The electronic device of claim 6 , wherein the determining an audio frame rendering time of each audio frame in the audio data based on an audio data sampling rate and an audio frame rendering time of the audio reference frame in the supplemental enhancement information, and a sending timestamp of each audio frame in the audio data comprises:

calculating, for each audio frame in the audio data, a second time difference between a sending timestamp of the each audio frame and the audio reference frame;

determining a second rendering time difference between the each audio frame and the audio reference frame based on the second time difference and the audio data sampling rate; and

determining the video frame rendering time of the each audio frame by adding the second rendering time difference and the video frame rendering time of the audio reference frame.

9 . The electronic device of claim 6 , wherein the monitoring the video data and the audio data for synchronized playback based on the rendering time of the video frame in the video data and the rendering time of the audio frame in the audio data comprises:

determining an arrival time difference of the video data relative to the audio data based on a video rendering time of a latest video frame in the video data and an arrival timestamp of the video data, and an audio frame rendering time of a latest audio frame in the audio data and an arrival timestamp of the audio data; and

monitoring the video data and the audio data for synchronized playback based on the arrival time difference.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2026
From: ZHANG, WEI
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073525/0337 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2026
From: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073525/0347 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2026
From: ZENG, XIANHUA
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 074459/0461 →
Priority Claims (1)
CN 202111241413.5 · Oct 25, 2021 · national
Continuity (1)
Related Publication 20240292044A1 · Aug 29, 2024
References Cited (28)
US 6642966B1 · Limaye · 2003 [cited by examiner]
US 20060272000A1 · Kwak · 2006 [cited by examiner]
US 20080170564A1 · Shi · 2008 [cited by examiner]
US 20150062353A1 · Dalal · 2015 [cited by examiner]
US 20150350716A1 · Kruglick · 2015 [cited by examiner]
US 20170111680A1 · Schneider et al. · 2017 [cited by applicant]
US 20180227164A1 · Wu et al. · 2018 [cited by applicant]
US 20200322670A1 · Zhang et al. · 2020 [cited by applicant]
US 20200336781A1 · Ramaswamy · 2020 [cited by applicant]
CN 103167320A · 2013 [cited by applicant]
CN 104410894A · 2015 [cited by applicant]
CN 109218794A · 2019 [cited by applicant]
CN 109660843A · 2019 [cited by applicant]
CN 110062277A · 2019 [cited by applicant]
CN 110234028A · 2019 [cited by applicant]
CN 110753202A · 2020 [cited by applicant]
CN 111294634A · 2020 [cited by applicant]
CN 111464256A · 2020 [cited by applicant]
CN 111654736A · 2020 [cited by applicant]
CN 112272327A · 2021 [cited by applicant]
CN 112291498A · 2021 [cited by applicant]
CN 112291498B · 2022 [cited by examiner]
JP 2010252151A · 2010 [cited by applicant]
WO 2017107516A1 · 2017 [cited by applicant]
WO 2020056877A1 · 2020 [cited by applicant]
English translation version of CN 112291498 B Nov. 4, 2022 (Year: 2022). [cited by examiner]
Chinese Office Action, issued in Chinese patent application No. 202111241413.5, dated Feb. 2, 2024, 15 pages (translation enclosed). [cited by applicant]
International Search Report (with English translation) and Written Opinion issued in PCT/CN2022/119419, dated Nov. 28, 2022, 11 pages provided. [cited by applicant]