IP Library Granted Patent US 12713102
Granted Patent B2
US 12713102 · App. 18/042,259 · Granted Aug 18, 2026

Subtitle rendering method and device for virtual reality space, equipment and medium

Inventor: Yepeng Chen (Beijing, CN)
Assignee: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
H04N21/4884G06F3/013G06T17/00H04N21/4312
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12713102
App. No.
18/042,259
Granted
Aug 18, 2026
Kind
B2
Abstract

The embodiment of the disclosure relates to a caption rendering method and device for a virtual reality space, equipment and a medium, and the method includes the steps: separating caption content and picture content on a currently played virtual reality video frame, and mapping and rendering the picture content to a virtual reality panoramic space; determining a target space position in the virtual reality panoramic space according to the current sight direction of the user; and rendering the subtitle content at the target spatial position to generate spatial subtitles.

Claims (80)

1 . A method for caption rendering in a virtual reality space, comprising:

separating a caption content and a picture content on a currently displayed virtual reality (VR) video frame, and mapping and rendering the picture content to a VR panoramic space;

determining a target spatial position in the VR panoramic space according to a user's current LOS (line-of-sight) direction, wherein the user's current LOS direction is a ray starting from the user; and

rendering the caption content at the target spatial position to generate a spatial caption,

wherein determining a target spatial position in the VR panoramic space according to a user's current LOS direction comprises:

determining an initial position of a VR device worn by the user in the VR panoramic space as a center point position of the VR panoramic space;

still taking the initial position as a center point position of the VR panoramic space in response to the VR device being moved;

obtaining a preset radius distance; and

starting from the center point position, taking a position extending to the preset radius distance in the user's current LOS direction as the target spatial position.

2 . The method according to claim 1 , wherein rendering the caption content at the target spatial position to generate a spatial caption comprises:

displaying a background interface matching the caption content synchronously, in response to the caption content input in real time in a corresponding area of the target spatial position, wherein the background interface changes with the caption content input in real time.

3 . The method according to claim 2 , wherein displaying a background interface matching the caption content comprises synchronously, in response to a caption content input in real time in a corresponding area of the target spatial position:

displaying a background interface having a real-time width matching the caption content according to a preset per-caption-unit width, a per-caption-unit-background width, and a width of a caption input in real time synchronously, in response to a width change of the caption content input in real time in the corresponding area of the target spatial position; and/or

displaying a background interface having a real-time height matching the caption content according to a preset per-caption-unit height, a per-caption-unit-background height, and a height of a caption input in real-time synchronously, in response to a height change of the caption content input in real time in the corresponding area of the target spatial position.

4 . The method according to claim 1 , wherein determining a target spatial position in the VR panoramic space according to a user's current LOS direction comprises:

acquiring a historical spatial position corresponding to a caption content of a previous frame displayed in the VR panoramic space;

acquiring LOS change information of the user's current LOS direction with respect to the user's LOS direction when the previous frame is viewed; and

determining the target spatial position according to the LOS change information and the historical spatial position.

5 . The method according to claim 4 , wherein acquiring LOS change information of the user's current LOS direction with respect to the user's LOS direction when the previous frame is viewed comprises:

acquiring a horizontal axis rotation angle of a camera in the VR device relative to the previous frame in a horizontal direction, wherein the horizontal axis rotation angle is change information from the user's horizontal LOS direction when the previous frame is viewed to the user's horizontal LOS direction when the currently displayed virtual reality video frame is viewed.

6 . The method according to claim 5 , wherein determining the target spatial position according to the LOS change information and the historical spatial position comprises:

acquiring a center position preset in the VR panoramic space, and using the horizontal axis rotation angle as a center rotation angle from the previous frame to the currently displayed VR video frame;

determining a historical spatial position of the caption content of the previous frame relative to the center position; and

determining the target spatial position according to the center position, the historical spatial position, and the center rotation angle.

7 . The method according to claim 6 , further comprising:

acquiring an initial position of the VR device in the VR panoramic space, wherein the initial position is a center point position in the VR panoramic space; and

using the center point position as the center position.

8 . The method according to claim 7 , further comprising:

using the center point position as the center position constantly in a case that the VR device is moved in the VR panoramic space.

9 . The method according to claim 6 , further comprising:

obtaining a preset radius distance; and

determining an initial spatial position of a caption content of an initial frame relative to the center position according to the center position, the radius distance, and the user's initial LOS direction.

10 . The method according to claim 6 , further comprising:

providing an engine architecture for spatial caption rendering, wherein the engine architecture comprises a camera node, a rotation root node, and an interface group node, the rotation root node is a parent node of the interface group node, the camera node and the rotation root node are nodes at the same level, and the interface group node varies with a center position and/or center angle of the rotation root node through a parent-child relationship;

wherein:

invoking the camera node to acquire the horizontal axis rotation angle of the camera relative to the previous frame in a case that frame switching occurs, and using the horizontal axis rotation angle as the center rotation angle relative to the previous frame in a case that frame switching occurs;

invoking the rotation root node to acquire a center position in real-time and the center rotation angle relative to the previous frame in a case that frame switching occurs; and

invoking the interface group node to acquire a real-time position of a caption content of each frame relative to the center position.

11 . The method according to claim 10 , wherein the engine architecture comprises:

a caption content node and a caption background node, wherein the interface group node is the parent node of the caption content node and the caption background node, the caption content node and the caption background node are nodes at the same level; the caption content node and the caption background node surround the rotation root node following the interface group node and facing the center,

the method further comprising:

invoking the caption content node to obtain real-time input caption information; and

invoking the caption background node to obtain background information matching the real-time input caption information.

12 . An electronic device, comprising:

a processor;

a memory for storing processor executable instructions;

wherein the processor is used to read the executable instructions from the memory and execute the instructions to implement the following steps of:

separating a caption content and a picture content on a currently displayed virtual reality (VR) video frame, and mapping and rendering the picture content to a VR panoramic space;

determining a target spatial position in the VR panoramic space according to a user's current LOS (line-of-sight) direction, wherein the user's current LOS direction is a ray starting from the user; and

rendering the caption content at the target spatial position to generate a spatial caption,

wherein determining a target spatial position in the VR panoramic space according to a user's current LOS direction comprises:

determining an initial position of a VR device worn by the user in the VR panoramic space as a center point position of the VR panoramic space;

still taking the initial position as a center point position of the VR panoramic space in response to the VR device being moved;

obtaining a preset radius distance; and

starting from the center point position, taking a position extending to the preset radius distance in the user's current LOS direction as the target spatial position.

13 . The electronic device according to claim 12 , wherein the processor is configured to:

acquire a historical spatial position corresponding to a caption content of a previous frame displayed in the VR panoramic space;

acquire LOS change information of the user's current LOS direction with respect to the user's LOS direction when the previous frame is viewed; and

determine the target spatial position according to the LOS change information and the historical spatial position.

14 . The electronic device according to claim 13 , wherein the processor is configured to:

acquire a horizontal axis rotation angle of a camera in the VR device relative to the previous frame in a horizontal direction, wherein the horizontal axis rotation angle is change information from the user's horizontal LOS direction when the previous frame is viewed to the user's horizontal LOS direction when the currently displayed virtual reality video frame is viewed.

15 . The electronic device according to claim 14 , wherein the processor is configured to:

acquire a center position preset in the VR panoramic space, and using the horizontal axis rotation angle as a center rotation angle from the previous frame to the currently displayed VR video frame;

determine a historical spatial position of the caption content of the previous frame relative to the center position; and

determine the target spatial position according to the center position, the historical spatial position, and the center rotation angle.

16 . The electronic device according to claim 15 , wherein the processor is configured to:

acquire an initial position of the VR device in the VR panoramic space, wherein the initial position is a center point position in the VR panoramic space; and

use the center point position as the center position.

17 . The electronic device according to claim 15 , wherein the processor is configured to:

obtain a preset radius distance; and

determine an initial spatial position of a caption content of an initial frame relative to the center position according to the center position, the radius distance, and the user's initial LOS direction.

18 . A non-transitory computer-readable storage medium on which a computer program is stored, wherein the computer program is used to perform the following steps of:

separating a caption content and a picture content on a currently displayed virtual reality (VR) video frame, and mapping and rendering the picture content to a VR panoramic space;

determining a target spatial position in the VR panoramic space according to a user's current LOS (line-of-sight) direction, wherein the user's current LOS direction is a ray starting from the user; and

rendering the caption content at the target spatial position to generate a spatial caption,

wherein determining a target spatial position in the VR panoramic space according to a user's current LOS direction comprises:

determining an initial position of a VR device worn by the user in the VR panoramic space as a center point position of the VR panoramic space;

still taking the initial position as a center point position of the VR panoramic space in response to the VR device being moved;

obtaining a preset radius distance; and

starting from the center point position, taking a position extending to the preset radius distance in the user's current LOS direction as the target spatial position.