IP Library Granted Patent US 12705838
Granted Patent B2
US 12705838 · App. 18/044,361 · Granted Aug 11, 2026

Virtual image displaying method and apparatus, electronic device and storage medium

Inventor: Liyou Xu (Beijing, CN)
Assignee: Beijing Bytedance Network Technology Co., Ltd.
G06T19/006G06T5/50H04N13/327G06T2207/20221G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705838
App. No.
18/044,361
Granted
Aug 11, 2026
Kind
B2
Abstract

Provided are a virtual image displaying method and apparatus, an electronic device, and a storage medium; by detecting a target object in a real-scene captured image captured by a terminal camera to determine a real capturing direction of the terminal camera according to a position of the target object in the real-scene captured image, calibrating, according to the real-scene capturing direction, a virtual capturing direction of a virtual camera, and performing, according to the calibrated virtual capturing direction, rendering processing on a shared virtual image corresponding to the target object to superimpose and display the processed shared virtual image on the real-scene captured image, the virtual capturing direction of the virtual camera is processed, which can be more suitable for a complex virtual image to making it can be applied to more interactive scenarios.

Claims (72)

1 . A virtual image displaying method, applied in a terminal, comprising:

obtaining a real-scene captured image captured by a terminal camera;

detecting a target object in the real-scene captured image;

determining, according to a position of the target object in the real-scene captured image, a real capturing direction of the terminal camera;

calibrating, according to the real capturing direction, a virtual capturing direction of a virtual camera;

obtaining a shared virtual image corresponding to the target object from a server; and

performing, according to the calibrated virtual capturing direction, rendering processing on the shared virtual image corresponding to the target object, to superimpose and display the processed shared virtual image on the real-scene captured image;

wherein the determining, according to the position of the target object in the real-scene captured image, the real capturing direction of the terminal camera comprises:

determining a depth of field of the target object in the real-scene captured image, wherein the depth of field is used to represent a distance of the target object relative to the terminal camera;

determining, according to the depth of field and the position of the target object in the real-scene captured image, a direction of the target object relative to the terminal camera; and

determining a current phase angle of the terminal camera, wherein the current phase angle of the terminal camera represents the distance of the target object relative to the terminal camera and the direction of the target object relative to the terminal camera;

wherein the obtaining the shared virtual image corresponding to the target object from the server comprises:

sending, to the server, an acquisition request for the shared virtual image, the acquisition request comprising the detected target object, for the server to retrieve, according to the target object comprised in the acquisition request, a corresponding scene image and a character image preset by at least one terminal user associated with the target object, wherein the at least one terminal user associated with the target object is a terminal user within a circular area with a geographical location of the target object as an origin and a preset distance as a radius, and the character image preset by the at least one terminal user is a character image established by the at least one terminal user prior to detecting the target object to reflect a user personality of the at least one terminal user; and

receiving, the shared virtual image delivered by the server,

wherein the shared virtual image is obtained by the server through image fusion processing on the corresponding scene image of the target object and the character image preset by the at least one terminal user associated with the target object.

2 . The displaying method according to claim 1 , wherein the calibrating, according to the real capturing direction, the virtual capturing direction of the virtual camera comprises:

calibrating, according to the current phase angle of the terminal camera, a phase angle of the virtual camera to obtain a calibrated phase angle of the virtual camera; wherein the calibrated phase angle of the virtual camera is consistent with the current phase angle of the terminal camera;

the performing, according to the calibrated virtual capturing direction, rendering processing on the shared virtual image corresponding to the target object comprises:

performing, according to the calibrated phase angle of the virtual camera, rendering processing on the shared virtual image.

3 . The displaying method according to claim 2 , wherein the calibrating, according to the current phase angle of the terminal camera, the phase angle of the virtual camera to obtain the calibrated phase angle of the virtual camera comprises:

determining, according to the current phase angle of the terminal camera, an offset matrix;

performing inverse transformation on the offset matrix to obtain an inverse transformation matrix; and

performing, by using the inverse transformation matrix, matrix transformation on an initial phase angle of the virtual camera to obtain the calibrated phase angle of the virtual camera.

4 . The displaying method according to claim 1 , wherein in the received shared virtual image delivered by the server, the corresponding scene image of the target object and the character image preset by the at least one terminal user associated with the target object are established in a same virtual coordinate system.

5 . The displaying method according to claim 4 , wherein the performing, according to the calibrated virtual capturing direction, rendering processing on the shared virtual image corresponding to the target object comprises:

performing, according to the calibrated virtual capturing direction, spatially rendering processing on spatial coordinates of the shared virtual image in the virtual coordinate system.

6 . The displaying method according to claim 1 , wherein the displaying method further comprises:

receiving, a control operation triggered on a corresponding character image displayed in the real-scene captured image, and uploading the control operation to the server, for the server to update the corresponding character image in the shared virtual image according to the control operation to obtain an updated shared virtual image; and

receiving, the updated shared virtual image delivered by the server, and performing rendering processing on the updated shared virtual image to superimpose and display the updated shared virtual image on the real-scene captured image.

7 . The displaying method according to claim 1 , wherein the acquisition request, sent to the server, for the shared virtual image further comprises a current geographical location, for the server to determine, according to a geographical location sharing range corresponding to the target object, whether the current geographical location belongs to the geographical location sharing range;

if it is determined that the current geographical location belongs to the geographical location sharing range, receiving the shared virtual image delivered by the server; if it is determined that the current geographical location does not belong to the geographical location sharing range, receiving a message of acquisition request failure delivered by the server.

8 . The displaying method according to claim 7 , wherein the geographical location sharing range corresponding to the target object is determined according to a geographical location of a terminal which detects the target object first.

9 . An electronic device, comprising:

at least one processor; and

a memory;

wherein the memory stores a computer execution instruction;

when the at least one processor executes the computer execution instruction stored in the memory, the one or more processors are caused to:

obtain a real-scene captured image captured by a terminal camera;

detect a target object in the real-scene captured image;

determine a real capturing direction of the terminal camera according to a position of the target object in the real-scene captured image; calibrate a virtual capturing direction of a virtual camera according to the real capturing direction; obtain a shared virtual image corresponding to the target object from a server; and perform rendering processing on the shared virtual image corresponding to the target object according to the calibrated virtual capturing direction; and

superimpose and display the processed shared virtual image on the real-scene captured image;

wherein the one or more processors are caused to:

determine a depth of field of the target object in the real-scene captured image, wherein the depth of field is used to represent a distance of the target object relative to the terminal camera;

determine, according to the depth of field and the position of the target object in the real-scene captured image, a direction of the target object relative to the terminal camera; and

determine a current phase angle of the terminal camera, wherein the current phase angle of the terminal camera represents the distance of the target object relative to the terminal camera and the direction of the target object relative to the terminal camera;

wherein the one or more processors are caused to:

send, to the server, an acquisition request for the shared virtual image, the acquisition request comprising the detected target object, for the server to retrieve, according to the target object comprised in the acquisition request, a corresponding scene image and a character image preset by at least one terminal user associated with the target object, wherein the at least one terminal user associated with the target object is a terminal user within a circular area with a geographical location of the target object as an origin and a preset distance as a radius, and the character image preset by the at least one terminal user is a character image established by the at least one terminal user prior to detecting the target object to reflect a user personality of the at least one terminal user; and

receive the shared virtual image delivered by the server,

wherein the shared virtual image is obtained by the server through image fusion processing on the corresponding scene image of the target object and the character image preset by the at least one terminal user associated with the target object.

10 . The electronic device according to claim 9 , wherein the one or more processors are caused to:

calibrate a phase angle of the virtual camera to obtain a calibrated phase angle of the virtual camera according to the current phase angle of the terminal camera; wherein the calibrated phase angle of the virtual camera is consistent with the current phase angle of the terminal camera; and

perform rendering processing on the shared virtual image according to the calibrated phase angle of the virtual camera.

11 . The electronic device according to claim 10 , wherein the one or more processors are caused to:

determine an offset matrix according to the current phase angle of the terminal camera;

performing inverse transformation on the offset matrix to obtain an inverse transformation matrix; and

perform matrix transformation on an initial phase angle of the virtual camera to obtain the calibrated phase angle of the virtual camera by using the inverse transformation matrix.

12 . The electronic device according to claim 9 , wherein in the received shared virtual image delivered by the server, the corresponding scene image of the target object and the character image preset by the at least one terminal user associated with the target object are established in a same virtual coordinate system.

13 . The electronic device according to claim 12 , wherein the one or more processors are caused to:

perform spatially rendering processing on spatial coordinates of the shared virtual image in the virtual coordinate system according to the calibrated virtual capturing direction.

14 . A non-transitory computer-readable storage medium having a computer execution instruction stored thereon, wherein when the computer execution instruction is executed by a processor, the processor is caused to:

obtain a real-scene captured image captured by a terminal camera;

detect a target object in the real-scene captured image;

determine a real capturing direction of the terminal camera according to a position of the target object in the real-scene captured image; calibrate a virtual capturing direction of a virtual camera according to the real capturing direction; obtain a shared virtual image corresponding to the target object from a server; and perform rendering processing on the shared virtual image corresponding to the target object according to the calibrated virtual capturing direction; and

superimpose and display the processed shared virtual image on the real-scene captured image;

wherein the one or more processors are caused to:

determine a depth of field of the target object in the real-scene captured image, wherein the depth of field is used to represent a distance of the target object relative to the terminal camera;

determine, according to the depth of field and the position of the target object in the real-scene captured image, a direction of the target object relative to the terminal camera; and

determine a current phase angle of the terminal camera, wherein the current phase angle of the terminal camera represents the distance of the target object relative to the terminal camera and the direction of the target object relative to the terminal camera;

wherein the one or more processors are caused to:

send, to the server, an acquisition request for the shared virtual image, the acquisition request comprising the detected target object, for the server to retrieve, according to the target object comprised in the acquisition request, a corresponding scene image and a character image preset by at least one terminal user associated with the target object, wherein the at least one terminal user associated with the target object is a terminal user within a circular area with a geographical location of the target object as an origin and a preset distance as a radius, and the character image preset by the at least one terminal user is a character image established by the at least one terminal user prior to detecting the target object to reflect a user personality of the at least one terminal user; and

receive the shared virtual image delivered by the server,

wherein the shared virtual image is obtained by the server through image fusion processing on the corresponding scene image of the target object and the character image preset by the at least one terminal user associated with the target object.