IP Library › Granted Patent US 12,033,241
Granted Patent B2
US 12,033,241 · App. 17/666,081 · Granted Jul 9, 2024

Scene interaction method and apparatus, electronic device, and computer storage medium

Inventor: Yuxuan Liang (Guangdong, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G06T11/00G06V10/44G06V20/20G06V20/41G06V20/46G06V40/16G06V40/20G06T2210/61G10L15/08H04L67/02H04L69/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,033,241
App. No.
17/666,081
Granted
Jul 9, 2024
Kind
B2
Abstract

A method for scene interaction includes identifying a first real scene interacting with a virtual scene and obtaining media information of the first real scene. The method also includes determining a scene feature associated with the first real scene based on a feature extraction on the media information, and mapping the scene feature associated with the first real scene to the virtual scene according to a correspondence between the virtual scene and the first real scene. Apparatus and non-transitory computer-readable storage medium counterpart embodiments are also contemplated.

Claims (58)

1. A method for scene interaction, the method comprising:

obtaining media information of a first environment scene, the media information including image information and audio information of the first environment scene;

determining, by processing circuitry of an electronic device, a scene feature associated with the first environment scene based on a feature extraction on the media information, the scene feature including (i) at least one of a scene image feature or a character image feature of the first environment scene and (ii) at least one audio feature extracted from the audio information of the first environment scene;

mapping the scene feature including (i) mapping the at least one of the scene image feature or the character image feature to at least one of a virtual background image or a virtual character image and (ii) mapping the at least one audio feature to at least one of a text content or a dynamic effect of a waveform audio feature; and

presenting (i) the at least one of the virtual background image or the virtual character image in a first region of a virtual scene and (ii) the at least one of the text content or the dynamic effect in a second region of the virtual scene, wherein the first region and the second region have a predefined positional relationship in the virtual scene.

2. The method according to claim 1 , wherein the scene feature comprises an image feature and the at least one audio feature.

3. The method according to claim 1 , wherein the determining the scene feature comprises:

performing an image feature extraction on the image information, to obtain an image feature of the first environment scene; and

performing an audio feature extraction on the audio information, to obtain the at least one audio feature of the first environment scene.

4. The method according to claim 3 , wherein the performing the image feature extraction on the image information comprises:

performing a scene recognition on the image information, to obtain the scene image feature of the first environment scene;

performing a face recognition on the image information, to obtain the character image feature of the first environment scene; and

performing a character action recognition on the image information, to obtain an action image feature of the first environment scene.

5. The method according to claim 3 , wherein the performing the image feature extraction on the image information comprises:

obtaining regional images of the first environment scene respectively corresponding to regional image acquisition parameters;

combining the regional images of a same time interval into an integrated image of the first environment scene; and

performing a feature extraction on the integrated image, to obtain the image feature of the first environment scene.

6. The method according to claim 5 , wherein the regional image acquisition parameters comprise at least one of an image acquisition angle or an image acquisition range.

7. The method according to claim 5 , wherein the performing the feature extraction on the integrated image comprises:

performing an edge detection on the integrated image, to obtain a feature region in the integrated image; and

performing a feature extraction on the feature region in the integrated image, to obtain the image feature of the first environment scene.

8. The method according to claim 3 , wherein the performing the audio feature extraction on the audio information comprises:

performing a speech recognition on the audio information, to obtain the text content of the first environment scene; and

performing a waveform detection on the audio information, to obtain the waveform audio feature of the first environment scene.

9. The method according to claim 1 , wherein the mapping the scene feature comprises:

determining, in the virtual scene, a feature mapping region corresponding to the first environment scene, according to the correspondence between the virtual scene and the first environment scene; and

presenting first scene content that maps with the scene feature of the first environment scene in the feature mapping region.

10. The method according to claim 1 , wherein the obtaining the media information comprises:

establishing a real-time communication link of a full-duplex communication protocol (WebSocket) based on a transmission control protocol (TCP) that performs a real-time communication between the processing circuitry of the electronic device and a media capturing device for the first environment scene; and

obtaining the media information of the first environment scene by using the real-time communication link.

11. The method according to claim 1 , wherein and the predefined positional relationship defines that the first region and the second region are completely overlapping.

12. The method according to claim 1 , wherein the predefined positional relationship defines that the first region and the second region are non-overlapping.

13. An apparatus, comprising:

processing circuitry configured to:

obtain media information of a first environment scene, the media information including image information and audio information of the first environment scene;

determine a scene feature associated with the first environment scene based on a feature extraction on the media information, the scene feature including (i) at least one of a scene image feature or a character image feature of the first environment scene and (ii) at least one audio feature extracted from the audio information of the first environment scene;

map the scene feature including (i) map the at least one of the scene image feature of the character image feature to at least one of a virtual background image or a virtual character image and (ii) map the at least one audio feature to at least one of a text content or a dynamic effect of a waveform audio feature; and

present (i) the at least one of the virtual background image or the virtual character image in a first region of a virtual scene and (ii) the at least one of the text content or the dynamic effect in a second region of the virtual scene, wherein the first region and the second region have a predefined positional relationship in the virtual scene.

14. The apparatus according to claim 13 , wherein the processing circuitry is configured to:

perform image feature extraction on the image information, and an audio feature extraction on the audio information.

15. The apparatus according to claim 13 , wherein the processing circuitry is configured to:

obtain regional images of the first environment scene;

combine the regional images in a same time interval into an integrated image of the first environment scene; and

perform a feature extraction on the integrated image, to obtain an image feature of the first environment scene.

16. The apparatus according to claim 15 , wherein the processing circuitry is configured to:

perform an edge detection on the integrated image, to obtain a feature region in the integrated image; and

perform a feature extraction on the feature region in the integrated image, to obtain the image feature of the first environment scene.

17. The apparatus according to claim 13 , wherein the processing circuitry is configured to:

perform a speech recognition on the audio information of the first environment scene, to obtain the text content of the first environment scene; and

perform a waveform detection on the audio information, to obtain the waveform audio feature of the first environment scene.

18. The apparatus according to claim 13 , wherein the processing circuitry is configured to:

determine, in the virtual scene, a feature mapping region corresponding to the first environment scene, according to the correspondence between the virtual scene and the first environment scene; and

present first scene content that maps with the scene feature of the first environment scene in the feature mapping region.

19. A non-transitory computer-readable medium storing instructions which when executed by a computer cause the computer to perform:

obtaining media information of a first environment scene, the media information including image information and audio information of the first environment scene;

determining a scene feature associated with the first environment scene based on a feature extraction on the media information, the scene feature including (i) at least one of a scene image feature or a character image feature of the first environment scene and (ii) at least one audio feature extracted from the audio information of the first environment scene;

mapping the scene feature including (i) mapping the at least one of the scene image feature of the character image feature to at least one of a virtual background image or a virtual character image and (ii) mapping the at least one audio feature to at least one of a text content or a dynamic effect of a waveform audio feature; and

presenting (i) the at least one of the virtual background image or the virtual character image in a first region of a virtual scene and (ii) the at least one of the text content or the dynamic effect in a second region of the virtual scene, wherein the first region and the second region have a predefined positional relationship in the virtual scene.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2022
From: LIANG, YUXUAN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 058963/0471 →
Priority Claims (1)
CN 202010049112.1 · Jan 16, 2020 · national
Continuity (2)
Continuation PCTCN2020127750 · Nov 10, 2020
Related Publication 20220156986A1 · May 19, 2022