IP Library › Granted Patent US 12,307,545
Granted Patent B2
US 12,307,545 · App. 17/985,578 · Granted May 20, 2025

Method, electronic device, and computer program product for processing virtual avatar

Inventors: Zijia Wang (WeiFang, CN); Zhisong Liu (Shenzhen, CN); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06T1/0021G06T13/40G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,545
App. No.
17/985,578
Granted
May 20, 2025
Kind
B2
Abstract

Embodiments of the present disclosure relate to a method, an electronic device, and a computer program product for processing a virtual avatar. The method includes generating an image feature of the virtual avatar based on a plurality of image blocks of the virtual avatar and corresponding positions of the plurality of image blocks in the virtual avatar. The method further includes generating, based on a watermark to be added to the virtual avatar, a text feature associated with text of the watermark. The method further includes generating a watermarked virtual avatar based on the image feature and the text feature, wherein the watermark is invisible to human beings and identifies an identity of a user of the virtual avatar in a metaverse. Through embodiments of the present disclosure, identification of a virtual avatar, illustratively for verification and/or tracking purposes can be achieved without affecting the appearance of the virtual avatar.

Claims (88)

1. A method of processing a virtual avatar, comprising:

generating an image feature of the virtual avatar based on a plurality of image blocks of the virtual avatar and corresponding positions of the plurality of image blocks in the virtual avatar;

generating, based on a watermark to be added to the virtual avatar, a text feature associated with text of the watermark; and

generating a watermarked virtual avatar based on the image feature and the text feature, wherein the watermark is invisible to human beings and identifies an identity of a user of the virtual avatar in a metaverse;

wherein generating the watermarked virtual avatar based on the image feature and the text feature comprises:

determining a key-value pair based on the image feature and the text feature, the key-value pair comprising (i) a key that represents a semantic embedding parameter of a self-attention mechanism in a machine learning model, and (ii) a value that represents the text feature;

determining the image feature as a query vector;

determining a weight set based on a similarity between the query vector and a key in the key-value pair; and

determining, based on the weight set, a watermarked image feature;

wherein generating the watermarked virtual avatar further comprises processing the watermarked virtual avatar through an invariant layer of the machine learning model, the invariant layer comprising at least one fully-connected layer of the machine learning model and implementing a mapping function that converts a first number of image channels of an input image space of the watermarked virtual avatar into a second number of image channels of a transformed image space of an updated watermarked virtual avatar, the second number being greater than the first number.

2. The method according to claim 1 , wherein generating the image feature of the virtual avatar comprises:

dividing the virtual avatar into a plurality of visual marks, wherein each visual mark in the plurality of visual marks is represented by a vector of a predetermined length;

generating a corresponding position code for each of the visual marks; and

generating the image feature based on the plurality of visual marks and the corresponding position codes.

3. The method according to claim 1 , wherein generating the text feature associated with the text of the watermark comprises:

dividing the watermark into a plurality of text lemmas, wherein each text lemma in the plurality of text lemmas comprises a word; and

generating the text feature based on the plurality of text lemmas.

4. The method according to claim 1 , wherein the key-value pair is a first key-value pair, the query vector is a first query vector, the weight set is a first weight set, and the method further comprises:

determining a second key-value pair based on the watermarked image feature and the text feature;

determining the watermarked image feature as a second query vector;

determining a second weight set based on a similarity between the second query vector and a key in the second key-value pair; and

determining, based on the second weight set, an updated watermarked image feature.

5. The method according to claim 4 , wherein generating the watermarked virtual avatar based on the image feature and the text feature comprises:

generating the watermarked virtual avatar based on the updated watermarked image feature and the text feature.

6. The method according to claim 1 , further comprising:

sending the updated watermarked virtual avatar to a metaverse platform; and

sending an instruction for extracting the watermark in the watermarked virtual avatar to the metaverse platform.

7. The method according to claim 6 , further comprising:

causing, in response to the watermark being extracted, the metaverse platform to verify the watermark based on information of the user; and

causing, based on a result of the verification of the watermark, the metaverse platform to determine at least one of authenticity and ownership of the virtual avatar.

8. The method according to claim 1 , wherein the method is performed by a trained machine learning model, and the method further comprises:

training the machine learning model by using sample data, wherein the sample data comprises the virtual avatar, the watermark and a corresponding labeled watermarked virtual avatar, and a labeled extracted watermark.

9. An electronic device, comprising:

a processor; and

a memory coupled to the processor, wherein the memory has instructions stored therein which, when executed by the processor, cause the electronic device to execute actions comprising:

generating an image feature of a virtual avatar based on a plurality of image blocks of the virtual avatar and corresponding positions of the plurality of image blocks in the virtual avatar;

generating, based on a watermark to be added to the virtual avatar, a text feature associated with text of the watermark; and

generating a watermarked virtual avatar based on the image feature and the text feature, wherein the watermark is invisible to human beings and identifies an identity of a user of the virtual avatar in a metaverse;

wherein generating the watermarked virtual avatar based on the image feature and the text feature comprises:

determining a key-value pair based on the image feature and the text feature, the key-value pair comprising (i) a key that represents a semantic embedding parameter of a self-attention mechanism in a machine learning model, and (ii) a value that represents the text feature;

determining the image feature as a query vector;

determining a weight set based on a similarity between the query vector and a key in the key-value pair; and

determining, based on the weight set, a watermarked image feature;

wherein generating the watermarked virtual avatar further comprises processing the watermarked virtual avatar through an invariant layer of the machine learning model, the invariant layer comprising at least one fully-connected layer of the machine learning model and implementing a mapping function that converts a first number of image channels of an input image space of the watermarked virtual avatar into a second number of image channels of a transformed image space of an updated watermarked virtual avatar, the second number being greater than the first number.

10. The electronic device according to claim 9 , wherein generating the image feature of the virtual avatar comprises:

dividing the virtual avatar into a plurality of visual marks, wherein each visual mark in the plurality of visual marks is represented by a vector of a predetermined length;

generating a corresponding position code for each of the visual marks; and

generating the image feature based on the plurality of visual marks and the corresponding position codes.

11. The electronic device according to claim 9 , wherein generating the text feature associated with the text of the watermark comprises:

dividing the watermark into a plurality of text lemmas, wherein each text lemma in the plurality of text lemmas comprises a word; and

generating the text feature based on the plurality of text lemmas.

12. The electronic device according to claim 9 , wherein the key-value pair is a first key-value pair, the query vector is a first query vector, the weight set is a first weight set, and the actions further comprise:

determining a second key-value pair based on the watermarked image feature and the text feature;

determining the watermarked image feature as a second query vector;

determining a second weight set based on a similarity between the second query vector and a key in the second key-value pair; and

determining, based on the second weight set, an updated watermarked image feature.

13. The electronic device according to claim 12 , wherein generating the watermarked virtual avatar based on the image feature and the text feature comprises:

generating the watermarked virtual avatar based on the updated watermarked image feature and the text feature.

14. The electronic device according to claim 9 , wherein the actions further comprise:

sending the updated watermarked virtual avatar to a metaverse platform; and

sending an instruction for extracting the watermark in the watermarked virtual avatar to the metaverse platform.

15. The electronic device according to claim 14 , wherein the actions further comprise:

causing, in response to the watermark being extracted, the metaverse platform to verify the watermark based on information of the user; and

causing, based on a result of the verification of the watermark, the metaverse platform to determine at least one of authenticity and ownership of the virtual avatar.

16. A computer program product tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a device, cause the device to implement a method of processing a virtual avatar, the method comprising:

generating an image feature of the virtual avatar based on a plurality of image blocks of the virtual avatar and corresponding positions of the plurality of image blocks in the virtual avatar;

generating, based on a watermark to be added to the virtual avatar, a text feature associated with text of the watermark; and

generating a watermarked virtual avatar based on the image feature and the text feature, wherein the watermark is invisible to human beings and identifies an identity of a user of the virtual avatar in a metaverse;

wherein generating the watermarked virtual avatar based on the image feature and the text feature comprises:

determining a key-value pair based on the image feature and the text feature, the key-value pair comprising (i) a key that represents a semantic embedding parameter of a self-attention mechanism in a machine learning model, and (ii) a value that represents the text feature;

determining the image feature as a query vector;

determining a weight set based on a similarity between the query vector and a key in the key-value pair; and

determining, based on the weight set, a watermarked image feature;

wherein generating the watermarked virtual avatar further comprises processing the watermarked virtual avatar through an invariant layer of the machine learning model, the invariant layer comprising at least one fully-connected layer of the machine learning model and implementing a mapping function that converts a first number of image channels of an input image space of the watermarked virtual avatar into a second number of image channels of a transformed image space of an updated watermarked virtual avatar, the second number being greater than the first number.

17. The computer program product of claim 16 , wherein generating the image feature of the virtual avatar comprises:

dividing the virtual avatar into a plurality of visual marks, wherein each visual mark in the plurality of visual marks is represented by a vector of a predetermined length;

generating a corresponding position code for each of the visual marks; and

generating the image feature based on the plurality of visual marks and the corresponding position codes.

18. The computer program product of claim 16 , wherein generating the text feature associated with the text of the watermark comprises:

dividing the watermark into a plurality of text lemmas, wherein each text lemma in the plurality of text lemmas comprises a word; and

generating the text feature based on the plurality of text lemmas.

19. The computer program product of claim 16 , wherein the key-value pair is a first key-value pair, the query vector is a first query vector, the weight set is a first weight set, and the method further comprises:

determining a second key-value pair based on the watermarked image feature and the text feature;

determining the watermarked image feature as a second query vector;

determining a second weight set based on a similarity between the second query vector and a key in the second key-value pair; and

determining, based on the second weight set, an updated watermarked image feature.

20. The computer program product of claim 19 , wherein generating the watermarked virtual avatar based on the image feature and the text feature comprises:

generating the watermarked virtual avatar based on the updated watermarked image feature and the text feature.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2022
From: WANG, ZIJIA; LIU, ZHISONG; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 061742/0315 →
Priority Claims (1)
CN 202211275814.7 · Oct 18, 2022 · national
Continuity (1)
Related Publication 20240135482A1 · Apr 25, 2024
References Cited (18)
US 6408082B1 · Rhoads · 2002 [cited by examiner]
US 11281928B1 · Hoehne · 2022 [cited by examiner]
US 20080059801A1 · Cohen · 2008 [cited by examiner]
US 20110066860A1 · Dettinger · 2011 [cited by examiner]
US 20210178274A1 · St-Pierre · 2021 [cited by examiner]
US 20210365677A1 · Anzenberg · 2021 [cited by examiner]
WO WO0039954A1 · 1998 [cited by examiner]
S. Baluja, “Hiding Images in Plain Sight: Deep Steganography,” 31st Conference on Neural Information Processing Systems, Dec. 2017, 11 pages. [cited by applicant]
N. Papernot et al., “The Limitations of Deep Learning in Adversarial Settings,” IEEE European Symposium on Security & Privacy, arXiv:1511.07528v1, Nov. 24, 2015, 16 pages. [cited by applicant]
A. Dosovitskiy et al., “An Image is Worth 16×16 Words Transformers for Image Recognition at Scale,” The International Conference on Learning Representations, arXiv:2010.11929v2, Jun. 3, 2021, 22 pages. [cited by applicant]
D. A. Hudson et al., “Generative Adversarial Transformers,” Proceedings of the 38th International Conference on Machine Learning, Jul. 2021, 13 pages. [cited by applicant]
T. Zong et al., “Robust Histogram Shape Based Method for Image Watermarking,” IEEE Transactions on Circuits and Systems for Video Technology, May 2015, 14 pages, vol. 25, No. 5. [cited by applicant]
M. Zareian et al., “Robust Quantisation Index Modulation-based Approach for Image Watermarking,” IET Image Process, Jul. 2013, pp. 432-441, vol. 7, No. 5. [cited by applicant]
W. Tang et al., “Automatic Steganographic Distortion Learning using a Generative Adversarial Network,” IEEE Signal Processing Letters, Oct. 2017, pp. 1547-1551, vol. 24, No. 10. [cited by applicant]
H. Kandi et al., “Exploring the Learning Capabilities of Convolutional Neural Networks for Robust Image Watermarking,” Computers & Security, Mar. 2017, pp. 247-268, vol. 65. [cited by applicant]
D. Li et al., “A Novel CNN Based Security Guaranteed Image Watermarking Generation Scenario for Smart City Applications,” Information Sciences, Accepted Manuscript, Feb. 26, 2018, 28 pages. [cited by applicant]
S.-M. Mun et al., “Finding Robust Domain from Attacks: A Learning Framework for Blind Watermarking,” Neurocomputing, Apr. 2019, pp. 191-202, vol. 337. [cited by applicant]
J. Ouyang et al., “Color Image Watermarking Based on Quaternion Fourier Transform and Improved Uniform Log-polar Mapping,” Computers and Electrical Engineering, Aug. 2015, pp. 419-432, vol. 46. [cited by applicant]