IP Library Granted Patent US 12,267,623
Granted Patent B2
US 12,267,623 · App. 17/751,405 · Granted Apr 1, 2025

Camera-less representation of users during communication sessions

Inventors: Justin G. Binder (Oakland, CA); Abhimanyu Yadav (San Francisco, CA); Ahmed S. Hussen Abdelaziz (Cupertino, CA); Abhishek Walia (Santa Clara, CA); Anushree Prasanna Kumar (San Jose, CA)
Assignee: Apple Inc.
H04N7/157G06V40/176G10L15/25H04L12/1822H04N7/152
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,267,623
App. No.
17/751,405
Granted
Apr 1, 2025
Kind
B2
Abstract

An example process includes receiving, from a user, an input corresponding to a request to render, without using a camera, and during a communication session with an external electronic device, an avatar associated with the user; and in accordance with receiving the input: in accordance with a determination that the electronic device is coupled to an external accessory device: during the communication session with the external electronic device, and while a camera corresponding to the communication session is disabled: receiving, from the external accessory device, a first data stream detected by a first type of sensor of the external accessory device; determining, based on the first data stream, a first set of data representing a first type of visual feature of the avatar; and rendering the avatar using the first set of data.

Claims (121)

1. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to:

receive, from a user, an input corresponding to a request to render, without using camera data, and during a communication session with an external electronic device, an avatar associated with the user; and

in accordance with receiving the input:

in accordance with a determination that the electronic device is coupled to an external accessory device:

during the communication session with the external electronic device, and while a camera corresponding to the communication session is disabled:

receive, from the external accessory device separate from the electronic device, a first data stream detected by a first type of sensor of the external accessory device, wherein the first type of sensor includes a vibration sensor for detecting bone conduction data;

determine, based on the first data stream detected by the first type of sensor, a first set of data representing a first type of visual feature of the avatar; and

render the avatar using the first set of data.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the external accessory device does not include a camera.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the external accessory device includes a headset.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the communication session corresponds to at least one of:

a textual communication session;

an audio communication session;

a video communication session; and

a virtual or mixed reality communication session.

5. The non-transitory computer-readable storage medium of claim 1 , wherein the camera corresponding to the communication session is configured to transmit image data to one or more participants of the communication session.

6. The non-transitory computer-readable storage medium of claim 1 , wherein:

the input corresponds to a selection of an affordance displayed in a user interface of a second application configured to provide the communication session.

7. The non-transitory computer-readable storage medium of claim 6 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in response to receiving the input corresponding to the selection of the affordance, disable a camera accessible by the second application.

8. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

in accordance with receiving the input:

in accordance with a determination that the electronic device is coupled to the external accessory device:

during the communication session with the external electronic device, and while the camera corresponding to the communication session is disabled:

receive, from the external accessory device, a second data stream detected by a second type of sensor of the external accessory device; and

determine, based on the second data stream, a second set of data representing a second type of visual feature of the avatar; and

wherein rendering the avatar using the first set of data includes rendering the avatar using the second set of data.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the second type of sensor includes an audio sensor and the second type of visual feature includes mouth movement of the avatar, the mouth movement corresponding to user speech.

10. The non-transitory computer-readable storage medium of claim 8 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

determine, based on the second data stream, a third set of data representing facial movement corresponding to non-speech sound, wherein rendering the avatar using the first set of data includes rendering the avatar using the third set of data.

11. The non-transitory computer-readable storage medium of claim 8 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

determine, based on the second data stream, a fourth set of data representing an emotional state of the user, wherein rendering the avatar using the first set of data includes rendering the avatar using the fourth set of data.

12. The non-transitory computer-readable storage medium of claim 8 , wherein rendering the avatar using the first set of data includes:

synchronizing displayed mouth movement of the avatar with user speech included in the second data stream.

13. The non-transitory computer-readable storage medium of claim 1 , wherein the first type of visual feature of the avatar includes facial movement of the avatar, the facial movement corresponding to a predetermined non-speech sound.

14. The non-transitory computer-readable storage medium of claim 1 , wherein the first type of visual feature of the avatar includes an emotional state of the avatar, the emotional state corresponding to the a visual feature of the user.

15. The non-transitory computer-readable storage medium of claim 1 , wherein the one or more programs further comprise instructions, which when executed by the one or more processors, cause the electronic device to:

determine, based on the first data stream detected by the first type of sensor, whether the user is speaking; and

in accordance with a determination that the user is not speaking, forgo rendering the avatar using the first set of data.

16. An electronic device, comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving, from a user, an input corresponding to a request to render, without using camera data, and during a communication session with an external electronic device, an avatar associated with the user; and

in accordance with receiving the input:

in accordance with a determination that the electronic device is coupled to an external accessory device:

during the communication session with the external electronic device, and while a camera corresponding to the communication session is disabled:

receiving, from the external accessory device separate from the electronic device, a first data stream detected by a first type of sensor of the external accessory device, wherein the first type of sensor includes a vibration sensor for detecting bone conduction data;

determining, based on the first data stream detected by the first type of sensor, a first set of data representing a first type of visual feature of the avatar; and

rendering the avatar using the first set of data.

17. The electronic device of claim 16 , wherein the external accessory device does not include a camera.

18. The electronic device of claim 16 , wherein the external accessory device includes a headset.

19. The electronic device of claim 16 , wherein the communication session corresponds to at least one of:

a textual communication session;

an audio communication session;

a video communication session; and

a virtual or mixed reality communication session.

20. The electronic device of claim 16 , wherein the camera corresponding to the communication session is configured to transmit image data to one or more participants of the communication session.

21. The electronic device of claim 16 , wherein:

the input corresponds to a selection of an affordance displayed in a user interface of a second application configured to provide the communication session.

22. The electronic device of claim 21 , wherein the one or more programs further include instructions for:

in response to receiving the input corresponding to the selection of the affordance, disabling a camera accessible by the second application.

23. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

in accordance with receiving the input:

in accordance with a determination that the electronic device is coupled to the external accessory device:

during the communication session with the external electronic device, and while the camera corresponding to the communication session is disabled:

receiving, from the external accessory device, a second data stream detected by a second type of sensor of the external accessory device; and

determining, based on the second data stream, a second set of data representing a second type of visual feature of the avatar; and

wherein rendering the avatar using the first set of data includes rendering the avatar using the second set of data.

24. The electronic device of claim 23 , wherein the second type of sensor includes an audio sensor and the second type of visual feature includes mouth movement of the avatar, the mouth movement corresponding to user speech.

25. The non electronic device of claim 23 , wherein the one or more programs further include instructions for:

determining, based on the second data stream, a third set of data representing facial movement corresponding to non-speech sound, wherein rendering the avatar using the first set of data includes rendering the avatar using the third set of data.

26. The electronic device of claim 23 , wherein the one or more programs further include instructions for:

determining, based on the second data stream, a fourth set of data representing an emotional state of the user, wherein rendering the avatar using the first set of data includes rendering the avatar using the fourth set of data.

27. The electronic device of claim 23 , wherein rendering the avatar using the first set of data includes:

synchronizing displayed mouth movement of the avatar with user speech included in the second data stream.

28. The electronic device of claim 16 , wherein the first type of visual feature of the avatar includes facial movement of the avatar, the facial movement corresponding to a predetermined non-speech sound.

29. The electronic device of claim 16 , wherein the first type of visual feature of the avatar includes an emotional state of the avatar, the emotional state corresponding to the a visual feature of the user.

30. The electronic device of claim 16 , wherein the one or more programs further include instructions for:

determining, based on the first data stream detected by the first type of sensor, whether the user is speaking; and

in accordance with a determination that the user is not speaking, forging rendering the avatar using the first set of data.

31. A method, comprising:

at an electronic device with one or more processors and memory:

receiving, from a user, an input corresponding to a request to render, without using camera data, and during a communication session with an external electronic device, an avatar associated with the user; and

in accordance with receiving the input:

in accordance with a determination that the electronic device is coupled to an external accessory device:

during the communication session with the external electronic device, and while a camera corresponding to the communication session is disabled:

 receiving, from the external accessory device separate from the electronic device, a first data stream detected by a first type of sensor of the external accessory device, wherein the first type of sensor includes a vibration sensor for detecting bone conduction data;

 determining, based on the first data stream detected by the first type of sensor, a first set of data representing a first type of visual feature of the avatar; and

 rendering the avatar using the first set of data.

32. The method of claim 31 , wherein the external accessory device does not include a camera.

33. The method of claim 31 , wherein the external accessory device includes a headset.

34. The method of claim 31 , wherein the communication session corresponds to at least one of:

a textual communication session;

an audio communication session;

a video communication session; and

a virtual or mixed reality communication session.

35. The method of claim 31 , wherein the camera corresponding to the communication session is configured to transmit image data to one or more participants of the communication session.

36. The method of claim 31 , wherein:

the input corresponds to a selection of an affordance displayed in a user interface of a second application configured to provide the communication session.

37. The method of claim 36 , further comprising:

in response to receiving the input corresponding to the selection of the affordance, disabling a camera accessible by the second application.

38. The method of claim 31 , further comprising:

in accordance with receiving the input:

in accordance with a determination that the electronic device is coupled to the external accessory device:

during the communication session with the external electronic device, and while the camera corresponding to the communication session is disabled:

receiving, from the external accessory device, a second data stream detected by a second type of sensor of the external accessory device; and

determining, based on the second data stream, a second set of data representing a second type of visual feature of the avatar; and

wherein rendering the avatar using the first set of data includes rendering the avatar using the second set of data.

39. The method of claim 38 , wherein the second type of sensor includes an audio sensor and the second type of visual feature includes mouth movement of the avatar, the mouth movement corresponding to user speech.

40. The method of claim 38 , further comprising:

determining, based on the second data stream, a third set of data representing facial movement corresponding to non-speech sound, wherein rendering the avatar using the first set of data includes rendering the avatar using the third set of data.

41. The method of claim 38 , further comprising:

determining, based on the second data stream, a fourth set of data representing an emotional state of the user, wherein rendering the avatar using the first set of data includes rendering the avatar using the fourth set of data.

42. The method of claim 38 , wherein rendering the avatar using the first set of data includes:

synchronizing displayed mouth movement of the avatar with user speech included in the second data stream.

43. The method of claim 31 , wherein the first type of visual feature of the avatar includes facial movement of the avatar, the facial movement corresponding to a predetermined non-speech sound.

44. The method of claim 31 , wherein the first type of visual feature of the avatar includes an emotional state of the avatar, the emotional state corresponding to the a visual feature of the user.

45. The method of claim 31 , further comprising:

determining, based on the first data stream detected by the first type of sensor, whether the user is speaking; and

in accordance with a determination that the user is not speaking, forgoing rendering the avatar using the first set of data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2022
From: HUSSEN ABDELAZIZ, AHMED S.; BINDER, JUSTIN G.; KUMAR, ANUSHREE PRASANNA; YADAV, ABHIMANYU; WALIA, ABHISHEK
To: APPLE INC.
Reel/Frame 060433/0110 →
Continuity (2)
Provisional Application 63308864 · Feb 10, 2022
Related Publication 20230254448A1 · Aug 10, 2023
References Cited (34)
US 8156060B2 · Borzestowski et al. · 2012 [cited by applicant]
US 9665567B2 · Li et al. · 2017 [cited by applicant]
US 10360716B1 · Van Der Meulen et al. · 2019 [cited by applicant]
US 10521946B1 · Roche et al. · 2019 [cited by applicant]
US 10586369B1 · Roche et al. · 2020 [cited by applicant]
US 10904488B1 · Weisz · 2021 [cited by examiner]
US 20090055187A1 · Leventhal et al. · 2009 [cited by applicant]
US 20100082345A1 · Wang et al. · 2010 [cited by applicant]
US 20100125785A1 · Moore et al. · 2010 [cited by applicant]
US 20100125811A1 · Moore et al. · 2010 [cited by applicant]
US 20110064388A1 · Brown et al. · 2011 [cited by applicant]
US 20150100537A1 · Grieves et al. · 2015 [cited by applicant]
US 20150312523A1 · Li · 2015 [cited by examiner]
US 20170185581A1 · Bojja et al. · 2017 [cited by applicant]
US 20180107945A1 · Gao et al. · 2018 [cited by applicant]
US 20180336184A1 · Bellegarda et al. · 2018 [cited by applicant]
US 20190172243A1 · Mishra et al. · 2019 [cited by applicant]
US 20190213774A1 · Jiao et al. · 2019 [cited by applicant]
US 20190340419A1 · Milman · 2019 [cited by examiner]
US 20200090393A1 · Shin et al. · 2020 [cited by applicant]
US 20200135226A1 · Mittal et al. · 2020 [cited by applicant]
US 20200357406A1 · York et al. · 2020 [cited by applicant]
US 20210090314A1 · Hussen et al. · 2021 [cited by applicant]
US 20210248804A1 · Hussen Abdelaziz et al. · 2021 [cited by applicant]
US 20210249009A1 · Manjunath et al. · 2021 [cited by applicant]
US 20210281965A1 · Malik et al. · 2021 [cited by applicant]
US 20220068278A1 · York et al. · 2022 [cited by applicant]
US 20220277505A1 · Baszucki · 2022 [cited by examiner]
US 20230199147A1 · Desserrey · 2023 [cited by examiner]
CN 110598671A · 2019 [cited by applicant]
CN 111316203A · 2020 [cited by applicant]
WO 2020010530A1 · 2020 [cited by applicant]
Fitzpatrick, Aidan, “Introducing Camo 1.5: AR modes”, Available Online at: “https://reincubate.com/blog/camo-ar-modes-release/”, Oct. 28, 2021, 8 pages. [cited by applicant]
Malcangi Mario, “Text-driven avatars based on artificial neural networks and fuzzy logic”, International Journal of Computers, vol. 4, No. 2, Dec. 31, 2010, pp. 61-69. [cited by applicant]