IP Library Granted Patent US 12,242,664
Granted Patent B2
US 12,242,664 · App. 18/204,892 · Granted Mar 4, 2025

Multimodal inputs for computer-generated reality

Inventors: Ranjit Desai (Cupertino, CA); Maneli Noorkami (Menlo Park, CA)
Assignee: Apple Inc.
G06F3/011G06F9/5011G06F16/907G06T19/006G06V10/25G06V20/20G06V40/174G06V40/176G06V40/20H04N21/4334
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,664
App. No.
18/204,892
Granted
Mar 4, 2025
Kind
B2
Abstract

Implementations of the subject technology provide determining an operating mode of an electronic device based at least in part on whether the electronic device is communicatively coupled to an associated base device. Based on the determined operating mode, the subject technology identifies a set of input modalities for initiating a recording of content within a field of view of the electronic device. The subject technology monitors sensor information generated by at least one sensor included in, or communicatively coupled to, the electronic device. Further, the subject technology initiates the recording of content within the field of view of the electronic device when the monitored sensor information indicates that at least one of the identified set of input modalities has been triggered.

Claims (52)

1. A method comprising:

determining a quality of service metric associated with operation of an electronic device;

based on the determined quality of service metric, selecting a first set of user input modalities or a second set of user input modalities for initiating a recording of content within a field of view of the electronic device, the second set of user input modalities differing from the first set of user input modalities;

monitoring sensor information generated by at least one sensor included in, or communicatively coupled to, the electronic device; and

initiating the recording of content within the field of view of the electronic device when the monitored sensor information indicates that a first user input corresponding to at least one of the first selected set of user input modalities has been received from a user when the first selected set of user input modalities is selected, or that a second user input corresponding to at least one of the second selected set of user input modalities has been received from the user when the second set of user input modalities is selected.

2. The method of claim 1 , further comprising:

detecting that the quality of service metric has changed; and

responsive to detecting the change, updating the set of user input modalities for initiating the recording of content based on the changed quality of service metric.

3. The method of claim 1 , wherein the quality of service metric is determined based at least in part on available computing resources, or available power in the electronic device, the available power including an amount of battery power.

4. The method of claim 1 , further comprising:

deactivating a particular input modality corresponding to at least one sensor in the electronic device, wherein the particular input modality comprises a facial expression, gaze direction, eye position, hand gesture, hardware input, speech, or identification of an object or person in a scene.

5. The method of claim 1 , further comprising:

deactivating a particular sensor in the electronic device when the quality of service metric falls below a threshold value, wherein the particular sensor comprises a camera, an inertial measurement unit, a microphone, or a touch sensor.

6. The method of claim 1 , further comprising:

determining a region of interest in the field of view of the electronic device.

7. The method of claim 6 , wherein the region of interest is determined based on a gesture or an indicator corresponding to the region of interest.

8. The method of claim 1 , further comprising:

generating an annotation corresponding to the recording of content; and

adding the annotation as metadata to the recording of content.

9. A system comprising;

a processor,

a memory device containing instructions, which when executed by the processor cause the processor to perform operations comprising:

determining a quality of service metric associated with operation of an electronic device;

based on the determined quality of service metric, selecting a first set of input modalities or a second set of input modalities for initiating a recording of content within a field of view of the electronic device, the second set of input modalities differing from the first set of input modalities;

monitoring sensor information generated by at least one sensor included in, or communicatively coupled to, the electronic device; and

initiating the recording of content within the field of view of the electronic device when the monitored sensor information indicates that a first user input corresponding to at least one of the first selected set of input modalities has been received from a user when the first selected set of input modalities is selected, or that a second user input corresponding to at least one of the second selected set of input modalities has been received from the user when the second set of input modalities is selected.

10. The system of claim 9 , wherein the memory device contains further instructions, which when executed by the processor further cause the processor to perform further operations further comprising:

detecting that the quality of service metric has changed; and

responsive to detecting the change, updating the set of input modalities for initiating the recording of content based on the changed quality of service metric.

11. The system of claim 9 , wherein the quality of service metric is determined based at least in part on available computing resources, or available power in the electronic device, the available power including an amount of battery power.

12. The system of claim 9 , wherein the memory device contains further instructions, which when executed by the processor further cause the processor to perform further operations further comprising:

deactivating a particular input modality corresponding to at least one sensor in the electronic device, wherein the particular input modality comprises a facial expression, gaze direction, eye position, hand gesture, hardware input, speech, or identification of an object or person in a scene.

13. The system of claim 9 , wherein the memory device contains further instructions, which when executed by the processor further cause the processor to perform further operations further comprising:

deactivating a particular sensor in the electronic device when the quality of service metric falls below a threshold value, wherein the particular sensor comprises a camera, an inertial measurement unit, a microphone, or a touch sensor.

14. The system of claim 9 , wherein the memory device contains further instructions, which when executed by the processor further cause the processor to perform further operations further comprising:

determining a region of interest in the field of view of the electronic device, wherein the region of interest is determined based on a gesture or an indicator corresponding to the region of interest.

15. The system of claim 9 , wherein the memory device contains further instructions, which when executed by the processor further cause the processor to perform further operations further comprising:

generating an annotation corresponding to the recording of content; and

adding the annotation as metadata to the recording of content.

16. A non-transitory machine-readable medium comprising instructions, which when executed by a computing device, cause the computing device to perform operations comprising:

determining a quality of service metric associated with operation of an electronic device;

based on the determined quality of service metric, selecting a first set of input modalities or a second set of input modalities for initiating a recording of content within a field of view of the electronic device, the second set of input modalities differing from the first set of input modalities;

monitoring sensor information generated by at least one sensor included in, or communicatively coupled to, the electronic device; and

initiating the recording of content within the field of view of the electronic device when the monitored sensor information indicates that a first user input corresponding to at least one of the first selected set of input modalities has been received from a user when the first selected set of input modalities is selected, or that a second user input corresponding to at least one of the second selected set of input modalities has been received from the user when the second set of input modalities is selected.

17. The non-transitory machine-readable medium of claim 16 , wherein the operations further comprise:

detecting that the quality of service metric has changed; and

responsive to detecting the change, updating the set of input modalities for initiating the recording of content based on the changed quality of service metric.

18. The non-transitory machine-readable medium of claim 16 , wherein the quality of service metric is determined based at least in part on available computing resources, or available power in the electronic device, the available power including an amount of battery power.

19. The non-transitory machine-readable medium of claim 16 , further comprising:

deactivating a particular sensor in the electronic device when the quality of service metric falls below a threshold value, wherein the particular sensor comprises a camera, an inertial measurement unit, a microphone, or a touch sensor.

20. The non-transitory machine-readable medium of claim 16 , further comprising:

determining a region of interest in the field of view of the electronic device.

Continuity (3)
Continuation 17016190 · Sep 9, 2020
Provisional Application 62897909 · Sep 9, 2019
Related Publication 20230315196A1 · Oct 5, 2023
References Cited (18)
US 10812711B2 · Sapienza · 2020 [cited by applicant]
US 10908694B2 · Potts · 2021 [cited by applicant]
US 20090218957A1 · Kraft et al. · 2009 [cited by applicant]
US 20130177296A1 · Geisner et al. · 2013 [cited by applicant]
US 20140267021A1 · Lee et al. · 2014 [cited by applicant]
US 20150268728A1 · Makela · 2015 [cited by applicant]
US 20150346932A1 · Nuthulapati · 2015 [cited by applicant]
US 20170337476A1 · Gordon et al. · 2017 [cited by applicant]
US 20180276841A1 · Krishnaswamy · 2018 [cited by examiner]
US 20180336332A1 · Singh · 2018 [cited by applicant]
US 20190141252A1 · Pallamsetty · 2019 [cited by examiner]
CN 103389798A · 2013 [cited by applicant]
CN 106815264A · 2017 [cited by applicant]
CN 109154861A · 2019 [cited by applicant]
International Search Report and Written Opinion from PCT/US2020/050002, dated Dec. 2, 2020, 17 pages. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 202080057569.5, dated Apr. 8, 2024, 17 pages including English language translation. [cited by applicant]
European Office Action from European Patent Application No. 20776028.1, dated Oct. 4, 2023, 6 pages. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 202080057569.5, dated Dec. 3, 2024, 10 pages including English language translation. [cited by applicant]