IP Library Granted Patent US 12,646,513
Granted Patent B2
US 12,646,513 · App. 18/212,499 · Granted Jun 2, 2026

Command disambiguation based on environmental context

Inventors: Devin W. Chalmers (Oakland, CA); Brian W. Temple (Santa Clara, CA); Carlo Eduardo C. Del Mundo (Bellevue, WA); Harry J. Saddler (Berkeley, CA); Jean-Charles Bernard Marcel Bazin (Sunnyvale, CA)
Assignee: APPLE INC.
G10L15/22G06F3/013G06V10/255G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,513
App. No.
18/212,499
Granted
Jun 2, 2026
Kind
B2
Abstract

In one implementation, a method of changing a state of an object is performed at a device including an image sensor, one or more processors, and non-transitory memory. The method includes receiving a vocal command. The method includes obtaining, using the image sensor, an image of a physical environment. The method includes detecting, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object. The method includes generating, based on the vocal command and detection of the object, an instruction including the object identifier of the object. The method includes effectuating the instruction to change a state of the object.

Claims (73)

1 . A method comprising:

at an electronic device including an image sensor, one or more processors, and non-transitory memory:

receiving a vocal command;

obtaining, using the image sensor, an image of a physical environment;

detecting, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object;

determining, based on the vocal command, a plurality of potential instructions;

selecting, based on the detection of the object, one of the plurality of potential instructions as an intended instruction; and

generating a data packet including the object identifier of the object to effectuate the intended instruction to change a state of the object.

2 . The method of claim 1 , wherein the object identifier of the object is further stored in association with an object type of the object and selecting the intended instruction is further based on the object type of the object.

3 . The method of claim 1 , wherein the object identifier of the object is further stored in association with a name of the object and selecting the intended instruction is further based on the name of the object.

4 . The method of claim 1 , wherein selecting the intended instruction is further based on determining that a gaze of a user is directed at a location of the detection of the object.

5 . The method of claim 1 , wherein selecting the intended instruction is further based on the state of the object.

6 . The method of claim 1 , further comprising:

receiving a second vocal command;

obtaining, using the image sensor, an image of a second physical environment;

detecting, in the image of the second physical environment, the object;

determining, based on the second vocal command, a second plurality of potential instructions;

selecting, based on the detection of the object, one of the second plurality of potential instructions as a second intended instruction; and

generating a second data packet including the object identifier of the object to effectuate the second intended instruction to change the state of the object.

7 . The method of claim 1 , further comprising:

obtaining a request to enroll the object;

obtaining, using the image sensor, one or more images of the object;

determining, based on the one or more images of the object, the visual model of the object; and

storing, in the non-transitory memory, the visual model of the object in association with the object identifier of the object.

8 . A device comprising:

an image sensor;

a non-transitory memory; and

one or more processors to:

receive a vocal command;

obtain, using the image sensor, an image of a physical environment;

detect, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object;

determine, based on the vocal command, a plurality of potential instructions;

select, based on the detection of the object, one of the plurality of potential instructions as an intended instructions; and

generate a data packet including the object identifier of the object to effectuate the intended instruction to change a state of the object.

9 . The device of claim 8 , wherein the object identifier of the object is further stored in association with an object type of the object and the one or more processors are to select the intended instruction based on the object type of the object.

10 . The device of claim 8 , wherein the object identifier of the object is further stored in association with a name of the object and the one or more processors are to select the intended instruction based on the name of the object.

11 . The device of claim 8 , wherein the one or more processors are to select the intended instruction further based on determining that a gaze of a user is directed at a location of the detection of the object.

12 . The device of claim 8 , wherein the one or more processors are to select the intended instruction based on the state of the object.

13 . The device of claim 8 , wherein the one or more processors are further to:

receive a second vocal command;

obtain, using the image sensor, an image of a second physical environment;

detect, in the image of the second physical environment, the object;

determine, based on the second vocal command, a second plurality of potential instructions;

select, based on the detection of the object, one of the second plurality of potential instructions as a second intended instructions; and

generate a second data packet including the object identifier of the object to effectuate the second intended instruction to change the state of the object.

14 . The device of claim 8 , wherein the one or more processors are further to:

obtain a request to enroll the object;

obtain, using the image sensor, one or more images of the object;

determine, based on the one or more images of the object, the visual model of the object; and

store, in the non-transitory memory, the visual model of the object in association with the object identifier of the object.

15 . A non-transitory memory storing one or more programs, which, when executed by one or more processors of a device including an image sensor, cause the device to:

receive a vocal command;

obtain, using the image sensor, an image of a physical environment;

detect, in the image of the physical environment, an object based on a visual model of the object stored in the non-transitory memory in association with an object identifier of the object;

determine, based on the vocal command, a plurality of potential instructions;

select, based on the detection of the object, one of the plurality of potential instructions as an intended instructions; and

generate a data packet including the object identifier of the object to effectuate the intended instruction to change a state of the object.

16 . The non-transitory memory of claim 15 , wherein the object identifier of the object is further stored in association with an object type of the object and the programs, when executed, cause the device to select the intended instruction based on the object type of the object.

17 . The non-transitory memory of claim 15 , wherein the object identifier of the object is further stored in association with a name of the object and the programs, when executed, cause the device to select the intended instruction based on the name of the object.

18 . The non-transitory memory of claim 15 , wherein the programs, when executed, cause the device to select the intended instruction based on determining that a gaze of a user is directed at a location of the detection of the object.

19 . The non-transitory memory of claim 15 , wherein the programs, when executed, cause the device to select the intended instruction based on the state of the object.

20 . The non-transitory memory of claim 15 , wherein the programs, when executed, further cause the device to:

receive a second vocal command;

obtain, using the image sensor, an image of a second physical environment;

detect, in the image of the second physical environment, the object;

determine, based on the second vocal command, a second plurality of potential instructions;

select, based on the detection of the object, one of the second plurality of potential instructions as a second intended instruction; and

generate a second data packet including the object identifier of the object to effectuate the second intended instruction to change the state of the object.

21 . The non-transitory memory of claim 15 , wherein the programs, when executed, further cause the device to:

obtain a request to enroll the object;

obtain, using the image sensor, one or more images of the object;

determine, based on the one or more images of the object, the visual model of the object; and

store, in the non-transitory memory, the visual model of the object in association with the object identifier of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2023
From: CHALMERS, DEVIN W.; DEL MUNDO, CARLO EDWARDO C.; TEMPLE, BRIAN W.; BAZIN, JEAN-CHARLES BERNARD MARCEL; SADDLER, HARRY J.
To: APPLE INC.
Reel/Frame 064020/0896 →
Continuity (2)
Provisional Application 63356626 · Jun 29, 2022
Related Publication 20240005921A1 · Jan 4, 2024
References Cited (18)
US 9116962B1 · Pance · 2015 [cited by applicant]
US 9911240B2 · Bedikian et al. · 2018 [cited by applicant]
US 10475446B2 · Gruber et al. · 2019 [cited by applicant]
US 10565256B2 · Badr et al. · 2020 [cited by applicant]
US 10963735B2 · Rhoads · 2021 [cited by applicant]
US 11087424B1 · Sarkar et al. · 2021 [cited by applicant]
US 11853651B2 · Francisco · 2023 [cited by examiner]
US 20130238326A1 · Kim · 2013 [cited by examiner]
US 20180232608A1 · Pradeep · 2018 [cited by examiner]
US 20190333233A1 · Hu · 2019 [cited by examiner]
US 20200193206A1 · Turkelson · 2020 [cited by examiner]
US 20200193976A1 · Cartwright · 2020 [cited by examiner]
US 20200204391A1 · Glaser et al. · 2020 [cited by applicant]
US 20210342047A1 · Badr et al. · 2021 [cited by applicant]
US 20210366472A1 · Lee et al. · 2021 [cited by applicant]
US 20220012547A1 · Zhu et al. · 2022 [cited by applicant]
US 20230089049A1 · Drummond · 2023 [cited by examiner]
US 20240411364A1 · Miller · 2024 [cited by examiner]