IP Library › Granted Patent US 12,749,270
Granted Patent B2
US 12,749,270 · App. 18/330,652 · Granted Sep 29, 2026

Smart interactivity for scanned objects using affordance regions

Inventors: Angela Blechschmidt (San Jose, CA); Gefen Kohavi (San Carlos, CA); Daniel Ulbricht (Sunnyvale, CA)
Assignee: Apple Inc.
G06T19/006G06T15/10G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,270
App. No.
18/330,652
Granted
Sep 29, 2026
Kind
B2
Abstract

Generating a virtual representation of an interaction includes determining a potential user interaction with a physical object in a physical environment, determining an object type associated with the physical object, and obtaining an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of the object type, a likelihood of user contact. The object-centric affordance region is mapped to a geometry of the physical object to obtain an instance-specific affordance region, is used to render the virtual representation of the interaction with the physical object.

Claims (57)

1 . A method comprising:

detecting, based on sensor data of a view of a physical environment, a first user interaction with a physical object in the physical environment;

determining an object type of the physical object;

obtaining an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of a geometry representative of the object type, a likelihood of user contact;

applying the object-centric affordance region to a virtual representation of the physical object to obtain an instance-specific affordance region; and

generating a virtual representation of a user performing the first user interaction with the virtual representation of the physical object in accordance with the instance-specific affordance region.

2 . The method of claim 1 , wherein the object-centric affordance region is further associated with an interaction type of the first user interaction.

3 . The method of claim 1 , wherein determining the instance-specific affordance region comprises:

obtaining a generic geometric representation for the object type, wherein the object-centric affordance region corresponds to the generic geometric representation; and

performing a correspondence mapping between the generic geometric representation and a geometry of the physical object.

4 . The method of claim 3 , wherein the geometry of the physical object is obtained by scanning, by a local device, the physical object.

5 . The method of claim 3 , wherein one or more faces of the generic geometric representation are associated with a likelihood of user contact for a corresponding face.

6 . The method of claim 1 , further comprising:

determining a second user interaction with the physical object in the physical environment;

obtaining a second object-centric affordance region of the object type, wherein the second object-centric affordance region indicates a likelihood of user contact for the second user interaction; and

applying the second object-centric affordance region to a geometry of the physical object.

7 . The method of claim 6 , further comprising:

in accordance with a determination that the first user interaction and the second user interaction are likely to occur concurrently, modifying the instance-specific affordance region in accordance with the second object-centric affordance region.

8 . The method of claim 1 , wherein the object-centric affordance region is generated by:

collecting training data of prior instances of the first user interaction with a plurality of objects of the object type;

for each of the plurality of objects:

generating a geometric representation of the physical object and a geometric representation of a user performing the first user interaction, and

assigning a contact value to each of a plurality of regions of the geometric representation based on the geometry representative of the object type and the geometric representation of the user performing the first user interaction; and

combining the contact values for each of the plurality of regions across each of the plurality of objects.

9 . The method of claim 1 , wherein the first user interaction is detected based on sensor data collected by a local device.

10 . A non-transitory computer readable medium comprising computer readable code executable by one or more processors to:

detect, based on sensor data of a view of a physical environment, a first user interaction with a physical object in the physical environment;

determine an object type of the physical object;

obtain an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of a geometry representative of the object type, a likelihood of user contact;

apply the object-centric affordance region to a virtual representation of the physical object to obtain an instance-specific affordance region; and

generate a virtual representation of a user performing the first user interaction with the virtual representation of the physical object in accordance with the instance-specific affordance region.

11 . The non-transitory computer readable medium of claim 10 , wherein the object-centric affordance region is further associated with an interaction type of the first user interaction.

12 . The non-transitory computer readable medium of claim 10 , wherein the computer readable code to apply the object-centric accordance region to the virtual representation of the physical object comprises computer readable code to:

obtain a generic geometric representation for the object type, wherein the object-centric affordance region corresponds to the generic geometric representation; and

perform a correspondence mapping between the generic geometric representation and a geometry of the physical object.

13 . The non-transitory computer readable medium of claim 12 , wherein the geometry of the physical object is obtained by scanning, by a local device, the physical object.

14 . The non-transitory computer readable medium of claim 12 , wherein one or more faces of the generic geometric representation are associated with a likelihood of user contact for corresponding face.

15 . A system comprising:

one or more processors; and

one or more computer readable media comprising computer readable code executable by the one or more processors to:

detect, based on sensor data of a view of a physical environment, a first user interaction with a physical object in the physical environment;

determine an object type of the physical object;

obtain an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of a geometry representative of the object type, a likelihood of user contact;

applying the object-centric affordance region to a virtual representation of the physical object to obtain an instance-specific affordance region; and

generate a virtual representation of a user performing the first user interaction with the virtual representation of the physical object in accordance with the instance-specific affordance region.

16 . The system of claim 15 , further comprising computer readable code to:

determine a second user interaction with the physical object in the physical environment;

obtain a second object-centric affordance region of the object type, wherein the second object-centric affordance region indicates a likelihood of user contact for the second user interaction; and

apply the second object-centric affordance region to a geometry of the physical object.

17 . The system of claim 16 , further comprising computer readable code to:

in accordance with a determination that the first user interaction and the second user interaction are likely to occur concurrently, modify the instance-specific affordance region in accordance with the second object-centric affordance region.

18 . The system of claim 17 , wherein the object-centric affordance region is generated by:

collecting training data of the training data of prior instances of the first user interaction with a plurality of objects of the object type;

for each of the plurality of objects:

generating a geometric representation of a corresponding object and a geometric representation of a user performing the user interaction, and

assigning a contact value to each of a plurality of regions of the geometric representation based on the geometric representation of the corresponding object and the geometric representation of the first user interaction; and

combining the contact values for each of the plurality of regions across each of the plurality of objects.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2023
From: BLECHSCHMIDT, ANGELA; KOHAVI, GEFEN; ULBRICHT, DANIEL
To: APPLE INC.
Reel/Frame 065139/0631 →
Continuity (2)
Provisional Application 63365956 · Jun 7, 2022
Related Publication 20230394773A1 · Dec 7, 2023
References Cited (9)
US 10354139B1 · Li · 2019 [cited by examiner]
US 11014246B2 · Kuwamura · 2021 [cited by examiner]
US 20210264669A1 · Shreve · 2021 [cited by examiner]
US 20220402125A1 · Moreno Noguer · 2022 [cited by examiner]
Priyanka et al., “Learning Dexterous Grasping with Object-Centric Visual Affordances”, Submitted Sep. 3, 2020, https://arxiv.org/abs/2009.01439 (Year: 2020). [cited by examiner]
Aldoma et al., “Supervised Learning of Hidden and Non-Hidden 0-order Affordances and Detection in Real Scenes”, Jun. 28, 2012, IEEE Xplore, https://doi.org/10.1109/ICRA.2012.6224931 (Year: 2012). [cited by examiner]
Deng et al., “3D AffordanceNet: A Benchmark for Visual Object Affordance Understanding”, Mar. 31, 2021, arXiv, https://doi.org/10.48550/arXiv.2103.16397 (Year: 2021). [cited by examiner]
Sansar Help, “Defining Grab Points,” Jul. 28, 2020, Retrieved from the Internet: URL: https://help.sansar.com/hc/en-us/articles/360001396786-Defining-grab-points [Retrieved on Dec. 15, 2021]. [cited by applicant]
Unity Manual Version 2020.3, “Inverse Kinematics,” Dec. 10, 2021, Retrieved from the Internet: URL: https://docs.unity3d.com/Manual/InverseKinematics.html [Retrieved on Dec. 15, 2021]. [cited by applicant]