Smart interactivity for scanned objects using affordance regions
Generating a virtual representation of an interaction includes determining a potential user interaction with a physical object in a physical environment, determining an object type associated with the physical object, and obtaining an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of the object type, a likelihood of user contact. The object-centric affordance region is mapped to a geometry of the physical object to obtain an instance-specific affordance region, is used to render the virtual representation of the interaction with the physical object.
1 . A method comprising:
detecting, based on sensor data of a view of a physical environment, a first user interaction with a physical object in the physical environment;
determining an object type of the physical object;
obtaining an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of a geometry representative of the object type, a likelihood of user contact;
applying the object-centric affordance region to a virtual representation of the physical object to obtain an instance-specific affordance region; and
generating a virtual representation of a user performing the first user interaction with the virtual representation of the physical object in accordance with the instance-specific affordance region.
2 . The method of claim 1 , wherein the object-centric affordance region is further associated with an interaction type of the first user interaction.
3 . The method of claim 1 , wherein determining the instance-specific affordance region comprises:
obtaining a generic geometric representation for the object type, wherein the object-centric affordance region corresponds to the generic geometric representation; and
performing a correspondence mapping between the generic geometric representation and a geometry of the physical object.
4 . The method of claim 3 , wherein the geometry of the physical object is obtained by scanning, by a local device, the physical object.
5 . The method of claim 3 , wherein one or more faces of the generic geometric representation are associated with a likelihood of user contact for a corresponding face.
6 . The method of claim 1 , further comprising:
determining a second user interaction with the physical object in the physical environment;
obtaining a second object-centric affordance region of the object type, wherein the second object-centric affordance region indicates a likelihood of user contact for the second user interaction; and
applying the second object-centric affordance region to a geometry of the physical object.
7 . The method of claim 6 , further comprising:
in accordance with a determination that the first user interaction and the second user interaction are likely to occur concurrently, modifying the instance-specific affordance region in accordance with the second object-centric affordance region.
8 . The method of claim 1 , wherein the object-centric affordance region is generated by:
collecting training data of prior instances of the first user interaction with a plurality of objects of the object type;
for each of the plurality of objects:
generating a geometric representation of the physical object and a geometric representation of a user performing the first user interaction, and
assigning a contact value to each of a plurality of regions of the geometric representation based on the geometry representative of the object type and the geometric representation of the user performing the first user interaction; and
combining the contact values for each of the plurality of regions across each of the plurality of objects.
9 . The method of claim 1 , wherein the first user interaction is detected based on sensor data collected by a local device.
10 . A non-transitory computer readable medium comprising computer readable code executable by one or more processors to:
detect, based on sensor data of a view of a physical environment, a first user interaction with a physical object in the physical environment;
determine an object type of the physical object;
obtain an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of a geometry representative of the object type, a likelihood of user contact;
apply the object-centric affordance region to a virtual representation of the physical object to obtain an instance-specific affordance region; and
generate a virtual representation of a user performing the first user interaction with the virtual representation of the physical object in accordance with the instance-specific affordance region.
11 . The non-transitory computer readable medium of claim 10 , wherein the object-centric affordance region is further associated with an interaction type of the first user interaction.
12 . The non-transitory computer readable medium of claim 10 , wherein the computer readable code to apply the object-centric accordance region to the virtual representation of the physical object comprises computer readable code to:
obtain a generic geometric representation for the object type, wherein the object-centric affordance region corresponds to the generic geometric representation; and
perform a correspondence mapping between the generic geometric representation and a geometry of the physical object.
13 . The non-transitory computer readable medium of claim 12 , wherein the geometry of the physical object is obtained by scanning, by a local device, the physical object.
14 . The non-transitory computer readable medium of claim 12 , wherein one or more faces of the generic geometric representation are associated with a likelihood of user contact for corresponding face.
15 . A system comprising:
one or more processors; and
one or more computer readable media comprising computer readable code executable by the one or more processors to:
detect, based on sensor data of a view of a physical environment, a first user interaction with a physical object in the physical environment;
determine an object type of the physical object;
obtain an object-centric affordance region for the object type, wherein the object-centric affordance region indicates, for each of one or more regions of a geometry representative of the object type, a likelihood of user contact;
applying the object-centric affordance region to a virtual representation of the physical object to obtain an instance-specific affordance region; and
generate a virtual representation of a user performing the first user interaction with the virtual representation of the physical object in accordance with the instance-specific affordance region.
16 . The system of claim 15 , further comprising computer readable code to:
determine a second user interaction with the physical object in the physical environment;
obtain a second object-centric affordance region of the object type, wherein the second object-centric affordance region indicates a likelihood of user contact for the second user interaction; and
apply the second object-centric affordance region to a geometry of the physical object.
17 . The system of claim 16 , further comprising computer readable code to:
in accordance with a determination that the first user interaction and the second user interaction are likely to occur concurrently, modify the instance-specific affordance region in accordance with the second object-centric affordance region.
18 . The system of claim 17 , wherein the object-centric affordance region is generated by:
collecting training data of the training data of prior instances of the first user interaction with a plurality of objects of the object type;
for each of the plurality of objects:
generating a geometric representation of a corresponding object and a geometric representation of a user performing the user interaction, and
assigning a contact value to each of a plurality of regions of the geometric representation based on the geometric representation of the corresponding object and the geometric representation of the first user interaction; and
combining the contact values for each of the plurality of regions across each of the plurality of objects.