IP Library Patent Application 17502438
Patent Application
App. No. 17/502,438

SYSTEMS, METHODS, AND MEDIA FOR ACTION RECOGNITION AND CLASSIFICATION VIA ARTIFICIAL REALITY SYSTEMS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/502,438
Abstract

In particular embodiments, a computing system may determine a user intent to perform a task in a physical environment surrounding the user. The system may send a query based on the user intent to a mapping server that stores a three-dimensional (3D) occupancy map containing spatial and semantic information of physical items in the physical environment. The mapping server may be configured to identify a subset of the physical items that are relevant to the user intent. The system may receive, from the mapping server, a response to the query comprising a portion of the 3D occupancy containing the subset of the physical items specific to the user intent. The system may capture a plurality of video frames of the physical environment. The system may process the plurality of video frames and the portion of the 3D occupancy map to provide one or more action labels associated with the task.

Claims (61)

1 . A method comprising, by a computing system:

determining a user intent to perform a task in a physical environment surrounding the user;

sending a query based on the user intent to a mapping server that stores a three-dimensional (3D) occupancy map containing spatial and semantic information of physical items in the physical environment surrounding the user, wherein the mapping server is configured to identify a subset of the physical items that are relevant to the user intent;

receiving, from the mapping server, a response to the query comprising a portion of the 3D occupancy containing the subset of the physical items specific to the user intent;

capturing a plurality of video frames of the physical environment using a camera associated with a device worn by the user; and

processing the plurality of video frames and the portion of the 3D occupancy map to provide one or more action labels associated with the task on the device worn by the user.

2 . The method of claim 1 , wherein processing the plurality of video frames and the portion of the 3D occupancy map comprises:

generating a first feature map based on processing of the plurality of video frames;

generating a second feature map based on processing of the portion of the 3D occupancy map;

processing the first feature map and the second feature map to generate an action region map, the action region map indicating a probability of action happening within each region of the portion of the 3D occupancy map;

filtering, via an attention pooling process, the second feature map associated with the portion of the 3D occupancy map based on the action region map; and

using the first feature map associated with the plurality of video frames and the filtered second feature map associated with the portion of the 3D occupancy map to generate the one or more action labels for display on the device worn by the user.

3 . The method of claim 2 , wherein the first and second feature maps are generated using a three-dimensional (3D) convolution network.

4 . The method of claim 2 , wherein processing the first feature map and the second feature map to generate the action region map comprises:

concatenating the first feature map and the second feature map using a first machine-learning model.

5 . The method of claim 4 , wherein the one or more action labels are generated using a second machine-learning model.

6 . The method of claim 2 , wherein the action region map is a heat map.

7 . The method of claim 1 , wherein the portion of the 3D occupancy is a parent-children semantic occupancy map comprising a parent voxel and a plurality of children voxels.

8 . The method of claim 7 , wherein each children voxel of the plurality of children voxels comprises a plurality of grids indicating a coarse location or feature of an item of the subset of the physical items specific to the user intent.

9 . The method of claim 1 , wherein the subset of the physical items specific to the user intent is identified, at the mapping server, using a scene graph or a knowledge graph.

10 . The method of claim 1 , wherein:

the task is an action direction task; and

the one or more action labels aid in performing the action direction task.

11 . The method of claim 1 , wherein:

the device worn by the user is an augmented-reality device; and

the one or more action labels are overlaid on a display screen of the augmented-reality device.

12 . The method of claim 1 , wherein the plurality of video frames and the portion of the 3D occupancy map are processed in parallel.

13 . The method of claim 1 , wherein the user intent is determined explicitly through a voice command of the user.

14 . The method of claim 1 , wherein the user intent is determined automatically, without explicit user input, based on one or more of a current location, time of day, or previous history of the user.

15 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

determine a user intent to perform a task in a physical environment surrounding the user;

send a query based on the user intent to a mapping server that stores a three-dimensional (3D) occupancy map containing spatial and semantic information of physical items in the physical environment surrounding the user, wherein the mapping server is configured to identify a subset of the physical items that are relevant to the user intent;

receive, from the mapping server, a response to the query comprising a portion of the 3D occupancy containing the subset of the physical items specific to the user intent;

capture a plurality of video frames of the physical environment using a camera associated with a device worn by the user; and

process the plurality of video frames and the portion of the 3D occupancy map to provide one or more action labels associated with the task.

16 . The media of claim 15 , wherein to process the plurality of video frames and the portion of the 3D occupancy map, the software is further operable when executed to:

generate a first feature map based on processing of the plurality of video frames;

generate a second feature map based on processing of the portion of the 3D occupancy map;

process the first feature map and the second feature map to generate an action region map, the action region map indicating a probability of action happening within each region of the portion of the 3D occupancy map;

filter, via an attention pooling process, the second feature map associated with the portion of the 3D occupancy map based on the action region map; and

use the first feature map associated with the plurality of video frames and the filtered second feature map associated with the portion of the 3D occupancy map to generate the one or more action labels for display on the device worn by the user.

17 . The media of claim 15 , wherein:

the task is an action direction task; and

the one or more action labels aid in performing the action direction task.

18 . A system comprising:

one or more processors; and

one or more computer-readable non-transitory storage media coupled to one or more of the processors and comprising instructions operable when executed by one or more of the processors to cause the system to:

determine a user intent to perform a task in a physical environment surrounding the user;

send a query based on the user intent to a mapping server that stores a three-dimensional (3D) occupancy map containing spatial and semantic information of physical items in the physical environment surrounding the user, wherein the mapping server is configured to identify a subset of the physical items that are relevant to the user intent;

receive, from the mapping server, a response to the query comprising a portion of the 3D occupancy containing the subset of the physical items specific to the user intent;

capture a plurality of video frames of the physical environment using a camera associated with a device worn by the user; and

process the plurality of video frames and the portion of the 3D occupancy map to provide one or more action labels associated with the task.

19 . The system of claim 18 , wherein to process the plurality of video frames and the portion of the 3D occupancy map, the one or more processors are further operable when executing the instructions to cause the system to:

generate a first feature map based on processing of the plurality of video frames;

generate a second feature map based on processing of the portion of the 3D occupancy map;

process the first feature map and the second feature map to generate an action region map, the action region map indicating a probability of action happening within each region of the portion of the 3D occupancy map;

filter, via an attention pooling process, the second feature map associated with the portion of the 3D occupancy map based on the action region map; and

use the first feature map associated with the plurality of video frames and the filtered second feature map associated with the portion of the 3D occupancy map to generate the one or more action labels for display on the device worn by the user.

20 . The system of claim 18 , wherein:

the task is an action direction task; and

the one or more action labels aid in performing the action direction task.

Assignments (2)
CHANGE OF NAME Recorded Jul 6, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060591/0848 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: LIU, MIAO; LI, CHAO; SOMASUNDARAM, KIRAN KUMAR
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 057999/0813 →