SYSTEMS AND METHODS FOR AUDIO-BASED AUGMENTED REALITY
Systems, methods, and non-transitory computer readable media are configured to receive a user request to identify at least one object of an environment in which a computing device is situated. A classification for the at least one object can be received. Subsequently, an audio tag based on the classification for the at least one object can be placed in a representation of the environment. The audio tag can be associated with a sound perceived by a user to be emanating from the least one object.
1 . A computer-implemented method comprising:
receiving, by a computing device, a user request to identify at least one object of an environment in which the computing device is situated;
receiving, by the computing device, a classification for the at least one object;
placing, by the computing device, an audio tag based on the classification for the at least one object in a representation of the environment, wherein the audio tag is associated with a sound perceived by a user to be emanating from the at least one object; and
generating, by the computing device, in response to the user request, the sound perceived by the user to be emanating from the at least one object.
2 . The computer-implemented method of claim 1 , wherein generating the sound perceived by the user is based at least in part on a location and an orientation of the computing device.
3 . The computer-implemented method of claim 1 , further comprising:
modeling, by the computing device, sound wave propagation within the representation of the environment in which the computing device is situated; and
generating, by the computing device, audio spatialization data based on the modeling of the sound wave propagation.
4 . The computer-implemented method of claim 3 , wherein generating the sound perceived by the user is based at least in part on the audio spatialization data.
5 . The computer-implemented method of claim 1 , further comprising:
receiving, by the computing device, one or more captured images of the environment in which the computing device is situated;
receiving, by the computing device, sensor data from one or more sensors of the computing device; and
generating, by the computing device, the representation of the environment in which the computing device is situated based on the one or more captured images and the sensor data.
6 . The computer-implemented method of claim 5 , wherein the one or more sensors include one or more of accelerometers, gyroscopes, or magnetometers.
7 . The computer-implemented method of claim 5 , wherein the representation of the environment in which the computing device is situated is a sparse three-dimensional map representation.
8 . The computer-implemented method of claim 1 , further comprising:
tracking, by the computing device, a location of the computing device within the environment in which the computing device is situated.
9 . The computer-implemented method of claim 1 , wherein the receiving a classification for the at least one object further comprises:
accessing, by the computing device, one or more machine learning models of a cascade of machine learning models.
10 . The computer-implemented method of claim 1 , further comprising:
receiving, by the computing device, a question regarding an object of the environment; and
generating, by the computing device, an answer to the question based on one or more of a cascade of machine learning models or a knowledge base graph.
11 . A system comprising:
at least one processor; and
a memory storing instructions that, when executed by the at least one processor, cause the system to perform:
receiving a user request to identify at least one object of an environment in which the system is situated;
receiving a classification for the at least one object;
placing an audio tag based on the classification for the at least one object in a representation of the environment, wherein the audio tag is associated with a sound perceived by a user to be emanating from the at least one object; and
generating, in response to the user request, the sound perceived by the user to be emanating from the at least one object.
12 . The system of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the system to perform:
modeling sound wave propagation within the representation of the environment in which the system is situated; and
generating audio spatialization data based on the modeling of the sound wave propagation.
13 . The system of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the system to perform:
receiving one or more captured images of the environment in which the system is situated;
receiving sensor data from one or more sensors of the system; and
generating the representation of the environment in which the system is situated based on the one or more captured images and the sensor data.
14 . The system of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the system to perform:
tracking a location of the system within the environment in which the system is situated.
15 . The system of claim 11 , wherein the instructions, when executed by the at least one processor, further cause the system to perform:
receiving a question regarding an object of the environment; and
generating an answer to the question based on one or more of a cascade of machine learning models or a knowledge base graph.
16 . A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform a method comprising:
receiving a user request to identify at least one object of an environment in which the computing system is situated;
receiving a classification for the at least one object;
placing an audio tag based on the classification for the at least one object in a representation of the environment, wherein the audio tag is associated with a sound perceived by a user to be emanating from the at least one object; and
generating, in response to the user request, the sound perceived by the user to be emanating from the at least one object.
17 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions, when executed by the at least one processor of the computing system, further cause the computing system to perform:
modeling sound wave propagation within the representation of the environment in which the computing system is situated; and
generating audio spatialization data based on the modeling of the sound wave propagation.
18 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions, when executed by the at least one processor of the computing system, further cause the computing system to perform:
receiving one or more captured images of the environment in which the computing system is situated;
receiving sensor data from one or more sensors of the computing system; and
generating the representation of the environment in which the computing system is situated based on the one or more captured images and the sensor data.
19 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions, when executed by the at least one processor of the computing system, further cause the computing system to perform:
tracking a location of the computing system within the environment in which the computing system is situated.
20 . The non-transitory computer-readable storage medium of claim 16 , wherein the instructions, when executed by the at least one processor of the computing system, further cause the computing system to perform:
receiving a question regarding an object of the environment; and
generating an answer to the question based on one or more of a cascade of machine learning models or a knowledge base graph.