IP Library › Granted Patent US 12,518,491
Granted Patent B2
US 12,518,491 · App. 18/097,040 · Granted Jan 6, 2026

Supplementing user perception and experience with augmented reality (AR) and artificial intelligence (AI) techniques utilizing an artificial intelligence (AI) agent

Inventors: Fan Zhang (Kirkland, WA); William Wong (Poughquag, NY); Anoop Kumar Sinha (Palo Alto, CA); Timothy Rosenberg (Arlington, VA); Chen Sun (Berkeley, CA); Gary Vu Nguyen (Edmonds, WA); Sofia Gallo Pavajeau (Seattle, WA); Agustya Mehta (San Carlos, CA); Johana Gabriela Coyoc Escudero (Milpitas, CA); Leonid Vladimirov (New York, NY); James Schultz (Redmond, WA)
Assignee: Meta Platforms, Inc.
G06T19/006G02B27/0172G06V10/82G06V10/84G02B2027/0178
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,491
App. No.
18/097,040
Granted
Jan 6, 2026
Kind
B2
Abstract

According to examples, a system for supplementing user perception and experience via augmented reality (AR), artificial intelligence (AI), and machine-learning (ML) techniques is described. The system may include a processor and a memory storing instructions. The processor, when executing the instructions, may cause the system to receive data associated with at least one of a location, context, or setting and determine, using at least one artificial intelligence (AI) model and at least one machine learning (ML) model, relationships between objects in the at least one of the location, context, or setting. The processor, when executing the instructions, may then apply an artificial intelligence (AI) agent to analyze the relationships and generate a three-dimensional (3D) mapping of the at least one of the location, context, or setting and provide an output to aid a user's perception and experience.

Claims (42)

1 . A system, comprising:

a processor; and

a memory storing instructions, which when executed by the processor, cause the processor to:

determine, using at least one artificial intelligence (AI) model, relationships between objects in a field of view of a camera of a headset worn by a user;

apply an artificial intelligence (AI) agent associated with the user to analyze the relationships determined by the at least one artificial intelligence (AI) model and generate a three-dimensional (3D) mapping that includes one or more objects within the field of view, wherein the artificial intelligence (AI) agent is personalized to the user based on previous user interactions with the artificial intelligence (AI) agent; and

provide an output via at least one output modality of a plurality of output modalities, the at least one output modality determined by the artificial intelligence (AI) agent, to aid a user's perception and experience of the field of view, wherein the output is based on the three-dimensional (3D) mapping and includes a prediction, made by the artificial intelligence (AI) agent, for a user activity that is associated with the field of view.

2 . The system of claim 1 , wherein the instructions, when executed by the processor, further cause the processor to implement the artificial intelligence (AI) agent to conduct a localization and mapping analysis to generate the three-dimensional (3D) mapping.

3 . The system of claim 1 , wherein the headset is a pair of augmented reality (AR) glasses.

4 . The system of claim 1 , wherein the at least one artificial intelligence (AI) model comprises at least one of a large language model (LLM), a generative adversarial network (GAN), a tree-based model, a Bayesian network, a support vector, clustering, a kernel method, a spline, or a knowledge graph.

5 . The system of claim 1 , wherein generating the three-dimensional (3D) mapping includes:

determining a location of each of the one or more objects relative to a location of the user based on, at least, the field of view.

6 . The system of claim 1 , wherein the instructions, when executed by the processor, further cause the processor to:

receive image data of the field of view from the camera of the headset; and

perform image analysis of the image data, wherein the determining the relationships between objects in the field of view is based on the image data.

7 . The system of claim 1 , wherein the instructions, when executed by the processor, further cause the processor to:

determine a risk associated with the one or more objects, wherein the output is further based on the risk.

8 . A method for supplementing user perception and experience, comprising:

determining, using at least one artificial intelligence (AI) model, relationships between objects in a field of view of a camera of a headset worn by a user;

applying an artificial intelligence (AI) agent associated with the user to analyze the relationships determined by the at least one artificial intelligence (AI) model and generating a three-dimensional (3D) mapping that includes one or more objects within the field of view, wherein the artificial intelligence (AI) agent is personalized to the user based on previous user interactions with the artificial intelligence (AI) agent; and

providing an output via at least one output modality of a plurality of output modalities, the at least one output modality determined by the artificial intelligence (AI) agent, to aid a user's perception and experience of the field of view wherein the output is based on the three-dimensional (3D) mapping and includes a prediction, made by the artificial intelligence (AI) agent, for a user activity that is associated with the field of view.

9 . The method of claim 8 , further comprising implementing the artificial intelligence (AI) agent to conduct a localization and mapping analysis to generate the three-dimensional (3D) mapping.

10 . The method of claim 8 , wherein the headset is a pair of augmented reality (AR) glasses.

11 . The method of claim 8 , wherein the at least one artificial intelligence (AI) model comprises at least one of a large language model (LLM), a generative adversarial network (GAN), a tree-based model, a Bayesian network, a support vector, clustering, a kernel method, a spline, or a knowledge graph.

12 . The method of claim 8 , wherein generating the three-dimensional (3D) mapping includes:

determining a location of each of the one or more objects relative to a location of the user based on, at least, the field of view.

13 . The method of claim 8 , further comprising:

receiving image data of the field of view from the camera of the headset; and

performing image analysis of the image data, wherein the determining the relationships between objects in the field of view is based on the image data.

14 . A non-transitory computer-readable storage medium having an executable stored thereon, which when executed instructs a processor to:

determine, using at least one artificial intelligence (AI) model, relationships between objects in a field of view of a camera of a headset worn by a user;

apply an artificial intelligence (AI) agent associated with the user to analyze the relationships determined by the at least one artificial intelligence (AI) model and generate a three-dimensional (3D) mapping that includes one or more objects within the field of view, wherein the artificial intelligence (AI) agent is personalized to the user based on previous user interactions with the artificial intelligence (AI) agent; and

provide an output via at least one output modality of a plurality of output modalities, the at least one output modality determined by the artificial intelligence (AI) agent, to aid a user's perception and experience of the field of view, wherein the output is based on the three-dimensional (3D) mapping and includes a prediction, made by the artificial intelligence (AI) agent, for a user activity that is associated with the field of view.

15 . The non-transitory computer-readable storage medium of claim 14 , wherein the executable when executed further instructs the processor to implement the artificial intelligence (AI) agent to conduct a localization and mapping analysis to generate the three-dimensional (3D) mapping.

16 . The non-transitory computer-readable storage medium of claim 14 , wherein generating the three-dimensional (3D) mapping includes:

determining a location of each of the one or more objects relative to a location of the user based on, at least, the field of view.

17 . The non-transitory computer-readable storage medium of claim 14 , wherein the executable when executed further instructs the processor to:

receive image data of the field of view from the camera of the headset; and

perform image analysis of the image data, wherein the determining the relationships between objects in the field of view is based on the image data.

18 . The non-transitory computer-readable storage medium of claim 14 , wherein the executable when executed further instructs the processor to:

determine a risk associated with the one or more objects, wherein the output is further based on the risk.

19 . The non-transitory computer-readable storage medium of claim 14 , wherein the headset is a pair of augmented reality (AR) glasses.

20 . The non-transitory computer-readable storage medium of claim 14 , wherein the at least one artificial intelligence (AI) model comprises at least one of a large language model (LLM), a generative adversarial network (GAN), a tree-based model, a Bayesian network, a support vector, clustering, a kernel method, a spline, or a knowledge graph.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: ZHANG, FAN; WONG, WILLIAM; SINHA, ANOOP KUMAR; ROSENBERG, TIMOTHY; SUN, CHEN; NGUYEN, GARY VU; GALLO PAVAJEAU, SOFIA; MEHTA, AGUSTYA; COYOC ESCUDERO, JOHANA GABRIELA; VLADIMIROV, LEONID; SCHULTZ, JAMES
To: META PLATFORMS, INC.
Reel/Frame 065926/0317 →
Continuity (1)
Related Publication 20240242442A1 · Jul 18, 2024
References Cited (5)
US 9576460B2 · Dayal · 2017 [cited by examiner]
US 11663781B1 · Alexander · 2023 [cited by examiner]
US 11955028B1 · Nash · 2024 [cited by examiner]
US 20120143808A1 · Karins · 2012 [cited by examiner]
US 20200409457A1 · Terrano · 2020 [cited by examiner]