IP Library Granted Patent US 12,311,265
Granted Patent B2
US 12,311,265 · App. 17/541,549 · Granted May 27, 2025

Systems and methods for training a model to determine a type of environment surrounding a user

Inventors: Brandon Sangston (Foster City, CA); Andrew Young (San Mateo, CA)
Assignee: Sony Interactive Entertainment Inc.
A63F13/54A63F13/215A63F13/355G06F3/165G06N20/00G06V20/50A63F2300/1081A63F2300/538A63F2300/6081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,311,265
App. No.
17/541,549
Filed
Dec 3, 2021
Granted
May 27, 2025
Kind
B2
Art Unit
2692
USPC
381/1
Abstract

A method for determining an environment in which a user is located is described. The method includes receiving a plurality of sets of audio data based on sounds emitted in a plurality of environments. Each of the plurality of environments has a different combination of objects. The method further includes receiving input data regarding the plurality of environments, and training an artificial intelligence (AI) model based on the plurality of sets of audio data and the input data. The method includes applying the AI model to audio data captured from an environment surrounding the first user to determine a type of the environment.

Claims (53)

1. A method for determining a real-world environment in which a first user is located, comprising:

receiving a plurality of sets of audio data generated from sounds emitted in a plurality of real-world environments, wherein each of the plurality of real-world environments has a different combination of objects;

extracting a plurality of features from the plurality of sets of audio data, wherein the plurality of features include a plurality of amplitudes of the plurality of sets of audio data, a plurality of frequencies of the plurality of sets of audio data, and a plurality of sense directions in which the sounds are sensed;

classifying the plurality of features to output associations between the plurality of features and a plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, a plurality of arrangements of the objects, and a plurality of states of the objects within the plurality of real-world environments;

receiving input data regarding the plurality of real-world environments;

training an artificial intelligence (AI) model based on the plurality of sets of audio data generated from the sounds in the plurality of real-world environments and based on the input data regarding the plurality of real-world environments, wherein said training the AI model includes:

providing, to the AI model, the associations between the plurality of features and the plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, the plurality of arrangements of the objects, and the plurality of states of the objects within the plurality of real-world environments; and

determining, by the AI model, a plurality of probabilities based on the associations between the plurality of features and the plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, the plurality of arrangements of the objects, and the plurality of states of the objects within the plurality of real-world environments; and

applying the AI model to audio data captured from the real-world environment surrounding the first user to determine that a type of the real-world environment in which the first user is located includes one of an indoor environment and an outdoor environment.

2. The method of claim 1 , further comprising:

receiving an indication of the type of the real-world environment to be simulated;

accessing, based on the type of the real-world environment to be simulated, the audio data captured from the real-world environment;

providing the audio data captured from the real-world environment to a client device for outputting a sound corresponding to the type of the real-world environment.

3. The method of claim 1 , wherein the plurality of sets of audio data include audio data that is generated from sounds emitted from one or more of the objects in the plurality of real-world environments and sounds that reflected from remaining of the objects.

4. The method of claim 1 , wherein the input data includes data identifying the objects in the plurality of real-world environments or image data captured by cameras in the plurality of real-world environments or a combination thereof.

5. The method of claim 1 , wherein the plurality of sets of audio data is captured when a second user moves from one location to another.

6. The method of claim 1 , wherein the plurality of sets of audio data is captured when a plurality of users including a second user and a third user are at a location.

7. The method of claim 1 , further comprising applying the AI model to the audio data captured from the real-world environment surrounding the first user to identify one or more objects within the real-world environment surrounding the first user or one or more states of the one or more objects or an arrangement of the one or more objects or a combination thereof.

8. A server for determining a real-world environment in which a user is located, comprising:

a processor configured to:

receive, via a computer network, a plurality of sets of audio data generated from sounds emitted in a plurality of real-world environments, wherein each of the plurality of real-world environments has a different combination of objects;

extract a plurality of features from the plurality of sets of audio data, wherein the plurality of features include a plurality of amplitudes of the plurality of sets of audio data, a plurality of frequencies of the plurality of sets of audio data, and a plurality of sense directions in which the sounds are sensed;

classify the plurality of features to output associations between the plurality of features and a plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, a plurality of arrangements of the objects, and a plurality of states of the objects within the plurality of real-world environments;

receive, via the computer network, input data regarding the plurality of real-world environments;

train an artificial intelligence (AI) model based on the plurality of sets of audio data generated from the sounds in the plurality of real-world environments and based on the input data regarding the plurality of real-world environments, wherein to train the AI model, the processor is configured to:

provide, to the AI model, the associations between the plurality of features and the plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, the plurality of arrangements of the objects, and the plurality of states of the objects within the plurality of real-world environments; and

determine, using the AI model, a plurality of probabilities based on the associations between the plurality of features and the plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, the plurality of arrangements of the objects, and the plurality of states of the objects within the plurality of real-world environments; and

apply the AI model to audio data captured from the real-world environment surrounding the user to determine that a type of the real-world environment in which the user is located includes one of an indoor environment and an outdoor environment; and

a memory device coupled to the processor.

9. The server of claim 8 , wherein the processor is configured to:

receive an indication of the type of the real-world environment to be simulated;

access, based on the type of the real-world environment to be simulated, the audio data captured from the real-world environment;

provide the audio data captured from the real-world environment to a client device for outputting a sound corresponding to the type of the real-world environment.

10. The server of claim 8 , wherein the plurality of sets of audio data include audio data that is generated from sounds emitted from one or more of the objects in the plurality of real-world environments and sounds that are reflected from remaining of the objects.

11. The server of claim 8 , wherein the input data includes data identifying the objects in the plurality of real-world environments or image data captured by cameras in the plurality of real-world environments or a combination thereof.

12. The server of claim 8 , wherein the processor is configured to apply the AI model to the audio data captured from the real-world environment surrounding the user to identify one or more objects within the real-world environment surrounding the user or one or more states of the one or more objects or an arrangement of the one or more objects or a combination thereof.

13. A system for determining a real-world environment in which a user is located, comprising:

a plurality of client devices configured to:

generate a plurality of sets of audio data generated from sounds emitted in a plurality of real-world environments, wherein each of the plurality of real-world environments has a different combination of objects; and

receive input data regarding the plurality of real-world environments; and

a server coupled to the plurality of client devices, wherein the server is configured to:

receive the plurality of sets of audio data from the plurality of client devices via a computer network;

extract a plurality of features from the plurality of sets of audio data, wherein the plurality of features include a plurality of amplitudes of the plurality of sets of audio data, a plurality of frequencies of the plurality of sets of audio data, and a plurality of sense directions of sensing the sounds;

classify the plurality of features to output associations between the plurality of features and a plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, a plurality of arrangements of the objects, and a plurality of states of the objects within the plurality of real-world environments;

receive the input data regarding the plurality of real-world environments from the plurality of client devices via the computer network;

train an artificial intelligence (AI) model based on the plurality of sets of audio data generated from the sounds in the plurality of real-world environments and based on the input data regarding the plurality of real-world environments, wherein to train the AI model, the server is configured to:

provide, to the AI model, the associations between the plurality of features and the plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, the plurality of arrangements of the objects, and the plurality of states of the objects within the plurality of real-world environments; and

determine, using the AI model, a plurality of probabilities based on the associations between the plurality of features and the plurality of types of the plurality of real-world environments, the objects within the plurality of real-world environments, the plurality of arrangements of the objects, and the plurality of states of the objects within the plurality of real-world environments; and

apply the AI model to audio data captured from the real-world environment surrounding the user to determine that a type of the real-world environment in which the user is located includes one of an indoor environment and an outdoor environment.

14. The system of claim 13 , wherein the server is configured to:

receive an indication of the type of the real-world environment to be simulated;

access, based on the type of real-world environment to be simulated, the audio data captured from the real-world environment;

provide the audio data captured from the real-world environment to a client device for outputting a sound corresponding to the type of real-world environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2021
From: SANGSTON, BRANDON; YOUNG, ANDREW
To: SONY INTERACTIVE ENTERTAINMENT INC.
Reel/Frame 058280/0890 →
Continuity (1)
Related Publication 20230173387A1 · Jun 8, 2023
References Cited (16)
US 8767968B2 · Flaks · 2014 [cited by examiner]
US 11551407B1 · Stanney · 2023 [cited by examiner]
US 20170287218A1 · Nuernberger · 2017 [cited by examiner]
US 20190228589A1 · Dascola · 2019 [cited by examiner]
US 20190236416A1 · Wang · 2019 [cited by examiner]
US 20190392212A1 · Sawhney · 2019 [cited by examiner]
US 20200388068A1 · Yeung · 2020 [cited by examiner]
US 20210058731A1 · Koike et al. · 2021 [cited by applicant]
US 20220101623A1 · Walsh et al. · 2022 [cited by applicant]
US 20220392478A1 · Hijazi · 2022 [cited by examiner]
US 20230147573A1 · Chien · 2023 [cited by examiner]
US 20230158409A1 · Gardner · 2023 [cited by examiner]
WO WO2021216060A1 · 2021 [cited by examiner]
PCT/US2022/050499, Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or the Declaration, PCT/ISA/220, and the International Search Report, P… [cited by applicant]
Chen et al. “SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning”, arxiv.org, Cornell University, Ithaca, NY, Jun. 16, 2022, XP091252433. [cited by applicant]
Chen et al. “SoundSpaces: Audio-visual Navigation in 3D Environments”, arxiv.org, Cornell University, Ithaca, NY Aug. 20, 2020, XP081744383. [cited by applicant]