IP Library Granted Patent US 10,977,525
Granted Patent B2
US 10,977,525 · App. 16/370,676 · Granted Apr 13, 2021

Indoor localization using real-time context fusion of visual information from static and dynamic cameras

Inventors: Chelhwon Kim (Palo Alto, CA); Chidansh Amitkumar Bhatt (Mountain View, CA); Miteshkumar Patel (San Jose, CA); Donald Kimber (Foster City, CA)
Assignee: FUJI XEROX CO., LTD.
G06K9/629G06F16/532G06F16/587G06K9/00771G06K9/6232G06N3/08H04N5/247H04W4/33
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,977,525
App. No.
16/370,676
Granted
Apr 13, 2021
Kind
B2
Abstract

A computer-implemented method of localization for an indoor environment is provided, including receiving, in real-time, a dynamic query from a first source, and static inputs from a second source; extracting features of the static inputs by applying a metric learning convolutional neural network (CNN), and aggregating the extracted features of the static inputs to generate a feature transformation; and iteratively extracting features of the dynamic query on a deep CNN as an embedding network and fusing the feature transformation into the deep CNN, and applying a triplet loss function to optimize the embedding network and provide a localization result.

Claims (32)

1. A computer-implemented method of localization for an indoor environment, comprising:

receiving, in real-time, a dynamic query from a first source, and static inputs from a second source;

extracting features of the dynamic query on a deep convolutional neural network (CNN) as an embedding network;

extracting features of the static inputs by applying a CNN as a condition network, and aggregating the extracted features of the static inputs to generate a feature transformation, and to modulate intermediate features of the embedding network by using the feature transformation; and

applying a triplet loss function to optimize the embedding network and the condition network, and to provide a localization result.

2. The computer-implemented method of claim 1 , wherein the localization result comprises a prediction indicative of a location of the first source in the indoor environment.

3. The computer-implemented method of claim 1 , wherein the dynamic query comprises an image and the first source is a mobile terminal device associated with a user, and the real-time static inputs comprise static images from the second source comprising a cameras networked in the indoor environment.

4. The computer-implemented method of claim 1 , wherein the static inputs are geo-tagged.

5. The computer-implemented method of claim 1 , wherein the localization result is provided during an unpredictable condition and/or an unstructured condition in the indoor environment.

6. The computer-implemented method of claim 1 , wherein the unpredictable condition comprises a change in objects and/or persons in the indoor environment, and the unstructured condition comprises a change in a layout of the indoor environment, and wherein the extracted features associated with the static inputs comprise high-level context information, and wherein the feature transformation comprises a scaling parameter and a shifting parameter.

7. The computer-implemented method of claim 1 , the extracting the features of the dynamic query on the deep CNN further comprises applying a metric learning CNN, and iteratively extracting the features of the dynamic query on the deep CNN and fusing the feature transformation into the deep CNN.

8. A server capable of localization for an indoor environment, the server configured to perform the operations of:

receiving, in real-time, a dynamic query from a first source, and static inputs from a second source;

extracting features of the dynamic query on a deep convolutional neural network (CNN) as an embedding network;

extracting features of the static inputs by applying a CNN as a condition network, and aggregating the extracted features of the static inputs to generate a feature transformation, and to modulate intermediate features of the embedding network by using the feature transformation; and

applying a triplet loss function to optimize the embedding network and the condition network, and to provide a localization result.

9. The server of claim 8 , wherein the localization result comprises a prediction indicative of a location of the first source in the indoor environment.

10. The server of claim 8 , wherein the dynamic query comprises an image and the first source is a mobile terminal device associated with a user, and the real-time static inputs comprise static images from the second source comprising a cameras networked in the indoor environment.

11. The server of claim 8 , wherein the static inputs are geo-tagged.

12. The server of claim 8 , wherein the localization result is provided during an unpredictable condition and/or an unstructured condition in the indoor environment, and wherein the unpredictable condition comprises a change in objects and/or persons in the indoor environment, and the unstructured condition comprises a change in a layout of the indoor environment.

13. The server of claim 8 , the extracting the features of the dynamic query on the deep CNN further comprises applying a metric learning CNN, and iteratively extracting the features of the dynamic query on the deep CNN and fusing the feature transformation into the deep CNN.

14. The server of claim 8 , wherein the extracted features associated with the static inputs comprise high-level context information, and wherein the feature transformation comprises a scaling parameter and a shifting parameter.

15. A non-transitory computer readable medium having a storage that stores instructions executed by a processor, the instructions comprising:

receiving, in real-time, a dynamic query from a first source, and static inputs from a second source;

extracting features of the dynamic query on a deep convolutional neural network (CNN) as an embedding network;

extracting features of the static inputs by applying a CNN as a condition network, and aggregating the extracted features of the static inputs to generate a feature transformation; and

applying a triplet loss function to optimize the embedding network and the condition network, and to provide a localization result.

16. The non-transitory computer readable medium of claim 15 , wherein the localization result comprises a prediction indicative of a location of the first source in the indoor environment.

17. The non-transitory computer readable medium of claim 15 , wherein the dynamic query comprises an image and the first source is a mobile terminal device associated with a user, and the real-time static inputs comprise static images from the second source comprising a cameras networked in the indoor environment and wherein the static inputs are geo-tagged.

18. The non-transitory computer readable medium of claim 15 , the extracting the features of the dynamic query on the deep CNN further comprises applying a metric learning CNN, and iteratively extracting the features of the dynamic query on the deep CNN and fusing the feature transformation into the deep CNN.

19. The non-transitory computer readable medium of claim 15 , wherein the localization result is provided during an unpredictable condition and/or an unstructured condition in the indoor environment, wherein the unpredictable condition comprises a change in objects and/or persons in the indoor environment, and the unstructured condition comprises a change in a layout of the indoor environment.

20. The non-transitory computer readable medium of claim 15 , wherein the extracted features associated with the static inputs comprise high-level context information, and wherein the feature transformation comprises a scaling parameter and a shifting parameter.

Assignments (2)
CHANGE OF NAME Recorded Oct 12, 2022
From: FUJI XEROX CO., LTD.
To: FUJIFILM BUSINESS INNOVATION CORP.
Reel/Frame 061657/0790 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2019
From: KIM, CHELHWON; BHATT, CHIDANSH AMITKUMAR; PATEL, MITESHKUMAR; KIMBER, DONALD
To: FUJI XEROX CO., LTD.
Reel/Frame 048746/0346 →
Continuity (1)
Related Publication 20200311468A1 · Oct 1, 2020