IP Library › Granted Patent US 10,437,879
Granted Patent B2
US 10,437,879 · App. 15/409,500 · Granted Oct 8, 2019

Visual search using multi-view interactive digital media representations

Inventors: Stefan Johannes Josef Holzer (San Mateo, CA); Abhishek Kar (Berkeley, CA); Alexander Jay Bruen Trevor (San Francisco, CA); Pantelis Kalogiros (San Francisco, CA); Ioannis Spanos (Larisa, GR); Radu Bogdan Rusu (San Francisco, CA)
Assignee: Fyusion, Inc.
G06F16/5838G06K9/00201G06K9/00671G06K9/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,437,879
App. No.
15/409,500
Filed
Jan 18, 2017
Granted
Oct 8, 2019
Kind
B2
Art Unit
2159
USPC
707/769
Abstract

Provided are mechanisms and processes for performing visual search using multi-view digital media representations, such as surround views. In one example, a process includes receiving a visual search query that includes a surround view of an object to be searched, where the surround view includes spatial information, scale information, and different viewpoint images of the object. The surround view is compared to stored surround views by comparing spatial information and scale information of the surround view to spatial information and scale information of the stored surround views. A correspondence measure is then generated indicating the degree of similarity between the surround view and a possible match. At least one search result is then transmitted with a corresponding image in response to the visual search query.

Claims (36)

1. A method comprising;

receiving a visual search query, the visual search query including a first multi-view digital media representation of an object to be searched; wherein the first multi-view digital media representation of the object includes spatial information, scale information, and a plurality of different viewpoint images of the object, and wherein the first multi-view digital media representation of the object is capable of being displayed as a three-dimensional model on a display device;

comparing the first multi-view digital media representation to a plurality of stored multi-view digital media representations by comparing spatial information and scale information of the first multi-view digital media representation to spatial information and scale information of the stored multi-view digital media representations, wherein the stored multi-view digital media representations include a second multi-view digital media representation generated by aggregating spatial information, scale information, and a plurality of different viewpoint images obtained when capturing the second multi-view digital media representation;

generating a correspondence measure indicating the degree of similarity between the first multi-view digital media representation and the second multi-view digital media representation; and

transmitting an image associated with second multi-view digital media representation in response to the visual search query.

2. The method of claim 1 , wherein the spatial information comprises depth information.

3. The method of claim 1 , wherein the spatial information comprises visual flow between the different viewpoints.

4. The method of claim 1 , wherein the spatial information comprises three-dimensional location information.

5. The method of claim 1 , wherein scale information is estimated using accelerometer information obtained when capturing the first multi-view digital media representation.

6. The method of claim 1 , wherein scale information is determined using inertial measurement unit (IMU) data obtained when capturing the first multi-view digital media representation.

7. The method of claim 1 , wherein the first multi-view digital media representation of the object further comprises three-dimensional shape information.

8. The method of claim 1 , wherein the first multi-view digital media representation of the object is compared to one or more stored multi-view digital media representations without generating a 3D model of the object.

9. The method of claim 1 , wherein the plurality of different viewpoint images obtained when capturing the second multi-view digital media representation are aggregated using an algorithm selected from the group consisting of pooling, bag of words, vector of locally aggregated descriptors (VLAD), and fisher vectors.

10. The method of claim 1 , wherein comparing the first multi-view digital media representation to a plurality of stored multi-view digital media representations further comprises performing selective visual search by comparing an area of focus in the first multi-view digital media representation to the plurality of stored multi-view digital media representations, the area of focus selected by a user submitting the visual search.

11. The method of claim 1 , wherein comparing the first multi-view digital media representation to a plurality of stored multi-view digital media representations further comprises performing a parameterized visual search by comparing attributes associated with the first multi-view digital media representation to attributes associated with the plurality of stored multi-view digital media representations, wherein the attributes are specified by a user submitting the visual search.

12. The method of claim 11 , wherein the attributes include parameter information that specify dimensions, color, texture, content/object category in various levels of detail.

13. The method of claim 1 , wherein comparing the first multi-view digital media representation to a plurality of stored multi-view digital media representations further comprises performing a viewpoint-aware visual search by comparing viewpoints associated with the first multi-view digital media representation to viewpoints associated with the plurality of stored multi-view digital media representations.

14. The method of claim 1 , wherein comparing the first multi-view digital media representation to a plurality of stored multi-view digital media representations further comprises filtering matches based on social information associated with a user submitting the visual search.

15. A system comprising:

an interface that receives a visual search query, the visual search query including a first multi-view digital media representation of an object to be searched, wherein the first multi-view digital media representation of the object includes spatial information, scale information, and a plurality of different viewpoint images of the object, wherein the first multi-view digital media representation of the object is capable of being displayed as a three-dimensional model on a display device, and wherein the interface further transmits an image associated with a search result in response to the visual search query;

a search server configured to compare the first multi-view digital media representation to a plurality of stored multi-view digital media representation by comparing spatial information and scale information of the first multi-view digital media representation to spatial information and scale information of the stored multi-view digital media representation, wherein the stored multi-view digital media representation include a second multi-view digital media representation generated by aggregating spatial information, scale information, and a plurality of different viewpoint images obtained when capturing the second multi-view digital media representation; and

a front end server configured to generate a correspondence measure indicating the degree of similarity between the first multi-view digital media representation and the second multi-view digital media representation.

16. The system of claim 15 , wherein the search server is further configured to perform selective visual search by comparing an area of focus in the first multi-view digital media representation to the plurality of stored multi-view digital media representation, the area of focus selected by a user submitting the visual search; wherein the search server is further configured to perform a parameterized visual search by comparing attributes associated with the first multi-view digital media representation to attributes associated with the plurality of stored multi-view digital media representations, wherein the attributes are specified by the user submitting the visual search; wherein the search server is further configured to perform a viewpoint-aware visual search by comparing viewpoints associated with the first multi-view digital media representation to viewpoints associated with the plurality of stored multi-view digital media representations; and

wherein the search server is further configured to filter matches based on social information associated with the user submitting the visual search.

17. A computer readable medium comprising:

computer code that receives a visual search query, the visual search query including a multi-view digital media representation of an object to be searched, wherein the multi-view digital media representation of the object includes spatial information, scale information, and a plurality of different viewpoint images of the object, and wherein the multi-view digital media representation of the object is capable of being displayed as a three-dimensional model on a display device;

computer code that compares the first multi-view digital media representation to a plurality of stored multi-view digital media representations by comparing spatial information and scale information of the first multi-view digital media representation to spatial information and scale information of the stored multi-view digital media representations, wherein the stored multi-view digital media representation include a second multi-view digital media representation generated by aggregating spatial information, scale information, and a plurality of different viewpoint images obtained when capturing the second multi-view digital media representation;

computer code that generates a correspondence measure indicating the degree of similarity between the first multi-view digital media representation and the second multi-view digital media representation; and

computer code that transmits an image associated with second multi-view digital media representation in response to the visual search query.

18. The computer readable medium of claim 17 , further comprising:

computer code that performs selective visual search by comparing an area of focus in the first multi-view digital media representation to the plurality of stored multi-view digital media representations, the area of focus selected by a user submitting the visual search; and

computer code that performs a parameterized visual search by comparing attributes associated with the first s multi-view digital media representations to attributes associated with the plurality of stored multi-view digital media representations, wherein the attributes are specified by the user submitting the visual search.

19. The computer readable medium of claim 17 , further comprising:

computer code that performs a viewpoint-aware visual search by comparing viewpoints associated with the first multi-view digital media representation to viewpoints associated with the plurality of stored multi-view digital media representations.

20. The computer readable medium of claim 17 , further comprising:

computer code that filters matches based on social information associated with a user submitting the visual search.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2017
From: HOLZER, STEFAN JOHANNES JOSEF; KAR, ABHISHEK; TREVOR, ALEXANDER JAY BRUEN; KALOGIROS, PANTELIS; SPANOS, IOANNIS; RUSU, RADU BOGDAN
To: FYUSION, INC.
Reel/Frame 041186/0964 →
Continuity (1)
Related Publication 20180203877A1 · Jul 19, 2018
Cited By (1)
US 12,235,890