IP Library Granted Patent US 12,277,160
Granted Patent B2
US 12,277,160 · App. 17/440,166 · Granted Apr 15, 2025

Search apparatus, training apparatus, search method, training method, and program

Inventors: Krishna Onkar (Tokyo, JP); Go Irie (Tokyo, JP); Xiaomeng Wu (Tokyo, JP); Takahito Kawanishi (Tokyo, JP); Kunio Kashino (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F16/434G06F16/433G06F18/2148G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,160
App. No.
17/440,166
Granted
Apr 15, 2025
Kind
B2
Abstract

A search apparatus for searching media data for a target region that matches query data includes a first feature extraction unit configured to extract a first feature vector from the query data using a first trained neural network; a second feature extraction unit configured to obtain a first region from the media data and extract a second feature vector from the first region using a second trained neural network; a localization unit configured to determine a candidate for the target region using a third trained neural network, based on the first feature vector, the second feature vector, and the first region or a location of the first region; and a control unit configured to repeat the operations of the second feature extraction unit and the localization unit until a predetermined condition is satisfied, by using the determined candidate for the target region as the first region to be used by the second feature extraction unit.

Claims (27)

1. An apparatus, comprising:

a processor; and

a memory that stores instructions, that when executed by the processor, cause the processor to

extract a first feature vector from query data by a first neural network;

obtain a first region from media data and extract a second feature vector from the first region by a second neural network;

determine a candidate for a target region by a third neural network, based on the first feature vector, the second feature vector, and the first region or a location of the first region;

repeat the operation of obtaining the first region and the operation of determining the candidate for the target region until a predetermined condition is satisfied, by using the determined candidate for the target region as the first region to be used in the operation of obtaining the first region, wherein the third neural network has been trained to minimize a number of executing the repeat operation according to reinforcement learning, and the candidate for the target region depends at least on a previously determined candidate during a previous iteration of the repeat operation;

determine whether the determined candidate for the target region captures the target region, and update parameters of the first neural network, the second neural network, and the third neural network, based on a result of the determination of the candidate capturing the target region;

extracting result media data from the media data according to the determined candidate for the target region to remove; and

presenting the result media data as a result of a query according to the query data.

2. The apparatus as claimed in claim 1 , wherein the first neural network, the second neural network, and the third neural network are trained using training query data and training media data based on a result of determining whether the candidate for the target region determined captures the target region, for the first feature vector extracted from the training query data and the first region obtained from the training media data.

3. The apparatus as claimed in claim 1 , wherein a parameter of the first neural network is the same as a parameter of the second neural network.

4. The apparatus as claimed in claim 1 , wherein the instructions further cause the processor to

down-sample the media data into down-sampled media data and obtain the first region based on the down-sampled media data using a fourth neural network.

5. The apparatus as claimed in claim 1 , wherein the updating parameters of the first neural network, the second neural network, and the third neural network, causes an increase of similarity between the first feature vector and the second feature vector.

6. The apparatus as claimed in claim 4 , wherein a parameter of the fourth neural network is updated together with the parameters of the first neural network, the second neural network, and the third neural network, based on the result of determining whether the candidate for the target region captures the target region.

7. A computer-implemented method, comprising:

extracting a first feature vector from query data by a first neural network;

obtaining a first region from media data and extracting a second feature vector from the first region by a second neural network;

determining a candidate for a target region using a third neural network, based on the first feature vector, the second feature vector, and the first region or a location of the first region;

repeating the obtaining step and the determining step until a predetermined condition is satisfied, by using the determined candidate for the target region as the first region to be used in the obtaining step, wherein the third neural network has been trained to minimize a number of executing the repeat operation according to reinforcement learning, and the candidate for the target region depends at least on a previously determined candidate during a previous iteration of the repeat operation;

determine whether the determined candidate for the target region captures the target region, and update parameters of the first neural network, the second neural network, and the third neural network, based on a result of the determination of the candidate capturing the target region;

extracting result media data from the media data according to the determined candidate for the target region to remove; and

presenting the result media data as a result of a query according to the query data.

8. The computer-implemented method according to claim 7 ,

wherein the updating parameters of the first neural network, the second neural network, and the third neural network causes an increase of similarity between the first feature vector and the second feature vector.

9. A non-transitory computer-readable storage medium that stores therein a program that causes a computer to function as an apparatus as claimed in claim 1 .

Assignments (2)
CHANGE OF NAME Recorded Oct 22, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073184/0535 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2021
From: ONKAR, KRISHNA; IRIE, GO; WU, XIAOMENG; KAWANISHI, TAKAHITO; KASHINO, KUNIO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 057508/0663 →
Priority Claims (1)
JP 2019-059437 · Mar 26, 2019 · national
Continuity (1)
Related Publication 20220188345A1 · Jun 16, 2022
References Cited (10)
US 10860898B2 · Yang · 2020 [cited by examiner]
US 20160379041A1 · Rhee · 2016 [cited by examiner]
Pele et al. (2007) “Accelerating pattern matching or how much can you slide?”, In Asian Conference on Computer Vision, Nov. 18, 2007, Springer, Berlin, Heidelberg, pp. 435-446. [cited by applicant]
Dekel et al. (2015) “Best-buddies similarity for robust template matching”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2021-2029. [cited by applicant]
Korman et al. (2013) “Fast-match: Fast affine template matching”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2331-2338. [cited by applicant]
Jiao et al. (2018) “A fast template matching algorithm based on principal orientation difference”, International Journal of Advanced Robotic Systems, 15(3):1729881418778223. [cited by applicant]
Ba et al. (2014) “Multiple object recognition with visual attention”, arXiv preprint arXiv:1412.7755, Dec. 24, 2014. [cited by applicant]
Ablavatski et al. (2017) “Enriched deep recurrent visual attention model for multiple object recognition”, In Applications of Computer Vision (WACV), IEEE Winter Conference on Mar. 24, 2017, pp. 971-978. [cited by applicant]
Mnih et al. (2014) “Recurrent models of visual attention”, In Advances in neural information processing systems, pp. 2204-2212. [cited by applicant]
Hel-Or et al. (2011) “Fast template matching in non-linear tone-mapped images”, In Computer Vision (ICCV), IEEE International Conference on Nov. 6, 2011, pp. 1355-1362. [cited by applicant]