IP Library › Granted Patent US 11,366,978
Granted Patent B2
US 11,366,978 · App. 16/295,400 · Granted Jun 21, 2022

Data recognition apparatus and method, and training apparatus and method

Inventors: Insoo Kim (Seongnam-si, KR); Kyuhong Kim (Seoul, KR); Chang Kyu Choi (Seongnam-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06K9/0055G06K9/00523G06K9/6269G06N3/08G06N20/10G10L15/02G10L15/16G10L17/02G10L17/04G10L17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,366,978
App. No.
16/295,400
Granted
Jun 21, 2022
Kind
B2
Abstract

A data recognition method includes: extracting a feature map from input data based on a feature extraction layer of a data recognition model; pooling component vectors from the feature map based on a pooling layer of the data recognition model; and generating an embedding vector by recombining the component vectors based on a combination layer of the data recognition model.

Claims (60)

1. A processor-implemented data recognition method, comprising:

extracting, by one or more processors, a feature map from input data based on a feature extraction layer of a data recognition model;

pooling, by the one or more processors, component vectors from the feature map based on a pooling layer of the data recognition model, wherein the pooling of the component vectors comprises extracting, from the feature map, component vectors that are orthogonal to each other;

generating, by the one or more processors, an embedding vector by recombining the component vectors based on a combination layer of the data recognition model; and

recognizing, by the one or more processors, the input data based on the embedding vector, the recognizing of the input data including one of a speaker recognition and an image recognition.

2. The data recognition method of claim 1 , wherein the recognizing of the input data comprises indicating, among registered users, a user mapped to a registration vector verified to match the embedding vector.

3. A processor-implemented data recognition method, comprising:

extracting, by one or more processors, a feature map from input data based on a feature extraction layer of a data recognition model;

pooling, by the one or more processors, component vectors from the feature map based on a pooling layer of the data recognition model;

generating, by the one or more processors, an embedding vector by recombining the component vectors based on a combination layer of the data recognition model; and

recognizing, by the one or more processors, the input data based on the embedding vector, the recognizing of the input data including one of a speaker recognition and an image recognition,

wherein the pooling of the component vectors comprises:

extracting, from the feature map, a first component vector based on the pooling layer; and

extracting, from the feature map, a second component vector having a cosine similarity to the first component vector that is less than a threshold similarity, based on the pooling layer.

4. The data recognition method of claim 1 , wherein the generating of the embedding vector comprises:

calculating an attention weight of the component vectors from the component vectors through an attention layer; and

calculating the embedding vector by applying the attention weight to the component vectors.

5. The data recognition method of claim 1 , wherein generating of the embedding vector comprises:

calculating an attention weight of the component vectors from the component vectors through an attention layer;

generating weighted component vectors by applying the attention weight to each of the component vectors; and

calculating the embedding vector by linearly combining the weighted component vectors.

6. The data recognition method of claim 1 , wherein the generating of the embedding vector comprises generating normalized component vectors by normalizing the component vectors.

7. The data recognition method of claim 1 , wherein the extracting of the feature map comprises extracting a number of feature maps using the feature extraction layer, wherein the feature extraction layer includes a convolutional layer, and wherein the number of feature maps is proportional to a number of component vectors to be pooled.

8. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the data recognition method of claim 1 .

9. A data recognition apparatus, comprising:

a memory configured to store a data recognition model comprising a feature extraction layer, a pooling layer, and a combination layer; and

a processor configured to:

extract a feature map from input data based on the feature extraction layer;

pool component vectors from the feature map based on the pooling layer, wherein the pooling of the component vectors comprises extracting, from the feature map, component vectors that are orthogonal to each other;

generate an embedding vector by recombining the component vectors based on the combination layer; and

recognize the input data based on the embedding vector, the recognizing of the input data including one of a speaker recognition and an image recognition.

10. A speaker recognition method, comprising:

extracting, based on a feature extraction layer of a speech recognition model, a feature map from speech data input by a speaker;

pooling component vectors from the feature map based on a pooling layer of the speech recognition model, wherein the pooling of the component vectors comprises extracting, from the feature map, component vectors that are orthogonal to each other;

generating an embedding vector by recombining the component vectors based on a combination layer of the speech recognition model;

recognizing the speaker based on the embedding vector; and

indicating a recognition result of the recognizing the speaker.

11. The speaker recognition method of claim 10 , wherein the indicating of the recognition result comprises verifying whether the speaker is a registered user based on a result of comparing the embedding vector to a registration vector corresponding to the registered user.

12. The speaker recognition method of claim 11 , wherein the indicating of the recognition result further comprises:

in response to a cosine similarity between the embedding vector and the registration vector exceeding a threshold recognition value, determining verification of the speech data to be successful; and

in response to the verification being determined to be successful, unlocking a device.

13. A speaker recognition apparatus, comprising:

a memory configured to store a speech recognition model including a feature extraction layer, a pooling layer, and a combination layer; and

a processor configured to

extract, based on the feature extraction layer, a feature map from speech data input by a speaker,

pool component vectors from the feature map based on the pooling layer, wherein the pooling of the component vectors comprises extracting, from the feature map, component vectors that are orthogonal to each other,

generate an embedding vector by recombining the component vectors based on the combination layer,

recognize the speaker based on the embedding vector, and

indicate a recognition result of the recognizing the speaker.

14. The speaker recognition apparatus of claim 13 , wherein the recognizing of the speaker based on the embedding vector comprises recognizing the speaker based on a similarity between the embedding vector and a registration vector corresponding to a registered user.

15. The speaker recognition apparatus of claim 13 , wherein the recombining the component vectors based on the combination layer comprises generating weighted component vectors by applying an attention weight to the component vectors, and linearly combining the weighted component vectors.

16. The speaker recognition apparatus of claim 13 , wherein the extracting of the feature map comprises extracting convolved data from the speech data based on a convolutional layer of the feature extraction layer, and pooling the convolved data.

17. An image recognition apparatus, comprising:

a memory configured to store an image recognition model including a feature extraction layer, a pooling layer, and a combination layer; and

a processor configured to

extract, based on the feature extraction layer, a feature map from image data obtained from a captured image of user,

pool component vectors from the feature map based on the pooling layer, wherein the pooling of the component vectors comprises extracting, from the feature map, component vectors that are orthogonal to each other,

generate an embedding vector by recombining the component vectors based on the combination layer,

recognize the user based on the embedding vector, and

indicate a recognition result of the recognizing the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2019
From: KIM, INSOO; KIM, KYUHONG; CHOI, CHANG KYU
To: SAMSUNG ELECTRONCIS CO., LTD.
Reel/Frame 048530/0413 →
Priority Claims (1)
KR 10-2018-0126436 · Oct 23, 2018 · national
Continuity (1)
Related Publication 20200125820A1 · Apr 23, 2020
Cited By (2)
US 12,567,414 US 12,670,912