IP Library › Granted Patent US 11,797,843
Granted Patent B2
US 11,797,843 · App. 16/796,739 · Granted Oct 24, 2023

Hashing-based effective user modeling

Inventors: Peng Zhou (Irvine, CA); Yingnan Zhu (Irvine, CA); Xiangyuan Zhao (Irvine, CA); Hong-hoe Kim (Aliso Viejo, CA); Hyun Chul Lee (Mountain View, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N3/08G06F16/2264G06F17/15G06N3/045G06N20/00G06Q30/0201G06Q30/0205G06Q30/0276
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,797,843
App. No.
16/796,739
Granted
Oct 24, 2023
Kind
B2
Abstract

In one embodiment, a method includes receiving user behavior data and contextual information associated with the user behavior data, the contextual information including a first data portion associated with a first context type. The method includes generating, from the user behavior data and the contextual information using a hashing algorithm, a first heterogeneous hash code including a first portion representing the user behavior data and a second hash code portion representing the first data portion associated with the first context type. The method includes accessing a second heterogeneous hash code including a third hash code portion representing a second data portion associated with the first context type. The method includes comparing the first heterogeneous hash code with the second heterogeneous hash code including determining similarity between the second hash code portion of the first heterogeneous hash code and the third hash code portion of the second heterogenous hash code.

Claims (65)

1. A computer-implemented method comprising, by a computer processor:

receiving user behavior data and contextual information associated with the user behavior data, the contextual information including a first data portion associated with a first context type;

generating, from the user behavior data and the contextual information, a first heterogeneous hash code including a first hash code portion representing the user behavior data and, separate from the first hash code portion, a second hash code portion representing the first data portion associated with the first context type, wherein generating the first heterogeneous hash code comprises:

generating, by a network layer of a machine-learning architecture:

a user-behavior representation based on the user behavior data and a weighted combination of at least (1) the user behavior data and (2) the contextual information; and

a first-context representation based on the first data portion and a weighted combination of at least (1) the first data portion and (2) the user behavior data;

embedding, by a hash layer of the machine-learning architecture, the user-behavior representation and the first-context representation;

binarizing the embedded user-behavior representation to generate the first hash code portion; and

binarizing the embedded first-context representation to generate the second hash code portion;

retrieving, from a heterogeneous hash code database, a second heterogeneous hash code including a third hash code portion representing a second data portion associated with the first context type, wherein, as stored in the heterogeneous hash code database, the third hash code portion is accessible separately from another portion of the second heterogeneous hash code;

comparing the first heterogeneous hash code with the second heterogeneous hash code including by determining a similarity between the second hash code portion of the first heterogeneous hash code and the third hash code portion of the second heterogeneous hash code; and

determining, based on the comparison, a similarity between at least two of a plurality of users.

2. The method of claim 1 , wherein the contextual information further includes a third data portion associated with a second context type; and wherein the first heterogeneous hash code further includes a fourth hash code portion representing the third data portion associated with the second context type.

3. The method of claim 2 , wherein the first context type includes at least one of a location type, a date/time type, or a demographic type, and wherein the first context type is different from the second context type.

4. The method of claim 2 , wherein the second heterogeneous hash code includes a fifth hash code portion representing a fourth data portion associated with the second context type; and

wherein comparing the first heterogeneous hash code with the second heterogeneous hash code further includes determining a similarity between the fourth hash code portion of the first heterogeneous hash code and the fifth hash code portion of the second heterogeneous hash code.

5. The method of claim 1 , wherein the second hash code portion is further based on a correlation between the first data portion and the third data portion; and

wherein the fourth hash code portion is further based on a correlation between the third data portion and the first data portion.

6. The method of claim 5 , wherein the correlation is measured by a gate layer of the machine learning architecture.

7. The method of claim 1 , wherein the first hash code portion is further based on a correlation between the user behavior data and the first data portion; and

wherein the second hash code portion is further based on a correlation between the first data portion and the user behavior data.

8. The method of claim 1 , further comprising:

identifying, based on comparing the first heterogeneous hash code with the second heterogeneous hash code, a new segment of one or more users within a threshold level of similarity to a given set of one or more users; and

expanding the given set by adding at least one user from the new segment.

9. The method of claim 1 , further comprising:

storing the first heterogeneous hash code in the heterogeneous hash code database, wherein, while stored, the first hash code portion is accessible separately from the second hash code portion, and wherein the second hash code portion is stored in association with the first context type.

10. The method of claim 9 , wherein comparing the first heterogeneous hash code with the second heterogeneous hash code comprises:

receiving a request to compare the first heterogeneous hash code and the second heterogeneous hash code based on the first context type;

retrieving hash code portions of the first heterogeneous hash code and the second heterogeneous hash code associated with the first context type from the heterogeneous hash code database.

11. The method of claim 1 , wherein the machine-learning architecture generates similar hash code values for similar data portions of the first context type and generates dissimilar hash code values for dissimilar data portions of the first context type.

12. The method of claim 1 , wherein the determining a similarity between the second hash code portion of the first heterogeneous hash code and the third hash code portion of the second heterogeneous hash code comprises:

calculating a bitwise distance between the second hash code portion and the third hash code portion.

13. An apparatus comprising:

one or more non-transitory computer-readable storage media embodying instructions; and

one or more processors coupled to the storage media and configured to execute the instructions to:

receive user behavior data and contextual information associated with the user behavior data, the contextual information including a first data portion associated with a first context type;

generate, from the user behavior data and the contextual information using a hashing algorithm, a first heterogeneous hash code including a first hash code portion representing the user behavior data and a second hash code portion representing the first data portion associated with the first context type, wherein generating the first heterogeneous hash code comprises:

generating, by a network layer of a machine-learning architecture:

a user-behavior representation based on the user behavior data and a weighted combination of at least (1) the user behavior data and (2) the contextual information; and

a first-context representation based on the first data portion and a weighted combination of at least (1) the first data portion and (2) the user behavior data;

embedding, by a hash layer of the machine-learning architecture, the user-behavior representation and the first-context representation;

binarizing the embedded user-behavior representation to generate the first hash code portion; and

binarizing the embedded first-context representation to generate the second hash code portion;

retrieve, from a heterogeneous hash code database, a second heterogeneous hash code including a third hash code portion representing a second data portion associated with the first context type, wherein, as stored in the heterogeneous hash code database, the third hash code portion is accessible separately from another portion of the second heterogeneous hash code;

compare the first heterogeneous hash code with the second heterogeneous hash code including by determining a similarity between the second hash code portion of the first heterogeneous hash code and the third hash code portion of the second heterogeneous hash code; and

determine, based on the comparison, a similarity between at least two of a plurality of users.

14. The apparatus of claim 13 , wherein the contextual information further includes a third data portion associated with a second context type; and wherein the first heterogeneous hash code further includes a fourth hash code portion representing the third data portion associated with the second context type.

15. The apparatus of claim 14 , wherein the first context type includes at least one of a location type, a date/time type, or a demographic type, and wherein the first context type is different from the second context type.

16. The apparatus of claim 14 , wherein the second heterogeneous hash code includes a fifth hash code portion representing a fourth data portion associated with the second context type; and

wherein comparing the first heterogeneous hash code with the second heterogeneous hash code further includes determining a similarity between the fourth hash code portion of the first heterogeneous hash code and the fifth hash code portion of the second heterogeneous hash code.

17. The apparatus of claim 13 , wherein one or more processors are configured to execute the instructions to store the first heterogeneous hash code in the heterogeneous hash code database, wherein, while stored, the first hash code portion is accessible separately from the second hash code portion, and wherein the second hash code portion is stored in association with the first context type.

18. One or more non-transitory computer-readable storage media embodying instructions that are operable when executed by one or more processors to:

receive user behavior data and contextual information associated with the user behavior data, the contextual information including a first data portion associated with a first context type;

generate, from the user behavior data and the contextual information using a hashing algorithm, a first heterogeneous hash code including a first hash code portion representing the user behavior data and a second hash code portion representing the first data portion associated with the first context type, wherein generating the first heterogeneous hash code comprises:

generating, by a network layer of a machine-learning architecture:

a user-behavior representation based on the user behavior data and a weighted combination of at least (1) the user behavior data and (2) the contextual information; and

a first-context representation based on the first data portion and a weighted combination of at least (1) the first data portion and (2) the user behavior data;

embedding, by a hash layer of the machine-learning architecture, the user-behavior representation and the first-context representation;

binarizing the embedded user-behavior representation to generate the first hash code portion; and

binarizing the embedded first-context representation to generate the second hash code portion;

retrieve, from a heterogeneous hash code database a second heterogeneous hash code including a third hash code portion representing a second data portion associated with the first context type, wherein, as stored in the heterogeneous hash code database, the third hash code portion is accessible separately from another portion of the second heterogeneous hash code;

compare the first heterogeneous hash code with the second heterogeneous hash code including by determining a similarity between the second hash code portion of the first heterogeneous hash code and the third hash code portion of the second heterogeneous hash code; and

determine, based on the comparison, a similarity between at least two of a plurality of users.

19. The one or more non-transitory computer-readable storage media of claim 18 , wherein the contextual information further includes a third data portion associated with a second context type; and wherein the first heterogeneous hash code further includes a fourth hash code portion representing the third data portion associated with the second context type.

20. The one or more non-transitory computer-readable storage media of claim 19 , wherein the first context type includes at least one of a location type, a date/time type, or a demographic type, and wherein the first context type is different from the second context type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2020
From: ZHOU, PENG; ZHU, YINGNAN; ZHAO, XIANGYUAN; KIM, HONG-HOE; LEE, HYUN CHUL
To: SAMSUNG ELECTRONICS COMPANY, LTD.
Reel/Frame 051882/0615 →
Continuity (2)
Provisional Application 62814418 · Mar 6, 2019
Related Publication 20200286112A1 · Sep 10, 2020