IP Library Granted Patent US 12,547,829
Granted Patent B2
US 12,547,829 · App. 17/581,515 · Granted Feb 10, 2026

Extended vocabulary including similarity-weighted vector representations

Inventors: Miquel Angel Farre Guiu (Bern, CH); Marc Junyent Martin (Barcelona, ES); Marcel Porta Valles (Balaguer, ES); Pablo Pernias (Sant Joan d'Alacant, ES); Francesc Josep Guitart Bravo (Lleida, ES); Christopher C. Stoafer (Seattle, WA); Mara Idai Lucien (Los Angeles, CA)
Assignee: Disney Enterprises, Inc.
G06F40/279G06F16/3344G06F40/169G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,829
App. No.
17/581,515
Granted
Feb 10, 2026
Kind
B2
Abstract

According to one implementation, a system includes a computing platform having processing hardware, and a system memory storing a software code. The processing hardware is configured to execute the software code to receive a vocabulary, identify words from the vocabulary for use in extending the vocabulary, pair each of those words with every other of those words to provide word pairs, and output the word pairs to a vocabulary administrator. The software code also receives word pair characterizations identifying each of the word pairs as one of similar, dissimilar, or neither similar nor dissimilar, configures, based on the word pair characterizations, a multi-dimensional vector space including multiple embedding vectors each corresponding respectively to one of the identified words, and cross-references each of those words with its corresponding embedding vector to produce an extended vocabulary corresponding to the received vocabulary.

Claims (45)

1 . A system comprising:

a computing platform including processing hardware and a system memory storing a software code;

the processing hardware configured to execute the software code to:

receive a vocabulary including a first plurality of words;

identify, from among the first plurality of words, a second plurality of words for extending the received vocabulary;

pair each word of the second plurality of words with every other word included among the second plurality of words to provide a plurality of word pairs;

output the plurality of word pairs to a human vocabulary administrator;

receive, from the human vocabulary administrator, a plurality of word pair characterizations each tagged to a respective one of the plurality of word pairs by the human vocabulary administrator, the plurality of word pair characterizations identifying each of the plurality of word pairs as one of similar, dissimilar, or neither similar nor dissimilar;

configure, based on the plurality of word pair characterizations, a multi-dimensional vector space including a plurality of embedding vectors each corresponding respectively to one of the second plurality of words, wherein all embedding vectors corresponding respectively to words included in a word pair characterized as neither similar nor dissimilar are orthogonal to one another in the multi-dimensional vector space and serve as basis vectors spanning the multi-dimensional vector space; and

cross-reference each of the second plurality of words with its corresponding embedding vector to produce an extended vocabulary corresponding to the received vocabulary.

2 . The system of claim 1 , wherein the plurality of embedding vectors provides basis vectors of the multi-dimensional vector space.

3 . The system of claim 1 , wherein the system memory further stores a machine learning (ML) model, and wherein the processing hardware is further configured to execute the software code to:

train, using the plurality of word pair characterizations, the ML model.

4 . The system of claim 3 , wherein the multi-dimensional vector space is configured using the trained ML model.

5 . The system of claim 1 , wherein the multi-dimensional vector space is configured using at least one of a triplet loss function or a cosine similarity loss function.

6 . The system of claim 1 , wherein the received vocabulary is included in a taxonomy having a hierarchical structure, and wherein the second plurality of words is identified based on a respective position of each word of the second plurality of words within the hierarchical structure.

7 . The system of claim 6 , wherein the taxonomy comprises a taxonomy of annotation tags configured for application to at least one of products, services, or digital media content.

8 . The system of claim 1 , further comprising a recommendation engine, wherein the processing hardware is further configured to execute the software code to provide the extended vocabulary as an input to the recommendation engine.

9 . The system of claim 8 , wherein the processing hardware is further configured to execute the recommendation engine to:

receive search data from a user system;

determine, using the extended vocabulary and the search data, a recommendation for a user of the user system; and

output the recommendation to the user system.

10 . The system of claim 9 , wherein the recommendation for the user identifies media content in the form of at least one of a movie, television (TV) content, a sports event, news, or a video game.

11 . A method for use by a system including a computing platform having processing hardware and a system memory storing a software code, the method comprising:

receiving, by the software code executed by the processing hardware, a vocabulary including a first plurality of words;

identifying, from among the first plurality of words by the software code executed by the processing hardware, a second plurality of words for extending the received vocabulary;

pairing, by the software code executed by the processing hardware, each word of the second plurality of words with every other word included among the second plurality of words to provide a plurality of word pairs;

outputting, by the software code executed by the processing hardware the plurality of word pairs to a human vocabulary administrator;

receiving from the human vocabulary administrator, by the software code executed by the processing hardware, a plurality of word pair characterizations each tagged to a respective one of the plurality of word pairs by the human vocabulary administrator, the plurality of word pair characterizations identifying each of the plurality of word pairs as one of similar, dissimilar, or neither similar nor dissimilar;

configuring, by the software code executed by the processing hardware based on the plurality of word pair characterizations, a multi-dimensional vector space including a plurality of embedding vectors each corresponding respectively to one of the second plurality of words, wherein all embedding vectors corresponding respectively to words included in a word pair characterized as neither similar nor dissimilar are orthogonal to one another in the multi-dimensional vector space and serve as basis vectors spanning the multi-dimensional vector space; and

cross-referencing, by the software code executed by the processing hardware, each of the second plurality of words with its corresponding embedding vector to produce an extended vocabulary corresponding to the received vocabulary.

12 . The method of claim 11 , wherein the plurality of embedding vectors provides the basis vectors of the multi-dimensional vector space.

13 . The method of claim 11 , wherein the system memory further stores a machine learning (ML) model, the method further comprising:

training, by the software code executed by the processing hardware and using the plurality of word pair characterizations, the ML model.

14 . The method of claim 13 , wherein configuring the multi-dimensional vector space is performed using the trained ML model.

15 . The method of claim 11 , wherein configuring the multi-dimensional vector space is performed using at least one of a triplet loss function or a cosine similarity loss function.

16 . The method of claim 11 , wherein the received vocabulary is included in a taxonomy having a hierarchical structure, and wherein the second plurality of words is identified based on a respective position of each word of the second plurality of words within the hierarchical structure.

17 . The method of claim 16 , wherein the taxonomy comprises a taxonomy of annotation tags configured for application to at least one of products, services, or digital media content.

18 . The method of claim 11 , wherein the system further comprises a recommendation engine, the method further comprising:

providing, by the software code executed by the processing hardware, the extended vocabulary as an input to the recommendation engine.

19 . The method of claim 18 , further comprising:

receiving, by the recommendation engine executed by the processing hardware, search data from a user system;

determining, by the recommendation engine executed by the processing hardware and using the extended vocabulary and the search data, a recommendation for a user of the user system; and

outputting, by the recommendation engine executed by the processing hardware, the recommendation to the user system.

20 . The method of claim 19 , wherein the recommendation for the user identifies media content in the form of at least one of a movie, television (TV) content, a sports event, news, or a video game.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2022
From: FARRE GUIU, MIQUEL ANGEL; MARTIN, MARC JUNYENT; VALLES, MARCEL PORTA; PERNIAS, PABLO; GUITART BRAVO, FRANCESC JOSEP
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 058729/0446 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2022
From: STOAFER, CHRISTOPHER C.; LUCIEN, MARA IDAI
To: DISNEY ENTERPRISES, INC.
Reel/Frame 058729/0532 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2022
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 058813/0557 →
Continuity (1)
Related Publication 20230237261A1 · Jul 27, 2023
References Cited (15)
US 6356864B1 · Foltz · 2002 [cited by examiner]
US 8606815B2 · Chen et al. · 2013 [cited by applicant]
US 9916381B2 · Antonelli et al. · 2018 [cited by applicant]
US 10019442B2 · Nefedov et al. · 2018 [cited by applicant]
US 11113308B1 · Singh et al. · 2021 [cited by applicant]
US 20090327336A1 · King et al. · 2009 [cited by applicant]
US 20110251839A1 · Achtermann · 2011 [cited by examiner]
US 20150220618A1 · Horesh et al. · 2015 [cited by applicant]
US 20170330363A1 · Song · 2017 [cited by examiner]
US 20190005049A1 · Mittal · 2019 [cited by examiner]
US 20190303465A1 · Shanmugamani · 2019 [cited by examiner]
US 20210004534A1 · Mizushima · 2021 [cited by examiner]
US 20210074171A1 · Agley · 2021 [cited by examiner]
US 20210110432A1 · Chen · 2021 [cited by examiner]
US 20210150346A1 · Bertinetto et al. · 2021 [cited by applicant]