IP Library › Granted Patent US 12,333,501
Granted Patent B2
US 12,333,501 · App. 17/935,060 · Granted Jun 17, 2025

Using unsupervised machine learning to identify attribute values as related to an input

Inventors: Liwei Wu (Davis, CA); Lichao Ni (Sunnyvale, CA); Mikaela Makalinao Guerrero (Downey, CA); Yanen Li (Hillsborough, CA)
Assignee: Microsoft Technology Licensing, LLC
G06Q10/1053
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,501
App. No.
17/935,060
Granted
Jun 17, 2025
Kind
B2
Abstract

Technologies for skill taxonomy management are described. Embodiments include extracting an input text from an online system and applying an unsupervised generative text machine learning model to the input text. The text generator generates a set of sentences based on a job title included in the input text. One or more skills are extracted from the set of sentences. The extracted one or more skills correspond to one or more skills in a skill taxonomy. A frequency distribution is generated over the extracted one or more skills. The one or more skills are ranked based on the frequency distribution. Based on the ranking, a subset of the extracted one or more skills is generated. The subset of the extracted one or more skills is provided to a downstream operation, process, or service of the online system.

Claims (72)

1. A method comprising:

extracting an input text from an online system, the input text comprising a job title;

applying an unsupervised generative text machine learning model to the input text;

generating, by the generative text machine learning model, a plurality of sentences based on the job title;

extracting one or more skills from the plurality of sentences, wherein the extracted one or more skills correspond to one or more skills in a skill taxonomy;

generating a frequency distribution over the extracted one or more skills;

ranking each skill of the extracted one or more skills based on the frequency distribution;

comparing the frequency distribution to a threshold skill distribution;

in response to determining that the frequency distribution does not satisfy the threshold skill distribution, generating additional sentences from the input text;

generating an additional frequency distribution using the plurality of sentences and the additional sentences;

comparing the additional frequency distribution to the threshold skill distribution;

in response to determining that the additional frequency distribution satisfies the threshold skill distribution, ranking each skill of the extracted one or more skills based on the additional frequency distribution;

generating a subset of the extracted one or more skills based on the ranking and a threshold number of skills; and

providing the subset of the extracted one or more skills to a downstream operation, process, or service of the online system.

2. The method of claim 1 , wherein extracting the one or more skills from the plurality of sentences comprises:

identifying one or more skills in the plurality of sentences that do not correspond to the one or more skills in the skill taxonomy; and

adding the identified one or more skills to the skill taxonomy.

3. The method of claim 1 , further comprising generating a skill recommendation based on comparing the subset of the extracted one or more skills with a set of skills identified in an entity profile.

4. The method of claim 1 , further comprising training the generative text machine learning model by applying unsupervised machine learning to a domain-independent set of unlabeled and unstructured training data.

5. The method of claim 1 , wherein extracting the input text from the online system comprises:

receiving an input seed phrase from a user system; and

determining the input text based on the input seed phrase.

6. The method of claim 1 , further comprising training the generative text machine learning model as a causal text generator that generates a set of sentences from a seed word or a job title.

7. A method comprising:

extracting, using a string search, an input text from an online system, the input text comprising a job title;

applying an unsupervised generative text machine learning model to the input text;

generating, by the generative text machine learning model, a plurality of sentences from the job title;

extracting one or more skills from the plurality of sentences, wherein the extracted one or more skills correspond to one or more skills in a skill taxonomy;

generating a frequency distribution over the extracted one or more skills, the frequency distribution generated by aggregating a number of occurrences for each skill of the extracted one or more skills;

ranking each skill of the extracted one or more skills based on the frequency distribution;

comparing the frequency distribution to a threshold skill distribution;

in response to determining that the frequency distribution does not satisfy the threshold skill distribution, generating additional sentences from the input text;

generating an additional frequency distribution using the plurality of sentences and the additional sentences;

comparing the additional frequency distribution to the threshold skill distribution;

in response to determining that the additional frequency distribution satisfies the threshold skill distribution, ranking each skill of the extracted one or more skills based on the additional frequency distribution;

generating a subset of the extracted one or more skills by selecting the subset using the ranking and a threshold number of skills, wherein the threshold number of skills define the number of skills in the subset; and

providing the subset of the extracted one or more skills to a downstream operation, process, or service of the online system.

8. The method of claim 7 , wherein extracting the one or more skills from the plurality of sentences comprises:

identifying one or more skills that do not correspond to one or more skills in the skill taxonomy; and

adding the identified one or more skills to the skill taxonomy.

9. The method of claim 7 , further comprising generating a recommended skill based on comparing the subset of the extracted one or more skills with a set of skills identified in an entity profile.

10. The method of claim 7 further comprising configuring the generative text machine learning model as a causal text generator that generates a set of sentences from a seed word or a job title.

11. The method of claim 7 , wherein extracting the input text from the online system comprises:

receiving an input seed phrase from a user system; and

determining the input text from the input seed phrase.

12. The method of claim 7 , further comprising training the generative text machine learning model by applying unsupervised machine learning to a domain-independent set of unlabeled and unstructured training data.

13. A system comprising:

at least one memory device; and

a processing device, operatively coupled to the at least one memory device, to:

extract an input text from an online system, the input text comprising a job title;

apply an unsupervised generative text machine learning model to the input text;

generate, by the unsupervised generative text machine learning model, a plurality of sentences based on the job title;

extract one or more skills from the plurality of sentences, wherein the extracted one or more skills correspond to one or more skills in a skill taxonomy;

generate a frequency distribution over the extracted one or more skills;

compare the frequency distribution to a threshold skill distribution;

in response to determining that the frequency distribution does not satisfy the threshold skill distribution, generate additional sentences from the input text;

generate an additional frequency distribution using the plurality of sentences and the additional sentences;

compare the additional frequency distribution to the threshold skill distribution; and

provide the extracted one or more skills to a downstream operation, process, or service of the online system.

14. The system of claim 13 , wherein to extract the input text from the online system, the processing device is to:

receive an input seed phrase from a user system; and

determine the input text from the input seed phrase.

15. The system of claim 13 , wherein the processing device trains the generative text machine learning model as a causal text generator that generates a set of sentences from a seed word or a job title.

16. The system of claim 13 , wherein the processing device generates a skill recommendation based on comparing the extracted one or more skills with a set of skills identified in an entity profile.

17. The system of claim 13 , wherein to extract the one or more skills from the plurality of sentences, the processing device is caused to:

identify one or more skills that do not correspond to a skill in the skill taxonomy; and

add the identified one or more skills to the skill taxonomy.

18. The system of claim 13 , wherein the processing device is caused to:

in response to determining that the additional frequency distribution satisfies the threshold skill distribution, rank each skill of the extracted one or more skills based on the additional frequency distribution; and

generate a subset of the extracted one or more skills based on the ranking.

19. The system of claim 13 , wherein the processing device is caused to train the generative text machine learning model by applying unsupervised machine learning to a domain-independent set of unlabeled and unstructured training data.

20. The system of claim 13 , wherein the processing device is caused to generate the frequency distribution by aggregating a number of occurrences for each skill of the extracted one or more skills.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: WU, LIWEI; NI, LICHAO; GUERRERO, MIKAELA MAKALINAO; LI, YANEN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061459/0983 →
Continuity (1)
Related Publication 20240104506A1 · Mar 28, 2024
References Cited (18)
US 7761320B2 · Fliess · 2010 [cited by examiner]
US 11961156B2 · Aguilar Achiaga · 2024 [cited by examiner]
US 20160092998A1 · Goel · 2016 [cited by examiner]
US 20180096306A1 · Wang · 2018 [cited by examiner]
US 20180121880A1 · Zhang · 2018 [cited by examiner]
US 20180181544A1 · Zhang · 2018 [cited by examiner]
US 20180261118A1 · Morris · 2018 [cited by examiner]
US 20190236718A1 · Rastkar · 2019 [cited by examiner]
US 20190385123A1 · Sawarkar · 2019 [cited by examiner]
US 20220327946A1 · Rushkin · 2022 [cited by examiner]
US 20220343250A1 · Tremblay · 2022 [cited by examiner]
US 20230214736A1 · Stavarache · 2023 [cited by examiner]
WO WO2022125096A1 · 2022 [cited by examiner]
Ioannis Konstantinidis et al., Knowledge-driven Unsupervised Skills Extraction for Graph-based Talent Matching, Sep. 7-9, 2022, In 12th Hellenic Conference on Artificial Intelligence (SETN 2022), Association for Computi… [cited by examiner]
“GPT-Neo 2.7B”, Retrieved from: https://web.archive.org/web/20220530040913/https://huggingface.co/EleutherAI/gpt-neo-2.7B, May 30, 2022, 5 Pages. [cited by applicant]
Gao, et al., “The Pile: An 800GB Dataset of Diverse Text for Language Modeling”, In Repository of arXiv:2101.00027v1, Dec. 31, 2020, 39 Pages. [cited by applicant]
Radford, et al., “Improving Language Understanding by Generative Pre-Training”, Retrieved from: https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf, 20… [cited by applicant]
Radford, et al., “Language Models Are Unsupervised Multitask Learners”, In Journal of OpenAI Blog, vol. 1, Issue 8, Feb. 24, 2019, 24 Pages. [cited by applicant]