IP Library › Granted Patent US 12,748,932
Granted Patent B2
US 12,748,932 · App. 18/057,018 · Granted Sep 29, 2026

Compression of word embeddings for natural language processing systems

Inventors: Xihui Lin (Montreal, CA); Andrew James Mcnamara (Cambridge, CA); Kaheer Suleman (Cambridge, CA)
Assignee: Microsoft Technology Licensing, LLC
G06F40/44G06F40/126G06F40/30G06N3/044G06N3/045G06N3/0455G06N3/0495G06N3/08G06N3/09G06N5/04G06N20/00G10L13/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,932
App. No.
18/057,018
Granted
Sep 29, 2026
Kind
B2
Abstract

Described herein are systems and methods that provide a natural language processing system (NLPS) that employs compressed word embeddings. An auto-encoder that includes encoder circuitry and decoder circuitry can be used to produce the compressed word embeddings. The decoder circuitry is trained to decompress the word embeddings with reduced or minimal differences between the original uncompressed word embeddings and the corresponding decompressed word embeddings. One or more parameters of the trained decoder circuitry are transferred to the NLPS, where the NLPS is then trained using the compressed word embeddings to improve the correctness of the responses or actions determined by the NLPS.

Claims (45)

1 . A computing system, comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, cause the system to perform a set of operations, the set of operations comprising:

receiving a spoken language input via a speech-to-text application;

obtaining original uncompressed word embeddings for the spoken language input, the original uncompressed word embeddings including a vector of real numbers;

generating input text compressed word embeddings based on the uncompressed word embeddings using an encoder of a trained auto-encoder, the input text compressed word embeddings including a vector of binary numbers;

processing the input text compressed word embeddings with a neural network to generate output compressed word embeddings including a vector of binary numbers;

decompressing the output compressed word embeddings to produce one or more decompressed word embeddings using a decoder of the trained auto-encoder; and

outputting a natural language output based on the one or more decompressed word embeddings, using a text-to-speech application.

2 . A method, comprising:

obtaining a first set of one or more uncompressed word embeddings;

obtaining a first set of one or more compressed word embeddings corresponding to the first set of one or more uncompressed word embeddings from an encoder of an auto-encoder;

adjusting a set of parameters to reduce differences between a first set of decompressed word embeddings corresponding to the first set of one or more compressed word embeddings and the first set of one or more uncompressed word embeddings;

training a decoder of the auto-encoder separately from the encoder using the set of parameters; and

storing, for subsequent processing by the natural language processor, the first set of one or compressed word embeddings and the set of parameters, wherein the method further comprises, during training:

training the natural language processor using a second set of one or more uncompressed word embeddings and generating one or more parameters of the natural language processor;

obtaining a second set of one or more compressed word embeddings that correspond to the second set of one or more uncompressed word embeddings and the set of parameters of the decoder of the auto-encoder;

updating at least one of one or more parameters of the natural language processor with at least one of the set of parameters of the decoder of the auto-encoder; and

retraining the natural language processor using the second set of one or more compressed word embeddings and the updated one or more parameters of the natural language processor.

3 . The method of claim 2 , wherein the auto-encoder further comprises a function activator operably connected to the encoder.

4 . The method of claim 3 , wherein the auto-encoder comprises a multi-layer neural network with the encoder comprising a first layer, the function activator a second layer, and the decoder a third layer.

5 . The method of claim 4 , wherein the function activator comprises a non-linear activation function.

6 . The method of claim 5 , wherein the encoder comprises a first linear transformation circuit.

7 . The method of claim 6 , wherein the decoder comprises a second linear transformation circuit.

8 . The method of claim 4 , wherein the encoder comprises one or more parameters that are randomly initialized.

9 . The method of claim 4 , wherein the encoder comprises one or more parameters that are determined through a training process.

10 . The method of claim 4 , wherein the decoder comprises the set of parameters that are determined through a training process.

11 . The system of claim 1 , wherein the auto-encoder comprises a multi-layer neural network including a first layer, a second layer, and a third layer, wherein the first layer includes the encoder, the second layer is operatively coupled to the encoder and includes a function activator, the third layer includes the decoder of the auto-encoder.

12 . A method, comprising:

receiving, by a natural language processor, an input;

obtaining, by the natural language processor and from a storage device, one or more compressed word embeddings based on the input, wherein the one or more compressed word embeddings were generated by an encoder of an auto-encoder prior to receiving the input;

decompressing, by the natural language processor, the one or more compressed word embeddings to produce one or more decompressed word embeddings, wherein the storage device further comprises one or more parameters that are generated by a decoder of the auto-encoder, and wherein the natural language processor and the decoder were trained separately from the encoder; and

determining, by the natural language processor, an action to be performed in response to the input based on the one or more decompressed word embeddings, wherein the method further comprises, during training:

training the natural language processor using one or more uncompressed word embeddings and generating one or more parameters of the natural language processor;

obtaining one or more compressed word embeddings that correspond to the one or more uncompressed word embeddings and a set of parameters of the decoder of the auto-encoder;

updating at least one of one or more parameters of the natural language processor with at least one of the set of parameters of the decoder of the auto-encoder; and

retraining the natural language processor using the one or more compressed word embeddings and the updated one or more parameters of the natural language processor.

13 . The method of claim 12 , wherein the decoder was trained using the compressed word embeddings generated by the encoder to generate the one or more parameters.

14 . The method of claim 13 , wherein the decoder was trained to reduce a difference between the one or more decompressed word embeddings and original word embeddings that were processed by the encoder.

15 . The method of claim 12 , wherein the encoder is remote from the storage device.

16 . The method of claim 12 , further comprising performing the determined action.

17 . The method of claim 12 , wherein the input comprises natural language input from a user.

18 . The computing system of claim 1 , wherein the neural network is a recurrent neural network.

19 . The computing system of claim 1 , further comprising:

a mobile electronic device that includes the processor and memory, and further includes a microphone configured to receive the spoken language input and a speaker configured to output audio of the natural language output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2022
From: LIN, XIHUI; MCNAMARA, ANDREW JAMES; SULEMAN, KAHEER
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061828/0129 →
Continuity (2)
Continuation 15685929 · Aug 24, 2017
Related Publication 20230083335A1 · Mar 16, 2023
References Cited (29)
US 9306597B1 · Daniel · 2016 [cited by examiner]
US 10445356B1 · Mugan · 2019 [cited by examiner]
US 20110078099A1 · Weston · 2011 [cited by examiner]
US 20130346443A1 · Kataoka · 2013 [cited by examiner]
US 20160321541A1 · Liu · 2016 [cited by examiner]
US 20160350288A1 · Wick · 2016 [cited by examiner]
US 20170010811A1 · Sato · 2017 [cited by examiner]
US 20180268806A1 · Chun · 2018 [cited by examiner]
US 20180336183A1 · Lee · 2018 [cited by examiner]
US 20190051292A1 · Na · 2019 [cited by examiner]
US 20190294980A1 · Laukien · 2019 [cited by examiner]
US 20190311002A1 · Paulus · 2019 [cited by examiner]
US 20200202846A1 · Bapna · 2020 [cited by examiner]
US 20210011904A1 · James · 2021 [cited by examiner]
CN 105043433A · 2015 [cited by applicant]
CN 106202010A · 2016 [cited by applicant]
CN 106448660A · 2017 [cited by applicant]
CN 106997370A · 2017 [cited by applicant]
Paula Lauren;Guangzhi Qu; Guang-Bin Huang; Paul Watta; Amaury Lendasse; A low-dimensional vector representation for words using an extreme learning machine; URL: https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=7966… [cited by examiner]
Pierre Baldi; Autoencoders, Unsupervised Learning, and Deep Architectures; 2012; URL: http://proceedings.mlr.press/v27/baldi12a/baldi12a.pdf (Year: 2012). [cited by examiner]
Martin Andrews; Compressing Word Embeddings; May 16, 2016; URL: https://arxiv.org/pdf/1511.06397v1 (Year: 2016). [cited by examiner]
Dayiheng Liu; Jiancheng Lv; Xiaofeng Qi; Jiangshu Wei; A neural words encoding model; Jul. 2016; URL: https://ieeexplore.ieee.org/document/7727245 (Year: 2016). [cited by examiner]
Y. Adi; E. Kermany; Y. Belinkov; O. Lavi; Y. Goldberg; Analysis of sentence embedding models using prediction tasks in natural language processing; Sep. 2017; URL: https://ieeexplore.ieee.org/abstract/document/8030297 (… [cited by examiner]
“Second Office Action and Search Report Issued in Chinese Patent Application No. 201880053393.9”, Mailed Date: Jul. 25, 2023, 11 Pages. [cited by applicant]
Xuanru, et al., “The Promotion Effects of Artificial Intelligence on e-Science”, in the Journal of Technology and Application of Scientific Research Informatization, Issue 6, Nov. 20, 2016, 14 Pages. [cited by applicant]
Notice to Grant Received for Chinese Application No. 201880053393.9, mailed on Jan. 4, 2024, 3 pages. [cited by applicant]
Communication under Rule 71(3) Received for European Application No. 18740062.7, mailed on Nov. 17, 2021, 7 pages. [cited by applicant]
Decision to Grant Received for European Application No. 18740062.7, mailed on Feb. 17, 2022, 2 pages. [cited by applicant]
Tissier, et al., “Near-lossless Binarization of Word Embeddings”, in Repository of arXiv:1803.09065v3, Nov. 15, 2018, 08 Pages. [cited by applicant]