IP Library Granted Patent US 12,639,554
Granted Patent B2
US 12,639,554 · App. 18/308,167 · Granted May 26, 2026

Combining generative aligners and transition-based parsers

Inventors: Ramon Fernandez Astudillo (White Plains, NY); Andrew Drozdov (New York, NY); Jiawei Zhou (Cambridge, MA); Radu Florian (Danbury, CT); Tahira Naseem (Briarcliff Manor, NY); Yoon Hyung Kim (Cambridge, MA)
Assignees: INTERNATIONAL BUSINESS MACHINES CORPORATION; Massachusetts Institute Of Technology
G06N3/0475G06N3/047G06N5/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,554
App. No.
18/308,167
Granted
May 26, 2026
Kind
B2
Abstract

Systems, computer-implemented methods, and computer program products to facilitate reducing error propagation when combining generative aligners and transition-based aligners are provided. According to an embodiment, a system can comprise a processor that executed components stored in memory. The computer executable components comprise a generative alignment component, an error propagation component, a discriminative parser, and a stochastic oracle policy component. The error propagation component can compute a posterior distribution over one or more hard alignments of parts given a pair of the generative alignment component. The discriminative parser can be trained via the stochastic oracle policy component to reduce error propagation when combining the generative alignment component with the discriminative parser.

Claims (43)

1 . A system, comprising:

a processor that executes computer-executable components stored in a non-transitory computer-readable memory, wherein the computer-executable components comprise:

a generative aligner approximates samples from a posterior distribution, wherein the processor computes the posterior distribution over one or more hard alignments of parts given a pair within a generative aligner; and wherein the system comprises

a discriminative parser trained via a stochastic oracle policy component to reduce error propagation when combining the generative alignment component with the discriminative parser.

2 . The system of claim 1 , wherein the stochastic oracle policy component samples the one or more hard alignments of parts from the posterior distribution and runs a deterministic oracle policy on one or more samples of the one or more hard alignments of parts.

3 . The system of claim 2 , wherein the stochastic oracle policy component performs a learning step on the discriminative parser over one or more epochs.

4 . The system of claim 3 , wherein the learning step includes learning via gradient descent over one or more epochs.

5 . The system of claim 3 , wherein the processor further:

applies a scaling factor to average over the one or more samples.

6 . The system of claim 1 , wherein the discriminative parser is exposed to 6 . one or more explanatory models.

7 . The system of claim 1 , wherein the discriminative parser and the generative alignment component are end-to-end systems.

8 . A computer implemented method for reducing error propagation when combining generative aligners and transition-based aligners, the computer implemented method comprising:

computing, by a device operatively coupled to a processor, a posterior distribution over one or more hard alignments of parts given a pair within a generative alignment; and

running, by the device, a stochastic oracle policy to train a discriminative parser, further comprising:

sampling, by the device, the one or more hard alignments of parts from the posterior distribution;

running, by the device, a deterministic oracle policy on one or more samples of the one or more hard alignments of parts;

performing, by the device, a learning step on the discriminative parser; and

applying, by the device, a scaling factor to average over the one or more samples.

9 . The computer implemented method of claim 8 , wherein an uncertainty of the one or more hard alignments of parts is incorporated into training the discriminative parser.

10 . The computer implemented method of claim 8 , further comprising:

exposing the discriminative parser to one or more explanatory models.

11 . The computer implemented method of claim 8 , wherein running the stochastic oracle policy prevents over-fitting the discriminative parser to bad alignments.

12 . The computer implemented method of claim 8 , wherein the learning step is performed via gradient descent.

13 . The computer implemented method of claim 8 , further comprising:

transferring embedding-level knowledge from the generative alignment into speech embeddings.

14 . The computer implemented method of claim 8 , wherein the pair includes a natural language portion and a formal representation portion associated with an input sentence.

15 . A computer program product for reducing error propagation when combining generative aligners and transition-based aligners, the computer program product comprising a non-transitory computer readable memory having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to:

compute a posterior distribution over one or more hard alignments of parts given a pair within a generative alignment, wherein the generative alignment is relative to approximate samples from a posterior distribution; and

run a stochastic oracle policy to train a discriminative parser,

wherein an uncertainty of alignment of the generative alignment is incorporated into training the discriminative parser.

16 . The computer program product of claim 15 , wherein the program instructions are further executable to cause the processor to:

sample the one or more hard alignments of parts from the posterior distribution;

run a deterministic oracle policy on one or more samples of the one or more hard alignments of parts;

perform a learning step on the discriminative parser; and

apply a scaling factor to average over the one or more samples.

17 . The computer program product of claim 16 , wherein the program instructions are further executable to cause the processor to:

expose the discriminative parser to one or more explanatory models.

18 . The computer program product of claim 16 , wherein the program instructions are further executable to cause the processor to:

prevent over-fitting of the discriminative parser to bad alignments via the stochastic oracle policy.

19 . The computer program product of claim 16 , wherein the program instructions are further executable to cause the processor to:

perform the learning step via gradient descent.

20 . The computer program product of claim 16 , wherein the program instructions are further executable to cause the processor to:

transfer embedding-level knowledge from the generative alignment into speech embeddings.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2023
From: FERNANDEZ ASTUDILLO, RAMON; DROZDOV, ANDREW; ZHOU, JIAWEI; FLORIAN, RADU; NASEEM, TAHIRA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 063464/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2023
From: KIM, YOON HYUNG
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 063464/0329 →
Continuity (1)
Related Publication 20240362466A1 · Oct 31, 2024
References Cited (34)
US 8874434B2 · Collobert et al. · 2014 [cited by applicant]
US 9536522B1 · Hall et al. · 2017 [cited by applicant]
US 10876847B2 · Wang · 2020 [cited by examiner]
US 10878188B2 · Zhang et al. · 2020 [cited by applicant]
US 11238095B1 · Burchard · 2022 [cited by applicant]
US 11314989B2 · Munoz Delgado · 2022 [cited by examiner]
US 11481637B2 · Malaya · 2022 [cited by examiner]
US 11551337B2 · Tagra · 2023 [cited by examiner]
US 20040243568A1 · Wang et al. · 2004 [cited by applicant]
US 20060277028A1 · Chen et al. · 2006 [cited by applicant]
US 20190188603A1 · Cirit · 2019 [cited by examiner]
US 20200082250A1 · Guan · 2020 [cited by examiner]
US 20210056429A1 · Gangeh · 2021 [cited by examiner]
US 20210073630A1 · Zhang · 2021 [cited by examiner]
US 20210103807A1 · Baker · 2021 [cited by examiner]
US 20210319240A1 · Demir · 2021 [cited by examiner]
US 20220083837A1 · Hashimoto et al. · 2022 [cited by applicant]
US 20220086174A1 · Helmsen · 2022 [cited by examiner]
US 20220114473A1 · Awasthy et al. · 2022 [cited by applicant]
US 20220245431A1 · Mehr · 2022 [cited by examiner]
US 20220253687A1 · Miyaguchi · 2022 [cited by examiner]
US 20220366264A1 · Moradi · 2022 [cited by examiner]
US 20220414430A1 · Li · 2022 [cited by examiner]
US 20230009814A1 · Hao · 2023 [cited by examiner]
US 20230128579A1 · Resnick · 2023 [cited by examiner]
US 20230281429A1 · Bandyopadhyay · 2023 [cited by examiner]
US 20230401469A1 · Papadopoulos · 2023 [cited by examiner]
US 20240362466A1 · Fernandez Astudillo · 2024 [cited by examiner]
US 20250045631A1 · Solmaz · 2025 [cited by examiner]
US 20250225401A1 · Durvasula · 2025 [cited by examiner]
US 20250278643A1 · Doshi · 2025 [cited by examiner]
US 20250322917A1 · Bengio · 2025 [cited by examiner]
Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR Parsing Jiawei Zhou Tahira Naseem RamónFernandezAstudillo Young-Suk Lee Radu Florian. [cited by examiner]
Drozdov, et al., “Inducing and Using Alignments for Transition-based AMR Parsing,” arXiv:2205.01464v1 [cs.CL] May 3, 2022. [cited by applicant]