IP Library › Granted Patent US 12,204,856
Granted Patent B1
US 12,204,856 · App. 17/482,548 · Granted Jan 21, 2025

Training and domain adaptation for supervised text segmentation

Inventors: Swapna Somasundaran (Plainsboro, NJ); Goran Glavaš (Heidelberg, DE)
Assignee: Educational Testing Service
G06F40/284G06F18/214G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,204,856
App. No.
17/482,548
Granted
Jan 21, 2025
Kind
B1
Abstract

Data such as unstructured text is received that includes a sequence of sentences. This received data is then tokenized into a plurality of tokens. The received data is segmented using a hierarchical transformer network model including a token transformer, a sentence transformer, and a segmentation classifier. The token transformer contextualizes tokens within sentences and yields sentence embeddings. The sentences transformer contextualizes sentence representations based on the sentence embedddings. The segmentation classifier predicts segments of the received data based on the contextualized sentence representations. Data can be provided which characterizes the segmentation of the received data. Related apparatus, systems, techniques and articles are also described.

Claims (49)

1. A computer-implemented method for text segmentation, comprising:

receiving an unstructured dataset built from a plurality of educational reading materials, wherein the unstructured dataset comprises a sequence of sentences;

tokenizing the unstructured dataset into a plurality of tokens;

contextualizing the plurality of tokens within the sequence of sentences by a token transformer and yielding sentence embeddings, wherein at least the token transformer was previously fine-tuned on a structured dataset not built from the plurality of educational reading materials;

contextualizing sentence representations by a sentence transformer based on the sentence embeddings, wherein the sentence transformer and the token transformer are hierarchically linked;

predicting at least one segment of the unstructured dataset by a segmentation classifier based on the contextualized sentence representations; and

providing data characterizing the at least one predicted segment of the unstructured dataset.

2. The method of claim 1 , wherein the providing data comprises at least one of: causing the provided data to be displayed in an electronic visual display, transmitting the provided data to a remote computing device, storing the provided data in physical persistence, or loading the provided data into memory.

3. The method of claim 1 , wherein the token transformer comprises a plurality of layers, wherein at least one layer of the plurality comprises a multi-head attention sublayer and a feed-forward sublayer.

4. The method of claim 3 , wherein the at least one layer of the plurality comprises a first bottleneck adapter positioned after the multi-head attention sublayer and a second bottleneck adapter positioned after the feed-forward sublayer.

5. The method of claim 4 , wherein the token transformer comprises a plurality of original parameters that are initialized with weights of a pretrained RoBERTa model.

6. The method of claim 5 , wherein the pretrained RoBERTa weights encode general-purpose distributional knowledge.

7. The method of claim 6 , wherein the plurality of original parameters remains unchanged throughout adapter-based fine-tuning, such that the distributional knowledge is preserved.

8. The method of claim 7 , wherein the at least one layer of the plurality is augmented with adapter parameters that are updated during adapter-based fine-tuning.

9. The method of claim 8 , wherein both the plurality of original parameters and the adapter parameters are configured to remain unchanged after the unstructured dataset is received.

10. The method of claim 1 , wherein the structured dataset comprises a segment-annotated sequence of sentences.

11. The method of claim 10 , wherein the annotated dataset is segmented by a heading structure.

12. The method of claim 1 , wherein the sentence transformer is randomly initiated.

13. The method of claim 1 , wherein the unstructured dataset is synthetically created at least by selecting and concatenating two samples from a same book, different books from a same content area, different books from different content areas, or a combination thereof.

14. The method of claim 1 , further comprising sliding a window over at least one sentence of the sequence to generate a plurality of probabilities that the at least one sentence starts a new segment.

15. The method of claim 14 , wherein the plurality of probabilities is averaged, and the segmentation classifier is configured to predict that the at least one sentence starts a new segment if the average is above a predetermined threshold.

16. A system for text segmentation, comprising:

at least one data processor; and

a memory including instructions which, when executed by the at least one data processor, result in operations comprising:

receiving an unstructured dataset built from a plurality of educational reading materials, wherein the unstructured dataset comprises a sequence of sentences;

tokenizing the unstructured dataset into a plurality of tokens;

contextualizing the plurality of tokens within the sequence of sentences by a token transformer and yielding sentence embeddings, wherein at least the token transformer was previously fine-tuned on a structured dataset not built from the plurality of educational reading materials;

contextualizing sentence representations by a sentence transformer based on the sentence embeddings, wherein the sentence transformer and the token transformer are hierarchically linked;

predicting at least one segment of the unstructured dataset by a segmentation classifier based on the contextualized sentence representations; and

providing data characterizing the at least one segment of the unstructured dataset.

17. A computer-implemented method of training a neural network for text segmentation, comprising:

receiving a first training dataset from a non-educational domain, the first dataset comprising a sequence of sentences;

training a hierarchical transformer network model out-of-domain, comprising:

tokenizing the first dataset into a plurality of tokens;

initializing weights of parameters of a token transformer;

contextualizing the plurality of tokens within the sequence of sentences by the token transformer and yielding a plurality of sentence embeddings;

contextualizing a plurality of sentence representations by a sentence transformer based on the plurality of sentence embeddings;

predicting, by a segmentation classifier, whether a sentence in the sequence starts a segment based on the plurality of contextualized sentence representations;

updating at least one adapter parameter of the token transformer;

receiving a second training dataset from an educational domain, the second dataset comprising a second sequence of sentences;

training the hierarchical transformer network model in-domain, comprising:

tokenizing the second dataset into a second plurality of tokens;

initializing the token transformer with the at least one updated adapter parameter;

contextualizing the second plurality of tokens within the second sequence of sentences by the token transformer and yielding a second plurality of sentence embeddings;

contextualizing a second plurality of sentence representations by the sentence transformer based on the second plurality of sentence embeddings; and

predicting, by the segmentation classifier, whether a second sentence in the second sequence starts a second segment based on the second plurality of contextualized sentence representations.

18. The method of claim 17 , wherein the weights of parameters of the token transformer comprise of a plurality of RoBERTa parameters.

19. The method of claim 18 , wherein the plurality of RoBERTa parameters are frozen throughout the in-domain training.

20. The method of claim 17 , wherein the educational domain is designed for grade-1 to college-level student populations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2021
From: SOMASUNDARAN, SWAPNA; GLAVAS, GORAN
To: EDUCATIONAL TESTING SERVICE
Reel/Frame 057571/0298 →
Continuity (1)
Provisional Application 63089724 · Oct 9, 2020
References Cited (35)
US 11862146B2 · Han · 2024 [cited by examiner]
US 11868895B2 · Huang · 2024 [cited by examiner]
US 20210183484A1 · Shaib · 2021 [cited by examiner]
Xing et al. “Improving Context Modeling in Neural Topic Segmentation”. arXiv:2010.03138v1 [cs.CL] Oct. 7, 2020 (Year: 2020). [cited by examiner]
Lukasik et al. “Text Segmentation by Cross Segment Attention”. arXiv:2004.14535v1 [cs.CL] Apr. 30, 2020 (Year: 2020). [cited by examiner]
Angheluta, Roxana, De Busser, Rik, Moens, Marie-Francine; The Use of Topic Segmentation for Automatic Summarization; Proceedings of the ACL-2002 Workshop on Automatic Summarization; 2002. [cited by applicant]
Bayomi, Mostafa, Lawless, Seamus; C-HTS: A Concept-Based Hierarchical Text Segmentation Approach; Proceedings of the 11th International Conference on Language Resources and Evaluation; Miyazaki, Japan; pp. 1519-1528; Ma… [cited by applicant]
Beeferman, Doug, Berger, Adam, Lafferty, John; Statistical Models for Text Segmentation; Machine Learning, 34(1-3); pp. 177-210; 1999. [cited by applicant]
Bokaei, Mohammad Hadi, Sameti, Hossein, Liu, Yang; Extractive Summarization of Multi-Party Meetings Through Discourse Segmentation; Natural Language Engineering, 22(1); pp. 41-72; 2016. [cited by applicant]
Brants, Thorsten, CHEN, Francine, Tsochantaridis, Ioannis; Topic-Based Document Segmentation with Probabilistic Latent Semantic Analysis; Proceedings of the 11th International Conference on Information and Knowledge Man… [cited by applicant]
Chen, Harr, Branavan, S.R.K., Barzilay, Regina, Karger, David; Global Models of Document Structure Using Latent Permutations; Proceedings of the Human Language Technologies: The 2009 Annual Conference of the North Ameri… [cited by applicant]
Choi, Freddy; Advances in Domain Independent Linear Text Segmentation; Proceedings of the 1st North American Chapter of the Association for Computational Linguistics; pp. 26-33; Apr. 2000. [cited by applicant]
Du, Lan, Buntine, Wray, Johnson, Mark; Topic Segmentation with a Structured Topic Model; Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn… [cited by applicant]
Eisenstein, Jacob; Hierarchical Text Segmentation from Multi-Scale Lexical Cohesion; Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog… [cited by applicant]
Fragkou, Pavlina, Petridis, V., Kehagias, Athanasios; A Dynamic Programming Algorithm for Linear Text Segmentation; Journal of Intelligent Information Systems, 23(2); pp. 179-197; 2004. [cited by applicant]
Glavas, Goran, Nanni, Federico, Ponzetto, Simone Paolo; Unsupervised Text Segmentation Using Semantic Relatedness Graphs; Proceedings of the 5th Joint Conference on Lexical and Computational Semantics; Berlin, Germany; … [cited by applicant]
Glavas, Goran, Somasundaran, Swapna; Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text Segmentation; Proceedings of the 34th AAAI Conference on Artificial Intelligence; pp. 7797-7804; 2020. [cited by applicant]
Hearst, Marti; Multi-Paragraph Segmentation of Expository Text; Proceedings of the 32nd Annual Meeting on Association for Computational Linguistics; pp. 9-16; Jun. 1994. [cited by applicant]
Hendrycks, Dan, Gimpel, Kevin; Gaussian Error Linear Units (GELUs); arXiv:1606.08415; 2016. [cited by applicant]
Houlsby, Neil, Giurgiu, Andrei, Jastrzebski, Stanislaw, Morrone, Bruna, de Laroussilhe, Quentin, Gesmundo, Andrea, Attariyan, Mona, Gelly, Sylvain; Parameter-Efficient Transfer Learning for NLP; International Conference… [cited by applicant]
Huang, Xiangji, Peng, Fuchun, Schuurmans, Dale, Cercone, Nick, Robertson, Stephen; Applying Machine Learning to Text Segmentation for Information Retrieval; Information Retrieval, 6(3-4); pp. 333-362; 2003. [cited by applicant]
Kingma, Diederik, BA, Jimmy Lei; ADAM: A Method for Stochastic Optimization; ICLR; 2015. [cited by applicant]
Koshorek, Omri, Cohen, Adir, Mor, Noam, Rotman, Michael, Berant, Jonathan; Text Segmentation as a Supervised Learning Task; Proceedings of the 2018 Conference of the North American Chapter of the Association for Computa… [cited by applicant]
Li, Jing, Chiu, Billy, Shang, Shuo, Shao, Ling; Neural Text Segmentation and Its Application to Sentiment Analysis; IEEE Transactions on Knowledge and Data Engineering; Mar. 2020. [cited by applicant]
Liu, Yinhan, Ott, Myle, Goyal, Naman, Du, Jingfei, Joshi, Mandar, Chen, Danqi, Levy, Omer, Lewis, Mike, Zettlemoyer, Luke, Stoyanov, Veselin; ROBERTa: A Robustly Optimized BERT Pretraining Approach; arXiv:1907.11692; Ju… [cited by applicant]
Misra, Hemant, Yvon, Francois, Jose, Joemon, Cappe, Olivier; Text Segmentation Via Topic Modeling: An Analytical Study; Proceedings of the 18th ACM Conference on Information and Knowledge Management; pp. 1553-1556; Nov.… [cited by applicant]
Pfeiffer, Jonas, Vulic, Ivan, Gurevych, Iryna, Ruder, Sebastian; MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer; arXiv:2005.00052; 2020. [cited by applicant]
Prince, Violaine, Labadie, Alexandre; Text Segmentation Based on Document Understanding for Information Retrieval; International Conference on Application of Natural Language to Information Systems; pp. 295-304; Jun. 20… [cited by applicant]
Rebuffi, Sylvestre-Alvise, Bilen, Hakan, Vedaldi, Andrea; Efficient Parametrization of Multi-Domain Deep Neural Networks; Computer Vision and Pattern Recognition, arXiv:1803.10082; 2018. [cited by applicant]
Riedl, Martin, Biemann, Chris; TopicTiling: A Text Segmentation Algorithm Based on LDA; Proceedings of the 2012 Student Research Workshop; Jeju, Republic of Korea; pp. 37-42; Jul. 2012. [cited by applicant]
Ruckle, Andreas, Pfeiffer, Jonas, Gurevych, Iryna; MultiCQA: Zero-Shot Transfer of Self-Supervised Test Matching Models on a Massive Scale; Proceedings of the 2020 Conference on Empirical Methods in Natural Language Pro… [cited by applicant]
Shtekh, Gennady, Kazakova, Polina, Nikitinsky, Nikita, Skachkov, Nikolay; Exploring Influence of Topic Segmentation on Information Retrieval Quality; International Conference on Internet Science; pp. 131-140; 2018. [cited by applicant]
Utiyama, Masao, Isahara, Hitoshi; A Statistical Model for Domain-Independent Text Segmentation; Proceedings of the 39th Annual Meeting of the Association for Computational Linguistics; pp. 499-506; Jul. 2001. [cited by applicant]
Vaswani, Ashish, Shazeer, Noam, Parmar, Niki, Uszkoreit, Jakob, Jones, Llion, Gomez, Aidan, Kaiser, Lukasz, Polosukhin, Illia; Attention Is All You Need; 31st Conference on Neural Information Processing Systems; Long Be… [cited by applicant]
Xia, Huosong, Tao, Min; Wang, Yi; Sentiment Text Classification of Customers Reviews on the Web Based on SVM; Sixth International Conference on Natural Computation; Yantai, China; pp. 3633-3637; Sep. 2010. [cited by applicant]
Cited By (2)
US 12,437,572 US 12,488,195