IP Library › Granted Patent US 12,494,205
Granted Patent B2
US 12,494,205 · App. 18/094,284 · Granted Dec 9, 2025

Data structure for task-oriented dialog modeling

Inventors: Taylor G. Smith (Allen, TX); Tori Salido (Austin, TX); Sebastin Ajay John Jeyaseelan (Thoraipakkam, IN); Kordel K. France (Allen, TX)
Assignee: TOYOTA CONNECTED NORTH AMERICA, INC.
G10L15/26G06F40/166G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,494,205
App. No.
18/094,284
Granted
Dec 9, 2025
Kind
B2
Abstract

An example operation includes one or more of receiving utterances from a user via an input device, identifying a plurality of sentences spoken by the user from the utterances, converting the utterances into a sequence of tokens and storing the sequence of tokens within a data structure, wherein the storing comprises adding padding tokens to the data structure in between the sequence of tokens to normalize a structure among the plurality of sentences within the data structure, executing a machine learning model on the data structure with the added padding tokens to determine to make a prediction, removing the padding tokens from the data structure, and executing a natural language processing (NLP) model on the sequence of tokens within the data structure with the padding tokens removed to determine a response to the user.

Claims (40)

1 . An apparatus comprising:

a receiver configured to receive utterances from a user via an input device; and

a processor configured to

identify a plurality of sentences spoken by the user from the utterances;

convert the utterances into a sequence of tokens and storing the sequence of tokens within a data structure, wherein the storing comprises adding padding tokens to the data structure in between the sequence of tokens to normalize a structure among the plurality of sentences within the data structure;

execute a machine learning model on the data structure with the added padding tokens to determine to make a prediction;

remove the padding tokens from the data structure; and

execute a natural language processing (NLP) model on the sequence of tokens within the data structure with the padding tokens removed to determine a response to the user,

wherein a first word in an input utterance and a last word in the input utterance are identified, wherein a token that corresponds to the first word and a token that corresponds to the last word are stored in different slots of the data structure.

2 . The apparatus of claim 1 , wherein the processor is configured to store one or more padding tokens between two sub-sequences of tokens that correspond to two sentences uttered by the user during a conversation to normalize a size of the two sub-sequences of tokens within the data structure.

3 . The apparatus of claim 1 , wherein the data structure comprises a tensor with slots that hold data and the processor is configured to store each sentence as a sequence of tokens within a sequence of slots of the tensor with one or more padding tokens inserted in slots between each sequence of tokens within the tensor.

4 . The apparatus of claim 1 , wherein the processor is configured to move the sequence of tokens within the data structure to collapse the sequence of tokens after the padding tokens have been removed and execute the NLP model on the collapsed sequence of tokens in the data structure.

5 . The apparatus of claim 1 , wherein the machine learning model comprises a sequence model that receives the data structure as an input and determines whether to control a system to take an action or to request more utterances from the user.

6 . The apparatus of claim 1 , wherein the processor is configured to insert a padding token into a slot of the data structure immediately before the token that corresponds to the first word and after a token that corresponds to a last word in a previous input utterance.

7 . The apparatus of claim 1 , wherein the processor is configured to insert a padding token into a slot of the data structure immediately after the token that corresponds to the last word in the input utterance and before a token that corresponds to a first word in a next input utterance.

8 . A method comprising:

receiving utterances from a user via an input device;

identifying a plurality of sentences spoken by the user from the utterances;

converting the utterances into a sequence of tokens and storing the sequence of tokens within a data structure, wherein the storing comprises adding padding tokens to the data structure in between the sequence of tokens to normalize a structure among the plurality of sentences within the data structure;

executing a machine learning model on the data structure with the added padding tokens to determine to make a prediction;

removing the padding tokens from the data structure; and

executing a natural language processing (NLP) model on the sequence of tokens within the data structure with the padding tokens removed to determine a response to the user,

wherein a first word in an input utterance and a last word in the input utterance are identified, wherein a token that corresponds to the first word and a token that corresponds to the last word are stored in different slots of the data structure.

9 . The method of claim 8 , wherein the storing comprises inserting one or more padding tokens between two sub-sequences of tokens corresponding to two sentences uttered by the user to normalize a size of the two sub-sequences of tokens within the data structure.

10 . The method of claim 8 , wherein the data structure comprises a tensor with slots that hold data and the storing comprises storing each sentence as a sequence of tokens within a sequence of slots of the tensor with one or more padding tokens inserted in slots between each sequence of tokens within the tensor.

11 . The method of claim 8 , wherein the removing further comprises collapsing the sequence of tokens within the data structure to concatenate the sequence of tokens after the padding tokens have been removed and executing the NLP model on the collapsed sequence of tokens in the data structure.

12 . The method of claim 8 , wherein the machine learning model comprises a sequence model that receives the data structure as an input and determines whether to control a system to take an action or to request more utterances from the user.

13 . The method of claim 8 , wherein the converting comprises inserting a padding token into a slot of the data structure immediately before the token corresponding to the first word and after a token corresponding to a last word in a previous input utterance.

14 . The method of claim 8 , wherein the converting comprises inserting a padding token into a slot of the data structure immediately after the token corresponding to the last word in the input utterance and before a token corresponding to a first word in a next input utterance.

15 . A non-transitory computer-readable storage medium comprising instructions, that when read by a processor, cause the processor to perform a method comprising:

receiving utterances from a user via an input device;

identifying a plurality of sentences spoken by the user from the utterances;

converting the utterances into a sequence of tokens and storing the sequence of tokens within a data structure, wherein the storing comprises adding padding tokens to the data structure in between the sequence of tokens to normalize a structure among the plurality of sentences within the data structure;

executing a machine learning model on the data structure with the added padding tokens to determine to make a prediction;

removing the padding tokens from the data structure; and

executing a natural language processing (NLP) model on the sequence of tokens within the data structure with the padding tokens removed to determine a response to the user,

wherein a first word in an input utterance and a last word in the input utterance are identified, wherein a token that corresponds to the first word and a token that corresponds to the last word are stored in different slots of the data structure.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the storing comprises inserting one or more padding tokens between two sub-sequences of tokens corresponding to two sentences uttered by the user to normalize a size of the two sub-sequences of tokens within the data structure.

17 . The non-transitory computer-readable storage medium of claim 15 , wherein the data structure comprises a tensor with slots that hold data and the storing comprises storing each sentence as a sequence of tokens within a sequence of slots of the tensor with one or more padding tokens inserted in slots between each sequence of tokens within the tensor.

18 . The non-transitory computer-readable storage medium of claim 15 , wherein the removing further comprises moving the sequence of tokens within the data structure to collapse the sequence of tokens after the padding tokens have been removed and executing the NLP model on the collapsed sequence of tokens in the data structure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2023
From: SMITH, TAYLOR G.; SALIDO, TORI; JEYASEELAN, SEBASTIN AJAY JOHN; FRANCE, KORDEL K.
To: TOYOTA CONNECTED NORTH AMERICA, INC.
Reel/Frame 062303/0664 →
Continuity (1)
Related Publication 20240233731A1 · Jul 11, 2024
References Cited (31)
US 6144938A · Surace · 2000 [cited by examiner]
US 6757362B1 · Cooper · 2004 [cited by examiner]
US 8112275B2 · Kennewick et al. · 2012 [cited by applicant]
US 8332224B2 · Cristo et al. · 2012 [cited by applicant]
US 8386261B2 · Mellott et al. · 2013 [cited by applicant]
US 8660849B2 · Gruber · 2014 [cited by examiner]
US 8914294B2 · Fabbrizio et al. · 2014 [cited by applicant]
US 9189742B2 · London · 2015 [cited by applicant]
US 9396727B2 · Tzirkel-Hancock et al. · 2016 [cited by applicant]
US 9894312B2 · Pontual et al. · 2018 [cited by applicant]
US 9977778B1 · Perez et al. · 2018 [cited by applicant]
US 9990591B2 · Gelfenbeyn et al. · 2018 [cited by applicant]
US 10274911B2 · Uppala et al. · 2019 [cited by applicant]
US 10908883B2 · Webster et al. · 2021 [cited by applicant]
US 11556722B1 · Ben Shahar · 2023 [cited by examiner]
US 20010016814A1 · Hauenstein · 2001 [cited by examiner]
US 20040181407A1 · Trinkel · 2004 [cited by examiner]
US 20050033582A1 · Gadd · 2005 [cited by examiner]
US 20080109224A1 · Dvorak et al. · 2008 [cited by applicant]
US 20150302002A1 · Mathias · 2015 [cited by examiner]
US 20180301151A1 · Mont-Reynaud et al. · 2018 [cited by applicant]
US 20210248996A1 · Itoh · 2021 [cited by examiner]
US 20220036005A1 · de Brébisson · 2022 [cited by examiner]
US 20240161735A1 · Salamon · 2024 [cited by examiner]
US 20240412042A1 · Savinov · 2024 [cited by examiner]
EP 2359362B1 · 2013 [cited by applicant]
EP 3616085A1 · 2020 [cited by applicant]
WO 2018236332A1 · 2018 [cited by applicant]
Colombo et al., Guiding attention in Sequence-to-Sequence models for Dialogue Act prediction, Association for the Advancement of Artificial Intelligence, Feb. 26, 2020. [cited by applicant]
Vlasov et al., Dialogue Transformers, May 1, 2020. [cited by applicant]
Yang et al., Hierarchical Attention Networks for Document Classification, Proceedings of NAACL-HLt 2016, pp. 1480-1489, San Diego California, Jun. 12-17, 2016, Association for Computational Linguistics. [cited by applicant]