User-specific text record-based format prediction
A method for training a machine learning model included generating training data for the machine learning model. Generating the training data includes generating first training input that includes candidate text portions of one or more electronic documents and generating a first target output for the first training input. The first target output identifies a formatting type for each of the candidate text portions. The training data is provided to train the machine learning model on (i) a set of training inputs including the first training input, and (ii) a set of target outputs including the first target output.
1 . A method for training a machine learning model comprising:
generating training data for the machine learning model, wherein generating the training data comprises:
generating first training input comprising candidate text portions of one or more electronic documents, wherein the candidate text portions correspond to user input to the one or more electronic documents; and
generating a first target output for the first training input, wherein the first target output (a) identifies a suggested formatting type for each of the candidate text portions, wherein the suggested formatting type is different from an original formatting type of a respective candidate text portion, and (b) identifies an indication whether the suggested formatting type was accepted by a respective user; and
providing the training data to train the machine learning model on (i) a set of training inputs comprising the first training input, and (ii) a set of target outputs comprising the first target output.
2 . The method of claim 1 , further comprising:
generating second training input comprising contextual information associated with the candidate text portions, wherein the first target output is generated for the first training input and the second training input and identifies the suggested formatting type for each of the candidate text portions in accordance with respective contextual information, and
wherein the set of training inputs comprises the second training input.
3 . The method of claim 2 , wherein the contextual information comprises text located before the candidate text portions in respective electronic documents of the one or more electronic documents.
4 . The method of claim 3 , wherein the contextual information comprises text located after the candidate text portions in the respective electronic documents of the one or more electronic documents.
5 . The method of claim 1 , wherein each training input of the set of training inputs is mapped to the first target output in the set of target outputs.
6 . A method, comprising:
receiving an indication of user input of a candidate text portion into an electronic document;
subsequent to receiving the indication of user input of the candidate text portion into the electronic document, identifying metadata associated with a word of a stored text record that corresponds to a word of the candidate text portion, wherein the word of the candidate text portion is a candidate for applying a formatting suggestion that suggests a change to a format of at least part of the candidate text portion of the electronic document;
annotating the candidate text portion with at least part of the metadata;
providing, to a trained machine learning model, first input comprising the annotated candidate text portion that includes the candidate text portion annotated with the at least part of the metadata; and
obtaining, from the trained machine learning model, one or more outputs identifying (i) a format identifier indicative of a suggested formatting type, and (ii) a level of confidence that the suggested formatting type is appropriate for the candidate text portion.
7 . The method of claim 6 , further comprising:
determining whether the level of confidence that the suggested formatting type is appropriate for the candidate text portion satisfies a threshold level; and
responsive to determining that the level of confidence that the suggested formatting type is appropriate for the candidate text portion satisfies the threshold level of confidence, generating a notification to a user of the electronic document of a formatting suggestion according to the suggested formatting type.
8 . The method of claim 6 , wherein the metadata indicates a heading level associated with the word of the stored text record.
9 . The method of claim 6 , wherein annotating the candidate text portion with the at least part of the metadata, comprises:
identifying a count associated with the word of the stored text record, wherein the count is indicative of a number of occurrences of the word in a plurality of stored text records; and
determining whether the count associated with the word of the stored text record satisfies a threshold number, wherein the candidate text portion is annotated with the at least part of the metadata associated with the word of the stored text record responsive to determining that the count satisfies the threshold number.
10 . The method of claim 6 , further comprising:
providing, to the trained machine learning model, second input comprising a remaining text portion of the electronic document.
11 . The method of claim 6 , further comprising:
identifying, among a plurality of stored text records, the stored text record that corresponds to the candidate text portion, wherein the plurality of stored text records comprises text regions that previously had been determined to satisfy have satisfied at least one of a plurality of predetermined patterns and comprises words of additional candidate text portions for which respective formatting suggestions were accepted by a user associated with the electronic document.
12 . The method of claim 11 , wherein the plurality of stored text records comprises text of one or more electronic documents edited by the user.
13 . A system, comprising:
a memory; and
a processing device, coupled to the memory, to perform operations comprising:
generating training data for a machine learning model, wherein generating the training data comprises:
generating first training input comprising candidate text portions of one or more electronic documents, wherein the candidate text portions correspond to user input to the one or more electronic documents; and
generating a first target output for the first training input, wherein the first target output (a) identifies a suggested formatting type for each of the candidate text portions, wherein the suggested formatting type is different from an original formatting type of a respective candidate text portion, and (b) identifies an indication whether the suggested formatting type was accepted by a respective user; and
providing the training data to train the machine learning model on (i) a set of training inputs comprising the first training input, and (ii) a set of target outputs comprising the first target output.
14 . The system of claim 13 , the operations further comprising:
generating second training input comprising contextual information associated with the candidate text portions, wherein the first target output is generated for the first training input and the second training input and identifies the suggested formatting type for each of the candidate text portions in accordance with respective contextual information, and
wherein the set of training inputs comprises the second training input.
15 . The system of claim 14 , wherein the contextual information comprises text located before the candidate text portions in respective electronic documents of the one or more electronic documents.
16 . The system of claim 15 , wherein the contextual information comprises text located after the candidate text portions in the respective electronic documents of the one or more electronic documents.
17 . The system of claim 13 , wherein each training input of the set of training inputs is mapped to the first target output in the set of target outputs.
18 . A system, comprising:
a memory; and
a processing device, coupled to the memory, to perform operations comprising:
receiving an indication of user input of a candidate text portion into an electronic document;
subsequent to receiving the indication of user input of the candidate text portion into the electronic document, identifying metadata associated with a word of a stored text record that corresponds to a word of the candidate text portion, wherein the word of the candidate text portion is a candidate for applying a formatting suggestion that suggests a change to a format of at least part of the candidate text portion of the electronic document;
annotating the candidate text portion with at least part of the metadata;
providing, to a trained machine learning model, first input comprising the annotated candidate text portion that includes the candidate text portion annotated with the at least part of the metadata; and
obtaining, from the trained machine learning model, one or more outputs identifying (i) a format identifier indicative of a suggested formatting type, and (ii) a level of confidence that the suggested formatting type is appropriate for the candidate text portion.
19 . The system of claim 18 , the operations further comprising:
determining whether the level of confidence that the suggested formatting type is appropriate for the candidate text portion satisfies a threshold level; and
responsive to determining that the level of confidence that the suggested formatting type is appropriate for the candidate text portion satisfies the threshold level of confidence, generating a notification to a user of the electronic document of a formatting suggestion according to the suggested formatting type.
20 . The system of claim 18 , wherein annotating the candidate text portion with the at least part of the metadata, comprises:
identifying a count associated with the word of the stored text record, wherein the count is indicative of a number of occurrences of the word in a plurality of stored text records; and
determining whether the count associated with the word of the stored text record satisfies a threshold number, wherein the candidate text portion is annotated with the at least part of the metadata associated with the word of the stored text record responsive to determining that the count satisfies the threshold number.