IP Library Granted Patent US 11,392,791
Granted Patent B2
US 11,392,791 · App. 16/556,225 · Granted Jul 19, 2022

Generating training data for natural language processing

Inventor: Waseem Alshikh (San Francisco, CA)
Assignee: Writer, Inc.
G06K9/6256G06F16/784G06F40/30G06N20/00G06V40/176G06V40/179G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,392,791
App. No.
16/556,225
Granted
Jul 19, 2022
Kind
B2
Abstract

A training data system enables the generation of training data based on video content received from one or more outside video sources. For example, the generated training data can include a transcript of a word or phrase alongside emotion, language style, and brand perception data associated with that word or phrase. To generate the training data from a video, the subtitles, video frame, metadata, and audio levels of the video can be analyzed by the training data system. The generated training data (potentially from a plurality of videos) can then be grouped into a set of training data and used to train machine learning modules for Natural Language Processing (NLP) techniques.

Claims (38)

1. A method comprising:

receiving, at a training data system, a video file comprising one or more timed subtitles and a plurality of video frames, each timed subtitle associated with a subset of the video frames of the plurality of video frames;

generating, by the training data system, a training data point for a first timed subtitle of the video file by:

performing facial recognition analysis on the subset of video frames associated with the first timed subtitle; and

determining emotion data for the training data point based on the subset of the video frames associated with the first timed subtitle and the facial recognition analysis; and

storing, by the training data system, the training data point.

2. The method of claim 1 , wherein generating a training data point for a first timed subtitle further comprises:

determining, by the training data system, language style data for the training data point based on the first timed subtitle and metadata of the video file, the language style data describing the tone or formality of the first timed subtitle.

3. The method of claim 2 , wherein generating a training data point for a first timed subtitle further comprises:

determining, by the training data system, brand perception data for the training data point based on a weighted combination of the emotion data and the language style data associated with the training data point.

4. The method of claim 1 , wherein receiving a video file comprising one or more timed subtitles and a plurality of video frames comprises:

receiving, at the training data system, a set of video files, each video file comprising one or more timed subtitles and a plurality of video frames.

5. The method of claim 1 , wherein receiving a video file comprising one or more timed subtitles and a plurality of video frames further comprises:

filtering, at the training data system, the set of video files based on metadata associated with each video file; and

associating one or more video files of the set of video files with a NLP training set.

6. The method of claim 1 , further comprising associating the training data point with a NLP training set.

7. The method of claim 6 , further comprising transmitting the NLP training set to an outside system, the outside system configured to train a machine learning model based on the NLP training set.

8. The method of claim 1 , wherein the training data point is associated with a target language.

9. The method of claim 1 , wherein determining emotion data for the training data point comprises performing emotion recognition analysis to associate faces recognized by the facial recognition analysis with one or more emotions.

10. A non-transitory computer readable storage medium comprising instructions which, when executed by a processor, cause the processor to perform the steps of:

receiving, at a training data system, a video file comprising one or more timed subtitles and a plurality of video frames, each timed subtitle associated with a subset of the video frames of the plurality of video frames;

generating, by the training data system, a training data point for a first timed subtitle of the video file by:

performing facial recognition analysis on the subset of video frames associated with the first timed subtitle; and

determining emotion data for the training data point based on the subset of the video frames associated with the first timed subtitle and the facial recognition analysis; and

storing, by the training data system, the training data point.

11. The non-transitory computer readable storage medium of claim 10 , wherein generating a training data point for a first timed subtitle further comprises:

determining, by the training data system, language style data for the training data point based on the first timed subtitle and metadata of the video file, the language style data describing the tone or formality of the first timed subtitle.

12. The non-transitory computer readable storage medium of claim 11 , wherein generating a training data point for a first timed subtitle further comprises:

determining, by the training data system, brand perception data for the training data point based on a weighted combination of the emotion data and the language style data associated with the training data point.

13. The non-transitory computer readable storage medium of claim 10 , wherein receiving a video file comprising one or more timed subtitles and a plurality of video frames comprises:

receiving, at the training data system, a set of video files, each video file comprising one or more timed subtitles and a plurality of video frames.

14. The non-transitory computer readable storage medium of claim 10 , wherein receiving a video file comprising one or more timed subtitles and a plurality of video frames further comprises:

filtering, at the training data system, the set of video files based on metadata associated with each video file; and

associating one or more video files of the set of video files with a NLP training set.

15. The non-transitory computer readable storage medium of claim 14 , further comprising associating the training data point with a NLP training set.

16. The non-transitory computer readable storage medium of claim 15 , further comprising transmitting the NLP training set to an outside system, the outside system configured to train a machine learning model based on the NLP training set.

17. The non-transitory computer readable storage medium of claim 10 , wherein the training data point is associated with a target language.

18. The non-transitory computer readable storage medium of claim 10 , wherein determining emotion data for the training data point comprises performing emotion recognition analysis to associate faces recognized by the facial recognition analysis with one or more emotions.

Assignments (3)
SECURITY INTEREST Recorded May 21, 2026
From: WRITER, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 074728/0276 →
CHANGE OF NAME Recorded May 2, 2022
From: QORDOBA, INC.
To: WRITER, INC.
Reel/Frame 059846/0765 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2019
From: ALSHIKH, WASEEM
To: QORDOBA, INC.
Reel/Frame 050568/0957 →
Continuity (2)
Provisional Application 62726193 · Aug 31, 2018
Related Publication 20200074229A1 · Mar 5, 2020