IP Library › Granted Patent US 11,593,702
Granted Patent B2
US 11,593,702 · App. 16/434,668 · Granted Feb 28, 2023

Methods, systems and computer program products for generating a training set for use during content generation

Inventors: François Pachet (Stockholm, SE); Pierre Roy (Stockholm, SE)
Assignee: Spotify AB
G06N20/00G06F30/20G06F2111/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,593,702
App. No.
16/434,668
Granted
Feb 28, 2023
Kind
B2
Abstract

A training set for use during content generation is generated by applying a first machine learning process P 1 to a first finite sequence s wherein s has a length Ls, to generate a first statistical model M(s). The first statistical model M(s) is sampled using a first sampling process G to generate a second finite sequence t wherein t has a length Lt. A second machine learning process P 2 is applied to the second finite sequence t to generate a second statistical model M(t), wherein no substring of the second finite sequence t of length d is identical to a substring of the first finite sequence s, wherein d is a predetermined number of elements in a sequence.

Claims (39)

1. A computer-implemented system for generating a training set for use during content generation, comprising:

a first machine learning processor configured to apply a first machine learning process P 1 to a first finite sequence s wherein s has a length L s , to generate a first statistical model M(s);

a first sampling processor configured to sample the first statistical model M(s) using a first sampling process G to generate a second finite sequence t wherein t has a length L t ;

a second machine learning processor configured to apply a second machine learning process P 2 to the second finite sequence t to generate a second statistical model M(t); and

a plagiarism tester operable to reject any substring in second finite sequence t of length d that is identical to a substring of first finite sequence s,

wherein no substring of the second finite sequence t of length d is identical to a substring of the first finite sequence s, wherein d is a predetermined number of elements in a sequence.

2. The system according to claim 1 , further comprising:

a counter configured to count the number of rejections made by the plagiarism tester and to communicate a signal to a user interface when the number of rejections exceeds a predetermined threshold.

3. The system according to claim 1 , further comprising:

wherein the first statistical model M(s) and the second statistical model M(t) are the same statistical model.

4. The system according to claim 1 , further comprising:

wherein the first statistical model M(s) and the second statistical model M(t) are different statistical models.

5. The system according to claim 1 , wherein a cost function defined by 1) a predefined distance of the first statistical model M(s) and the second statistical model M(t) and 2) the difference between Ls and Lt, is less than or equal to a predetermined distance.

6. A method for generating a training set for use during content generation, comprising:

applying a first machine learning process P 1 to a first finite sequence s wherein s has a length L s , to generate a first statistical model M(s);

sampling the first statistical model M(s) using a first sampling process G to generate a second finite sequence t wherein t has a length Lt;

applying a second machine learning process P 2 to the second finite sequence t to generate a second statistical model M(t); and

applying a plagiarism tester to reject any substring in second finite sequence t of length d that is identical to a substring of first finite sequence s,

wherein no substring of the second finite sequence t of length d is identical to a substring of the first finite sequence s, wherein d is a predetermined number of elements in a sequence.

7. The method according to claim 6 , further comprising:

counting the number of rejections; and

communicating a signal to a user interface when the number of rejections exceeds a predetermined threshold.

8. The method according to claim 6 , further comprising:

wherein the first statistical model M(s) and the second statistical model M(t) are the same statistical model.

9. The method according to claim 6 , further comprising:

wherein the first statistical model M(s) and the second statistical model M(t) are different statistical models.

10. The method according to claim 6 , wherein a cost function defined by (1) a predefined distance of the first statistical model M(s) and the second statistical model M(t) and (2) the difference between Ls and Lt, is less than or equal to a predetermined distance.

11. A non-transitory computer-readable medium having stored thereon sequences of instructions, the sequences of instructions including instructions which when executed by a computer system causes the computer system to perform:

applying a first machine learning process P 1 to a first finite sequence s wherein s has a length L s , to generate a first statistical model M(s);

sampling the first statistical model M(s) using a first sampling process G to generate a second finite sequence t wherein t has a length L t ;

applying a second machine learning process P 2 to the second finite sequence t to generate a second statistical model M(t); and

applying a plagiarism tester operable to reject any substring in second finite sequence t of length d that is identical to a substring of first finite sequence s,

wherein no substring of the second finite sequence t of length d is identical to a substring of the first finite sequence s, wherein d is a predetermined number of elements in a sequence.

12. The computer-readable medium according to claim 11 , further having stored thereon a sequence of instructions for causing the one or more processors to perform:

counting the number of rejections; and

communicating a signal to a user interface when the number of rejections exceeds a predetermined threshold.

13. The computer-readable medium according to claim 11 , wherein the first statistical model M(s) and the second statistical model M(t) are the same statistical model.

14. The computer-readable medium according to claim 11 , wherein the first statistical model M(s) and the second statistical model M(t) are different statistical models.

15. The computer-readable medium according to claim 11 , wherein a cost function defined by (1) a predefined distance of the first statistical model M(s) and the second statistical model M(t) and (2) the difference between Ls and Lt, is less than or equal to a predetermined distance.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2023
From: SPOTIFY AB
To: SOUNDTRAP AB
Reel/Frame 064315/0727 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2022
From: PACHET, FRANÇOIS; ROY, PIERRE
To: SPOTIFY AB
Reel/Frame 061686/0484 →
Continuity (1)
Related Publication 20200074343A1 · Mar 5, 2020