IP Library › Granted Patent US 11,929,095
Granted Patent B1
US 11,929,095 · App. 17/852,445 · Granted Mar 12, 2024

Speed adjustment of recorded audio and video to maximize desirable cognitive effects for the audience

Inventor: Phil Libin (Bentonville, AR)
Assignee: mmhmm inc.
G11B27/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,929,095
App. No.
17/852,445
Granted
Mar 12, 2024
Kind
B1
Abstract

Setting a replay speed of a pre-recorded video presentation includes determining a mood of a presenter of the pre-recorded video presentation, determining complexity of material that is presented in the pre-recorded video presentation, and setting a replay speed based on the mood of the presenter and the complexity of the material that is presented. Setting a replay speed of a pre-recorded video presentation may also include adjusting the replay speed based on determining a desired speech tempo for a listener. The desired speech tempo of the listener may be based on time of day, age of the listener, and/or comprehension level of the listener. Measuring the comprehension level of the listener may be based facial expressions of the listener, eye-tracking of the listener, and/or listener comprehension quizzes. Measuring the mood of the presenter may be based on facial recognition, sentiment recognition, and/or gesture recognition.

Claims (40)

1. A method of setting a replay speed of a pre-recorded video presentation, comprising:

determining a mood of a presenter of the pre-recorded video presentation;

determining complexity of material that is presented in the pre-recorded video presentation; and

setting a replay speed based on the mood of the presenter and the complexity of the material that is presented, wherein measuring the mood of the presenter is based on at least one of: facial recognition, sentiment recognition, and gesture recognition; and

accelerating the replay speed in response to the presenter changing from a serious and thoughtful mood to an excited and enthusiastic emotional state.

2. A method of setting a replay speed of a pre-recorded video presentation, comprising:

determining a mood of a presenter of the pre-recorded video presentation;

determining complexity of material that is presented in the pre-recorded video presentation; and

setting a replay speed based on the mood of the presenter and the complexity of the material that is presented, wherein the replay speed is optimized according to feedback from a plurality of users playing one or more test video presentations at a plurality of replay speeds and wherein the replay speed is optimized using an experimental space that is a multi-dimensional parallelepiped with a parameter subspace and an axis for revised values of the replay speed.

3. The method of claim 2 , further comprising:

adjusting the replay speed based on determining a desired speech tempo for a listener.

4. The method of claim 3 , wherein the desired speech tempo is the listener is based on at least one of: time of day, age of the listener, and comprehension level of the listener.

5. The method of claim 4 , wherein measuring the comprehension level of the listener is based on at least one of: facial expressions of the listener, eye-tracking of the listener, and listener comprehension quizzes.

6. The method of claim 2 , wherein the parameter space corresponds to the mood of the presenter of the pre-recorded video presentation and the complexity of the material.

7. The method of claim 6 , wherein the plurality of users is presented with different combinations of replay speeds, material complexity, and presenter moods.

8. The method of claim 7 , wherein the feedback from the plurality of users is aggregated into a quality function that represents preferences of the users for various combinations of replay speeds, material complexity, and presenter moods.

9. The method of claim 2 , wherein the listener chooses whether to replay the pre-recorded video presentation at a constant acceleration or at the replay speed that is set based on the feedback from the plurality of users.

10. The method of claim 2 , further comprising:

adjusting the replay speed based on at least one of: timber of speech of the presenter, intelligibility of the speech of the presenter, and intonation of the speech of the presenter.

11. A method of setting a replay speed of a pre-recorded video presentation, comprising:

determining a mood of a presenter of the pre-recorded video presentation;

determining complexity of material that is presented in the pre-recorded video presentation; and

setting a replay speed based on the mood of the presenter and the complexity of the material that is presented, wherein the replay speed is iteratively reset according to the mood of the presenter and the complexity of the material that is presented until an integrated consistency criteria is met.

12. The method of claim 11 , wherein the complexity of the material that is presented in the pre-recorded video presentation is based on at least one of: readability criteria for recognized text of the presentation, complexity of visual portions of the presentation, and intensity of interaction of the presenter with the visual portions of the presentation.

13. The method of claim 11 , wherein the pre-recorded video presentation is divided into a plurality of segments and wherein each of the segments is provided with a replay speed that is independent of a replay speed of different ones of the segments.

14. The method of claim 13 , wherein the segments are determined based on a relationship between an actual speech tempo of the presenter, an emotional state of the presenter and the complexity of the material that is presented.

15. The method of claim 14 , wherein the actual speech tempo of the presenter is determined using a sliding average window having a width between 20 seconds and 30 seconds.

16. The method of claim 11 , further comprising:

adjusting the replay speed based on at least one of: timber of speech of the presenter, intelligibility of the speech of the presenter, and intonation of the speech of the presenter.

17. The method of claim 11 , wherein the integrated consistency criteria is based, at least in part, on at least one of: a percentage of newly misrecognized words, an overall drop in speech recognition accuracy, and deviation in recognized emotional states.

18. A non-transitory computer readable software medium containing software that sets a replay speed of a pre-recorded video presentation, the software comprising:

executable code that determines a mood of a presenter of the pre-recorded video presentation;

executable code that determines complexity of material that is presented in the pre-recorded video presentation;

executable code that sets a replay speed based on the mood of the presenter and the complexity of the material that is presented, wherein measuring the mood of the presenter is based on at least one of: facial recognition, sentiment recognition, and gesture recognition; and

executable code that accelerates the replay speed in response to the presenter changing from a serious and thoughtful mood to an excited and enthusiastic emotional state.

19. A non-transitory computer readable software medium containing software that sets a replay speed of a pre-recorded video presentation, the software comprising:

executable code that determines a mood of a presenter of the pre-recorded video presentation;

executable code that determines complexity of material that is presented in the pre-recorded video presentation; and

executable code that sets a replay speed based on the mood of the presenter and the complexity of the material that is presented, wherein the replay speed is iteratively reset according to the mood of the presenter and the complexity of the material that is presented until an integrated consistency criteria is met.

20. The non-transitory computer readable software medium of claim 19 , wherein the integrated consistency criteria is based, at least in part, on at least one of: a percentage of newly misrecognized words, an overall drop in speech recognition accuracy, and deviation in recognized emotional states.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2023
From: LIBIN, PHIL
To: MMHMM INC.
Reel/Frame 062476/0222 →
Continuity (1)
Provisional Application 63223593 · Jul 20, 2021
Cited By (2)
US 12,253,998 US 12,262,068