IP Library Granted Patent US 12711674
Granted Patent B1
US 12711674 · App. 18/816,218 · Granted Aug 18, 2026

Static and dynamically created presets for immersive video presentations

Inventors: Rebekah Maggor (Paris, FR); Phil Libin (Bentonville, AR)
Assignee: mmhmm inc.
G06T11/00G06T19/006G06V20/41G06V20/46G06V20/48G06V20/49G11B27/036G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711674
App. No.
18/816,218
Granted
Aug 18, 2026
Kind
B1
Abstract

Providing a video presentation includes preparing a plurality of presentation presets for a plurality of episodes of the video presentation. Each of the presets includes a set of one or more presenters for the video presentation, materials for the video presentation, and a background environment for the video presentation. Providing a video presentation also includes choosing one of the presets for each of the episodes, actuating each of the presets for each of the episodes, determining a sentiment of the presenters during each of the episodes of the presentation, determining whether the background environment of each of the presets matches the sentiment of the presenters during each of the episodes, and changing the background environment of each of the presets that do not match the sentiment of the presenters during the video presentation. The background environment may include a structured list of visual components and parameters therefor.

Claims (27)

1 . A method of modifying a video presentation, comprising:

preparing a plurality of presentation presets for a plurality of episodes of the video presentation, each of the presentation presets including a set of one or more presenters for the video presentation, materials for the video presentation, and a background environment for the video presentation;

choosing an initial one of the presentation presets for each of the episodes;

analyzing the video presentation by quantifying a first difference between each scene of each episode and the presentation preset corresponding to the episode containing the scene and by quantifying a second difference between the scene and the presentation preset of either a subsequent one of the episodes or, for a last one of the episodes, a last scene of the last one of the episodes;

determining a particular scene of each episode having a maximum value of a minimum of the first difference and the second difference;

splitting at least one of the episodes into two different episodes at the particular scene in response to a minimum of the first difference and the second difference corresponding to the particular scene being greater than a predetermined threshold; and

choosing a new presentation preset for a second one of the two different episodes based on the particular scene, wherein the first difference is based on a value of a function aggregating metrics of distance between an arrangement of the presenters for each scene of each episode and an arrangement of the presenters for a particular one of the presentation presets corresponding to the episode and the second difference is based on a value of a function aggregating metrics of distance between an arrangement of the presenters for each scene of each episode and an arrangement of the presenters for either one of the presentation presets corresponding to a subsequent one of the episodes or, for a last one of the episodes, a last scene of the last one of the episodes.

2 . The method of claim 1 , wherein the background environment includes a structured list of visual components and parameters therefor.

3 . The method of claim 1 , wherein at least some of the presentation presets include optional components.

4 . The method of claim 3 , wherein the optional components include at least one of the following: at least one teleprompter for one or more of the presenters, an expiration clock for presentation time control within an episode, rules for changing a background environment, and an artificial intelligence component that monitors the video presentation and utilizes a technology stack.

5 . The method of claim 4 , wherein the technology stack includes voice emotion recognition, facial recognition, speech recognition for obtaining transcripts, and sentiment recognition to identify a sentiment of the presenter team and apply the rules for changing the background environment.

6 . The method of claim 1 , wherein the background environment is a physical background environment.

7 . The method of claim 6 , wherein a visual component is superimposed on at least one portion of the physical background environment.

8 . The method of claim 6 , wherein a plurality of virtual layers are superimposed with the physical background environment.

9 . The method of claim 1 , wherein the metrics of distance between the arrangements of the presenters are at least one of: a distance between full or partial orders of the presenters, a ratio of average Hausdorff distances between all pairs of images of the presenters, a distance between bounding rectangles, or cubes of the collections of presenter images.

10 . A non-transitory computer readable medium containing software that, when executed by a processor, modifies a video presentation, the software comprising:

executable code that chooses an initial one of a plurality of presentation presets for each episode of a plurality of episodes of the video presentation, each of the presentation presets including a set of one or more presenters for the video presentation, materials for the video presentation, and a background environment for the video presentation;

executable code that analyzes the video presentation by quantifying a first difference between each scene of each episode and the presentation preset corresponding to the episode containing the scene and by quantifying a second difference between the scene and the presentation preset of either a subsequent one of the episodes or, for a last one of the episodes, a last scene of the last one of the episodes;

executable code that determines a particular scene of each episode having a maximum value of a minimum of the first difference and the second difference;

executable code that splits at least one of the episodes into two different episodes at the particular scene in response to a minimum of the first difference and the second difference corresponding to the particular scene being greater than a predetermined threshold; and

executable code that chooses a new presentation preset for a second one of the two different episodes based on the particular scene, wherein the first difference is based on a value of a function aggregating metrics of distance between an arrangement of the presenters for each scene of each episode and an arrangement of the presenters for a particular one of the presentation presets corresponding to the episode and the second difference is based on a value of a function aggregating metrics of distance between an arrangement of the presenters for each scene of each episode and an arrangement of the presenters for either one of the presentation presets corresponding to a subsequent one of the episodes or, for a last one of the episodes, a last scene of the last one of the episodes.

11 . The non-transitory computer readable medium of claim 10 , wherein the background environment includes a structured list of visual components and parameters therefor.

12 . The non-transitory computer readable medium of claim 10 , wherein at least some of the presentation presets include optional components.

13 . The non-transitory computer readable medium of claim 12 , wherein the optional components include at least one of the following: at least one teleprompter for one or more of the presenters, an expiration clock for presentation time control within an episode, rules for changing a background environment, and an artificial intelligence component that monitors the video presentation and utilizes a technology stack.

14 . The non-transitory computer readable medium of claim 13 , wherein the technology stack includes voice emotion recognition, facial recognition, speech recognition for obtaining transcripts, and sentiment recognition to identify a sentiment of the presenter team and apply the rules for changing the background environment.

15 . The non-transitory computer readable medium of claim 10 , wherein the background environment is a physical background environment.

16 . The non-transitory computer readable medium of claim 15 , wherein at least one of: a visual component or a plurality of virtual layers are superimposed on at least one portion of the physical background environment.