IP Library Granted Patent US 11,729,476
Granted Patent B2
US 11,729,476 · App. 17/170,695 · Granted Aug 15, 2023

Reproduction control of scene description

Inventors: Brant Candelore (San Diego, CA); Mahyar Nejat (San Diego, CA); Peter Shintani (San Diego, CA); Robert Blanchard (San Diego, CA)
Assignee: SONY GROUP CORPORATION
H04N21/84G06N20/00H04N19/126H04N21/43074H04N21/4662H04N21/4882H04N21/6587H04N21/8133H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,729,476
App. No.
17/170,695
Granted
Aug 15, 2023
Kind
B2
Abstract

A media rendering device and method for reproduction control of scene description is provided. The media rendering device retrieves media content that includes a set of filmed scenes and text information. The text information includes video description information and timing information. The video description information describes a filmed scene in the set of filmed scenes. The media rendering device further extracts the timing information to reproduce the video description information from the text information of the filmed scene. The media rendering device further controls the reproduction of the video description information in either a textual representation or in a textual and audio representation at a first-time interval indicated by the extracted timing information of the filmed scene.

Claims (83)

1. A media rendering device, comprising:

a memory configured to store a trained machine learning (ML) model; and

circuitry configured to:

retrieve media content that comprises a set of filmed scenes and text information which includes video description information, speed information, and timing information, wherein

the video description information describes a filmed scene in the set of filmed scenes;

extract the timing information, to reproduce the video description information, from the text information of the filmed scene;

extract a first-time interval from the timing information, wherein

the first-time interval corresponds to a natural pause between consecutive audio portions of the filmed scene;

determine a set of second-time intervals of the filmed scene, wherein

each of the set of second-time intervals indicates a time interval for reproduction of an audio portion of the filmed scene in the set of filmed scenes;

determine a third-time interval which indicates a time duration required to reproduce an audio representation of the video description information of the filmed scene;

determine a multiplication factor based on a ratio of the determined third-time interval and the first-time interval;

determine a speed to reproduce the audio representation of the video description information based on the multiplication factor and an actual playback speed of the audio representation of the video description information; wherein the speed information indicates the speed for the reproduction of the audio representation of the video description information;

determine context information of the filmed scene based on an analysis of at least one characteristic of the filmed scene;

determine an audio characteristic to reproduce the audio representation of the video description information based on an application of the trained ML model on the determined context information of the filmed scene; and

control the reproduction of the audio representation of the video description information at the first-time interval indicated by the extracted timing information of the filmed scene, based on the speed information and the determined audio characteristic.

2. The media rendering device according to claim 1 , wherein

the circuitry is further configured to:

extract the speed information, to reproduce the video description information, from the text information of the filmed scene; and

control, based on the extracted speed information, the reproduction of the audio representation of the video description information at the first-time interval indicated by the extracted timing information of the filmed scene.

3. The media rendering device according to claim 1 , wherein the circuitry is further configured to:

determine a set of fourth-time intervals of the filmed scene, wherein each of the set of fourth-time intervals is different than the set of second-time intervals; and

select the first-time interval from the set of fourth-time intervals, wherein the first-time interval is higher than a time-interval threshold.

4. The media rendering device according to claim 1 , wherein the determined speed is lower than the actual playback speed of the audio representation.

5. The media rendering device according to claim 1 , wherein the determined speed is higher than the actual playback speed of the audio representation.

6. The media rendering device according to claim 1 , wherein

the circuitry is further configured to determine the speed to reproduce the audio representation of the video description information based on a defined speed setting associated with the media rendering device, and

the defined speed setting indicates a maximum speed to reproduce the audio representation of the video description information.

7. The media rendering device according to claim 6 , wherein the circuitry is further configured to:

receive the speed information with the text information; and

control playback of one of an image portion or the audio portion of the filmed scene based on the determined speed and the defined speed setting.

8. The media rendering device according to claim 6 , wherein the circuitry is further configured to:

receive a first user input which indicates profile information of a user to whom the media content is being rendered; and

determine the defined speed setting to reproduce the audio representation of the video description information based on the received first user input.

9. The media rendering device according to claim 1 , wherein the circuitry is further configured to:

receive a first user input which corresponds to a description of one of the set of filmed scenes;

search the received first user input in the video description information associated with each of the set of filmed scenes;

determine playback timing information to playback the media content based on the search; and

control the playback of the media content based on the determined playback timing information.

10. The media rendering device according to claim 1 , wherein the first-time interval is between a first dialogue word and a second dialogue word of the filmed scene.

11. The media rendering device according to claim 10 , wherein

the first dialogue word is a last word of a first shot of the filmed scene and the second dialogue word is a first word of a second shot of the filmed scene, and

the first shot and the second shot are consecutive shots of the filmed scene.

12. The media rendering device according to claim 1 , wherein

the video description information, that describes the filmed scene, includes cognitive information about animated or in-animated objects present in the filmed scene, and

the circuitry is further configured to control playback of the cognitive information included in the video description information of the filmed scene.

13. The media rendering device according to claim 1 , further comprising a display device configured to reproduce a textual representation of the video description information.

14. The media rendering device according to claim 1 , wherein

the media content further comprises closed caption information to represent the audio portion of each of the set of filmed scenes, and

the video description information which describes each of the set of filmed scenes is encoded with the closed caption information in the media content.

15. The media rendering device according to claim 1 , wherein the circuitry is further configured to control an audio rendering device, associated with the media rendering device, to reproduce the audio representation of the video description information and the audio portion of the filmed scene.

16. A method, comprising:

in a media rendering device:

storing a trained machine learning (ML) model in a memory;

retrieving media content that comprises a set of filmed scenes and text information which includes video description information, speed information, and timing information, wherein the video description information describes a filmed scene in the set of filmed scenes;

extracting the timing information to reproduce the video description information, from the text information of the filmed scene;

extracting a first-time interval from the timing information, wherein

the first-time interval corresponds to a natural pause between consecutive audio portions of the filmed scene;

determining a set of second-time intervals of the filmed scene, wherein

each of the set of second-time intervals indicates a time interval for reproduction of an audio portion of the filmed scene in the set of filmed scenes;

determining a third-time interval which indicates a time duration required to reproduce an audio representation of the video description information of the filmed scene;

determining a multiplication factor based on a ratio of the determined third-time interval and the first-time interval;

determining a speed to reproduce the audio representation of the video description information based on the multiplication factor and an actual playback speed of the audio representation of the video description information; wherein the speed information indicates the speed for the reproduction of the audio representation of the video description information;

determining context information of the filmed scene based on an analysis of at least one characteristic of the filmed scene;

determining an audio characteristic to reproduce the audio representation of the video description information based on an application of the trained ML model on the determined context information of the filmed scene; and

controlling the reproduction of the audio representation of the video description information at the first-time interval indicated by the extracted timing information of the filmed scene, based on the speed information and the determined audio characteristic.

17. The method according to claim 16 , further comprising:

extracting the speed information, to reproduce the video description information, from the text information of the filmed scene; and

controlling, based on the extracted speed information, the reproduction of the audio representation of the video description information at the first-time interval indicated by the extracted timing information of the filmed scene.

18. A non-transitory computer-readable medium having stored thereon, computer-executable instructions that when executed by a media rendering device, causes the media rendering device to execute operations, the operations comprising:

storing a trained machine learning (ML) model in a memory;

retrieving media content that comprises a set of filmed scenes and text information which includes video description information, speed information, and timing information, wherein the video description information describes a filmed scene in the set of filmed scenes;

extracting the timing information to reproduce the video description information, from the text information of the filmed scene;

extracting a first-time interval from the timing information, wherein

the first-time interval corresponds to a natural pause between consecutive audio portions of the filmed scene;

determining a set of second-time intervals of the filmed scene, wherein

each of the set of second-time intervals indicates a time interval for reproduction of an audio portion of the filmed scene in the set of filmed scenes;

determining a third-time interval which indicates a time duration required to reproduce an audio representation of the video description information of the filmed scene;

determining a multiplication factor based on a ratio of the determined third-time interval and the first-time interval;

determining a speed to reproduce the audio representation of the video description information based on the multiplication factor and an actual playback speed of the audio representation of the video description information, wherein the speed information indicates the speed for the reproduction of the audio representation of the video description information;

determining context information of the filmed scene based on an analysis of at least one characteristic of the filmed scene;

determining an audio characteristic to reproduce the audio representation of the video description information based on an application of the trained ML model on the determined context information of the filmed scene; and

controlling the reproduction of the audio representation of the video description information at the first-time interval indicated by the extracted timing information of the filmed scene, based on the speed information and the determined audio characteristic.

Assignments (2)
CHANGE OF NAME Recorded Oct 8, 2021
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 057886/0732 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2021
From: CANDELORE, BRANT; NEJAT, MAHYAR; SHINTANI, PETER; BLANCHARD, ROBERT
To: SONY CORPORATION
Reel/Frame 055688/0288 →
Continuity (1)
Related Publication 20220256156A1 · Aug 11, 2022
Cited By (1)
US 12,541,656