IP Library Granted Patent US 11,972,770
Granted Patent B2
US 11,972,770 · App. 17/668,347 · Granted Apr 30, 2024

Systems and methods for intelligent playback

Inventors: Yatish Jayant Naik Raikar (Bengaluru, IN); Varunkumar Tripathi (Bengaluru, IN); Karthik Mahabaleshwar Hegde (Uttara Kannada, IN)
Assignee: DISH NETWORK TECHNOLOGIES INDIA PRIVATE LIMITED
G10L21/043G10L21/055G10L25/48H04N21/4325
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,972,770
App. No.
17/668,347
Granted
Apr 30, 2024
Kind
B2
Abstract

Systems and methods for intelligent playback of media content may include an intelligent media playback system that, in response to determining the speech tempo in audio content by measuring syllable density of speech in the audio content, automatically adjusts a playback speed of the audio content as the audio content is being played based on the determined speech tempo. In some embodiments, the system may automatically and dynamically adjust the playback speed to result in a desired target speech tempo. In addition, the system may determine whether to automatically adjust playback speed of the audio content, as the media is being played, based on the detected speech tempo of the speech in the audio content and the determined type of content of media. Such automatic adjustments in playback speed result in more efficient playback of the audio content.

Claims (50)

1. A computer-implemented method for compressing digital media data, comprising:

receiving, by at least one computer processor, an audio signal representing audio content of the digital media data;

determining, by at least one computer processor, a speech tempo of speech in the audio content, wherein the speech tempo is a measure of a number of syllables of speech in the audio content per unit of time as the audio content is being played; and

in response to the determining the speech tempo of speech in the audio content, compressing, by at least one computer processor, the digital media data by re-encoding the digital media data content based on the determined speech tempo of the speech in the audio content of the digital media data; and

automatically adjusting a playback speed of the audio content as the audio content is being played based on the determined speech tempo of the speech in the audio content, wherein the automatically adjusting the playback speed of the audio content as the audio content is being played based on the determined speech tempo of the speech in the audio content includes:

storing a database including a plurality of selectable playback speeds, each selectable playback speed of the plurality of selectable playback speeds corresponding to a different speech tempo range of a plurality of different speech tempo ranges;

determining in which speech tempo range of the plurality of different speech tempo ranges the determined speech tempo of speech in the audio content falls;

selecting the speech tempo range of the plurality of different speech tempo ranges in which the determined speech tempo of speech in the audio content falls; and

changing the playback speed of the audio content as the audio content is being played to be the selectable playback speed corresponding to the selected speech tempo range of the plurality of different speech tempo ranges;

wherein:

the re-encoding of the audio content based on the determined speech tempo of the speech in the audio content of the digital media data includes:

detecting silent regions present in the audio content based on the determined speech tempo of speech in the audio content;

removing the detected silent regions from the audio content of the digital media data; and

re-encoding the digital media data content without the detected silent regions.

2. The method of claim 1 , wherein the detecting silent regions present in the audio content based on the determined speech tempo of speech in the audio content includes determining that regions in the audio content with a detected speech tempo of zero are silent regions.

3. The method of claim 1 further comprising:

receiving by at least one computer processor, a selection of a target speech tempo from a user; and

changing by at least one computer processor, the playback speed of the audio content as the audio content is being played in to have the audio played back with a resulting target speech tempo of the selected target speech tempo.

4. The method of claim 3 wherein the selection of the target speech tempo from the user is received user via a settings menu graphical user interface generated and provided by a receiving device operation and playback manager generated by the at least one computer processor.

5. The method of claim 3 further comprising:

continuously determining, by at least one computer processor, whether to increase or decrease playback speed of the audio content as the audio content is being played for each detectable corresponding incremental change in the current speech tempo of the audio content.

6. The method of claim 5 wherein a relationship between the detected speech tempo and a corresponding increase or decrease of playback speed is linear.

7. The method of claim 5 wherein a relationship between the detected speech tempo and a corresponding increase or decrease of playback speed is logarithmic.

8. The method of claim 5 wherein a relationship between the detected speech tempo and a corresponding increase or decrease of playback speed is exponential.

9. A system for compressing digital media data, comprising:

at least one processor; and

at least one memory coupled to the at least one processor, wherein the at least one memory has computer-executable instructions stored thereon that, when executed by the at least one processor, cause operations to be performed including:

receiving, by at least one computer processor, an audio signal representing audio content of the digital media data;

determining, by at least one computer processor, a speech tempo of speech in the audio content, wherein the speech tempo is a measure of a number of syllables of speech in the audio content per unit of time as the audio content is being played;

in response to the determining the speech tempo of speech in the audio content, compressing, by at least one computer processor, the digital media data by re-encoding the digital media data content based on the determined speech tempo of the speech in the audio content of the digital media data; and

automatically adjusting a playback speed of the audio content as the audio content is being played based on the determined speech tempo of the speech in the audio content, wherein the automatically adjusting the playback speed of the audio content as the audio content is being played based on the determined speech tempo of the speech in the audio content includes:

storing a database including a plurality of selectable playback speeds, each selectable playback speed of the plurality of selectable playback speeds corresponding to a different speech tempo range of a plurality of different speech tempo ranges;

determining in which speech tempo range of the plurality of different speech tempo ranges the determined speech tempo of speech in the audio content falls;

selecting the speech tempo range of the plurality of different speech tempo ranges in which the determined speech tempo of speech in the audio content falls; and

changing the playback speed of the audio content as the audio content is being played to be the selectable playback speed corresponding to the selected speech tempo range of the plurality of different speech tempo ranges;

wherein:

the re-encoding of the audio content based on the determined speech tempo of the speech in the audio content of the digital media data includes:

detecting silent regions present in the audio content based on the determined speech tempo of speech in the audio content;

removing the detected silent regions from the audio content of the digital media data; and

re-encoding the digital media data content without the detected silent regions.

10. The system of claim 9 , wherein the detecting silent regions present in the audio content based on the determined speech tempo of speech in the audio content includes determining that regions in the audio content with a detected speech tempo of zero are silent regions.

11. The system of claim 9 wherein the computer-executable instructions, when executed by the at least one processor, further cause operations to be performed including:

receiving by at least one computer processor, a selection of a target speech tempo from a user; and

changing by at least one computer processor, the playback speed of the audio content as the audio content is being played in to have the audio played back with a resulting target speech tempo of the selected target speech tempo.

12. The system of claim 11 wherein the selection of the target speech tempo from the user is received-use-F via a settings menu graphical user interface generated and provided by a receiving device operation and playback manager generated by the at least one computer processor.

13. The system of claim 11 wherein the computer-executable instructions, when executed by the at least one processor, further cause operations to be performed including:

continuously determining, by at least one computer processor, whether to increase or decrease playback speed of the audio content as the audio content is being played for each detectable corresponding incremental change in the current speech tempo of the audio content.

14. The system of claim 13 wherein a relationship between the detected speech tempo and a corresponding increase or decrease of playback speed is linear.

15. The system of claim 13 wherein a relationship between the detected speech tempo and a corresponding increase or decrease of playback speed is logarithmic.

16. The system of claim 13 wherein a relationship between the detected speech tempo and a corresponding increase or decrease of playback speed is exponential.

Assignments (2)
CHANGE OF NAME Recorded Jul 18, 2022
From: SLING MEDIA PRIVATE LIMITED
To: DISH NETWORK TECHNOLOGIES INDIA PRIVATE LIMITED
Reel/Frame 060535/0011 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2022
From: RAIKAR, YATISH JAYANT NAIK; TRIPATHI, VARUNKUMAR; HEGDE, KARTHIK MAHABALESHWAR
To: SLING MEDIA PVT. LTD
Reel/Frame 058985/0816 →
Continuity (2)
Continuation 16054910 · Aug 3, 2018
Related Publication 20220270632A1 · Aug 25, 2022