IP Library › Granted Patent US 10,102,867
Granted Patent B2
US 10,102,867 · App. 15/013,440 · Granted Oct 16, 2018

Television system, server apparatus, and television apparatus

Inventor: Naoki Yamanashi (Tokyo, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G10L21/028H04N21/233H04N21/439H04N21/4394H04N21/8113G10L21/034G10L21/0316
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,102,867
App. No.
15/013,440
Granted
Oct 16, 2018
Kind
B2
Abstract

According to one embodiment, a server apparatus includes a content provision module which selectively provides a plurality of items of content, and an audio source separation module which separates a voice component and a background sound component from an audio signal of the content, and sets different levels of volume, and a television apparatus connected to the server apparatus through a network includes an instruction module which instructs selection of the content to the content provision module of the server apparatus and instructs execution of an audio source separation process to the audio source separation module, and a playback module which plays back the content provided from the server apparatus in response to the instruction.

Claims (37)

1. A television system comprising:

a server apparatus comprising

a content provider configured to selectively provide a plurality of items of content in accordance with an instruction,

an audio source separator configured to separate an audio signal of the items of content to be provided into an audio source of a voice component and an audio source of a background sound component by a non-negative matrix factorization process in accordance with an instruction, and to set a volume of each of the audio sources in accordance with an instruction, and

a management information provider configured to provide a management information for performing an input operation of directing selection of the content to the content provider and an input operation of directing audio source separation processing and volume setting processing to the audio source separator; and

a television apparatus comprising

an instruction device connected to the server apparatus through a network, and configured to instruct the content provider of the server apparatus to select the content and to selectively instruct the audio source separator to execute a process for separating the audio sources and setting the volume based on the management information provided from the server apparatus, and

a reproduction device configured to reproduce the content provided from the server apparatus in response to the instruction, wherein

in the non-negative matrix factorization process,

a first basic matrix representing a feature of the audio source of the background sound component is created from a spectrogram of an audio signal in a zone in which a probability that the audio source of the background sound component will be included is high,

a second basic matrix representing the feature of the audio source of the background sound component is created by excluding a component that is highly related to the audio source of the voice component from the first basic matrix,

a third basic matrix representing a feature of the audio source of the voice component of the audio signal, and a first coefficient matrix are calculated by using the second basic matrix,

a spectrogram of the audio source of the voice component is estimated from a product of the third basic matrix and the first coefficient matrix, and

the audio source of the voice component is separated from the audio signal by converting the estimated spectrogram of the audio source of the voice component into a time signal.

2. The television system of claim 1 , wherein the instruction device of the television apparatus comprises a ratio instruction module configured to instruct the audio source separator to set a ratio of the volume of the audio source of the voice component and the volume of the audio source of the background sound component at a ratio specified by a user, and

the audio source separator of the server apparatus is configured to set the ratio of the volume of the audio source of the voice component and the volume of the audio source of the background sound component at the ratio specified by the ratio instruction module.

3. A server apparatus configured to connect to a television apparatus through a network, the server apparatus comprising:

a content provider configured to selectively provide a plurality of items of content to the television apparatus in accordance with an instruction from the television apparatus;

an audio source separator configured to separate an audio signal of the items of content provided to the television apparatus into an audio source of a voice component and an audio source of a background sound component by a non-negative matrix factorization process in accordance with an instruction from the television apparatus, and to set a volume of each of the audio sources in accordance with an instruction from the television apparatus, and

a management information provider configured to provide management information for performing an input operation of directing selection of the content to the content provider and an input operation of directing audio source separation processing and volume setting processing to the audio source separator wherein

in the non-negative matrix factorization process,

a first basic matrix representing a feature of the audio source of the background sound component is created from a spectrogram of an audio signal in a zone in which a probability that the audio source of the background sound component will be included is high,

a second basic matrix representing the feature of the audio source of the background sound component is created by excluding a component that is highly related to the audio source of the voice component from the first basic matrix,

a third basic matrix representing a feature of the audio source of the voice component of the audio signal, and a first coefficient matrix are calculated by using the second basic matrix,

a spectrogram of the audio source of the voice component is estimated from a product of the third basic matrix and the first coefficient matrix, and

the audio source of the voice component is separated from the audio signal by converting the estimated spectrogram of the audio source of the voice component into a time signal.

4. The server apparatus of claim 3 , wherein the audio source separator sets a ratio of the volume of the audio source of the voice component and the volume of the audio source of the background sound component at a ratio specified by the television apparatus.

5. A television apparatus configured to connect to a server apparatus through a network, the server apparatus comprising a content provider configured to selectively provide a plurality of items of content in accordance with an instruction, an audio source separator configured to separate an audio signal of the items of content to be provided into an audio source of a voice component and an audio source of a background sound component by a non-negative matrix factorization process in accordance with an instruction, and to set a volume of each of the audio sources in accordance with an instruction, and a management information provider configured to provide management information for performing an input operation of directing selection of the content to the content provider and an input operation of directing audio source separation processing and volume setting processing to the audio source separator, the television apparatus comprising:

an instruction device configured to instruct the content provider of the server apparatus to select the content and to selectively instruct the audio source separator to execute a process for separating the audio sources and setting the volume based on the management information provided from the server apparatus; and

a reproduction device configured to reproduce the content provided from the server apparatus in response to the instruction, wherein

in the non-negative matrix factorization process,

a first basic matrix representing a feature of the audio source of the background sound component is created from a spectrogram of an audio signal in a zone in which a probability that the audio source of the background sound component will be included is high,

a second basic matrix representing the feature of the audio source of the background sound component is created by excluding a component that is highly related to the audio source of the voice component from the first basic matrix,

a third basic matrix representing a feature of the audio source of the voice component of the audio signal, and a first coefficient matrix are calculated by using the second basic matrix,

a spectrogram of the audio source of the voice component is estimated from a product of the third basic matrix and the first coefficient matrix, and

the audio source of the voice component is separated from the audio signal by converting the estimated spectrogram of the audio source of the voice component into a time signal.

6. The television apparatus of claim 5 , wherein when the audio source separator of the server apparatus sets a ratio of the volume of the audio source of the voice component and the volume of the audio source of the background sound component at a specified ratio, the instruction device comprises a ratio instruction module configured to instruct the audio source separator to set the ratio of the volume of the audio source of the voice component and the volume of the audio source of the background sound component at a ratio specified by a user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2016
From: YAMANASHI, NAOKI
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 037650/0016 →
Continuity (2)
Continuation PCTJP2013084927 · Dec 26, 2013
Related Publication 20160148623A1 · May 26, 2016