CLOSED CAPTIONS FOR REAL TIME COMMUNICATION
The claimed subject matter provides systems and/or methods that facilitate yielding closed caption service associated with real time communication. For example, audio data and video data can be obtained from an active speaker in a real time teleconference. Moreover, the audio data can be converted into a set of characters (e.g., text data) that can be transmitted to other participants of the real time teleconference. Additionally, the real time teleconference can be a peer to peer conference (e.g., where a sending endpoint communicates with a receiving endpoint) and/or a multi-party conference (e.g., where an audio/video multi-point control unit (AVMCU) routes data such as the audio data, the video data, and the text data between endpoints).
1 . A system that facilitates providing closed captions for real time communications, comprising:
a real time conferencing component that communicates with at least one disparate real time conferencing component; and
a text streaming component that transmits text data utilized to render closed captions associated with a real time teleconference from the real time conferencing component to the at least one disparate real time conferencing component, the text data corresponding to audio data of the real time teleconference.
2 . The system of claim 1 , further comprising a speech to text conversion component that converts the audio data into the text data in real time.
3 . The system of claim 2 , further comprising a translation component that translates the text data from a first language into one or more disparate languages.
4 . The system of claim 1 , the text streaming component transmits the text data in a compressed form.
5 . The system of claim 1 , further comprising:
a video streaming component that transmits video data to the at least one disparate real time conferencing component; and
an audio streaming component that transmits audio data with the at least one disparate real time conferencing component.
6 . The system of claim 5 , further comprising a synchronization component that correlates the text data, the video data, and the audio data in time for presentation to listening participants in the real time teleconference, the synchronization component at least one of embeds the text data in the video data or employs timestamps with multiplexed streams associated with the text data, the video data, and the audio data.
7 . The system of claim 1 , the real time conferencing component negotiates with the at least one disparate real time conferencing component as to whether to transmit video data with the text data or the audio data.
8 . The system of claim 1 , the real time conferencing component transmits the text data to the at least one disparate real time conferencing component when the at least one real time conferencing component requests the text data.
9 . The system of claim 1 , the real time teleconference being a peer to peer conference where the real time conferencing component is a sending endpoint and the at least one disparate real time conferencing component is a receiving endpoint.
10 . The system of claim 1 , the real time teleconference being a multi-party conference where the real time conferencing component is a sending endpoint or an audio/video multi-point control unit (AVMCU) and the at least one disparate real time conferencing component is the AVMCU or a receiving endpoint.
11 . The system of claim 10 , the sending endpoint or the AVMCU further comprises a speech to text conversion component that converts the audio data into the text data.
12 . The system of claim 1 , the text streaming component transmits a text stream associated with a dominant speaker when a plurality of speakers are concurrently active or transmits a plurality of text streams corresponding with each of the concurrently active speakers.
13 . A method that facilitates routing data between endpoints in a multi-party real time conference, comprising:
identifying a sending endpoint associated with an active speaker at a particular time from a set of endpoints;
obtaining video data, audio data, and text data associated with a real time communication from the sending endpoint;
determining whether to send the video data with the audio data and/or the text data for each of the remaining endpoints in the set; and
transmitting the video data, the audio data, and/or the text data according to the respective determinations.
14 . The method of claim 13 , further comprising identifying disparate endpoints from the set as being associated with the active speaker at differing times.
15 . The method of claim 13 , further comprising obtaining the text data from the sending endpoint upon the text data being generated by the sending endpoint based upon the audio data.
16 . The method of claim 13 , further comprising converting the audio data into the text data in real time.
17 . The method of claim 13 , further comprising receiving a request for the text data from at least one of the remaining endpoints in the set.
18 . The method of claim 17 , the request being received in response to an output component associated with the at least one remaining endpoints being muted.
19 . The method of claim 13 , further comprising transmitting the text data in a selected language.
20 . A system that provides closed caption service associated with real time communications, comprising:
means for obtaining audio data and video data for transmission in a real time conference;
means for generating text data based upon the audio data, the text data enables presenting closed captions at a receiving endpoint; and
means for transmitting the audio data, the video data, and the text data.