Multimedia data processing method and apparatus, device and medium
View Patent ↗A multimedia data processing method and apparatus, a device and a medium, wherein the method includes: receiving text information input by a user; generating multimedia data based on the text information, in response to a processing instruction for the text information, and exhibiting a multimedia edit interface for performing an edition operation on the multimedia data, wherein the multimedia data includes a plurality of multimedia clips, and the multimedia edit interface includes a first edit track, a second edit track and a third edit track, and the first track clip, the second track clip and the third track clip whose timelines are aligned on the edit tracks respectively identify a text clip, a video image clip and a voice clip corresponding thereto.
1 . A multimedia data processing method, comprising:
receiving text information input by a user via a first interface, wherein the first interface comprises a control for generating a video;
splitting the text information into a plurality of text clips;
in response to receiving a selection of the control, generating a plurality of multimedia clips of the video based on the text information, wherein each of the plurality of multimedia clips corresponds to a different one of the plurality of text clips, and wherein each of the plurality of multimedia clips comprises a voice clip and a video image clip associated with the corresponding text clip;
displaying a second interface configured to perform editing operations on at least a portion of the generated plurality of multimedia clips, wherein the second interface comprises a first region and a second region;
displaying a first edit track, a second edit track, and a third edit track in the first region of the second interface, and simultaneously displaying a plurality of first track clips in the first edit track, a plurality of second track clips in the second edit track, and a plurality of third track clips in the third edit track, wherein the plurality of first track clips are configured to display the plurality of text clips, wherein the plurality of second track clips are configured to display video image clips corresponding to the plurality of text clips, wherein the plurality of third track clips are configured to display voice clips corresponding to the plurality of text clips, and wherein a timeline of the plurality of first track clips in the first edit track is aligned with a timeline of the plurality of second track clips in the second edit track and a timeline of the plurality of third track clips in the third edit track to respectively identify a text clip, a video image clip corresponding to the text clip, and a voice clip corresponding to the text clip;
determining a type of a clip selected from the plurality of text clips, the video image clips, and the voice clips displayed in the first region of the second interface; and
displaying a set of editing controls corresponding to the type of the selected clip in the second region of the second interface, wherein the second region of the second interface is configured to display one of a plurality sets of editing controls, and the plurality of sets of editing controls correspond to a plurality of clip types.
2 . The method according to claim 1 , wherein, the receiving text information input by a user, comprises:
receiving text information input by the user in a text region; and/or,
receiving link information input by the user in a link region, identifying the link information to acquire text information on a corresponding page, and displaying the same in the text region for edition by the user.
3 . The method according to claim 1 , further comprising:
displaying a timbre selection entry control;
display a candidate timbre menu, in response to a trigger operation of the user for the timbre selection entry control, wherein, the candidate timbre menu comprises candidate timbres, as well as an audition control corresponding to the candidate timbre;
determining a first target timbre according to a selection operation of the user for the candidate timbre menu; and
acquiring a plurality of voice clips generated through speeches of the plurality of text clips split from the text information based on the first target timbre.
4 . The method according to claim 1 , further comprising:
displaying a currently identified text clip on a first target track clip in a text edit region, in response to the first target track clip selected by the user on the first edit track;
updating to identify a target text clip on the first target track clip, based on the target text clip generated through modification performed by the user on the currently displayed text clip in the text edit region.
5 . The method according to claim 4 , further comprising:
determining a third target track clip corresponding to the first target track clip on the third edit track clip, in response to a text update operation performed on the target text clip on the first target track clip; and
acquiring a target voice clip corresponding to the target text clip, and updating to identify the target voice clip on the third target track clip.
6 . The method according to claim 5 , further comprising:
keeping the second edit track unchanged, in a case where it is detected that a first update time length corresponding to the target text clip on the first edit track is inconsistent with a time length corresponding to the text clip before modification, and displaying a first update track clip corresponding to the first update time length in a preset first candidate region, wherein, the target text clip is identified on the first update track clip;
keeping the second edit track unchanged, in a case where it is detected that a third update time length corresponding to the target voice clip on the third edit track is inconsistent with a time length corresponding to the voice clip before modification, and displaying a third update track clip corresponding to the third update time length in a preset second candidate region, wherein, the target voice clip is identified on the third update track clip.
7 . The method according to claim 5 , further comprising:
adjusting a length of the first target track clip according to a first update time, in a case where it is detected that a first update time length corresponding to the target text clip on the first edit track is inconsistent with a time length corresponding to the text clip before modification;
adjusting a length of the third target track clip according to a third update time, in a case where it is detected that a third update time length corresponding to the target voice clip on the third edit track is inconsistent with a time length corresponding to the voice clip before modification;
correspondingly adjusting a length of a second target track clip corresponding to the first target track clip and the third target track clip on the second edit track, so that timelines of the adjusted first target track clip, the adjusted second target track clip, and the adjusted third target track clip are aligned.
8 . The method according to claim 4 , further comprising:
determining a second target track clip corresponding to the first target track clip is on the second edit track, in response to a text update operation performed on the target text clip on the first target track clip;
acquiring a target video image clip matched with the target text clip, and updating to identify the target video image clip on the second target track clip.
9 . The method according to claim 1 , further comprising:
responding to a third target track clip selected by the user on the third edit track, wherein, the third target track clip correspondingly identifies the voice clip corresponding to the text clip displayed by the first target track clip;
displaying a current timbre used by the voice clip on the third target track clip in a preset audio edit region, and displaying an alternative candidate timbre;
updating to identify a target voice clip on the third target track clip, based on the second target timbre generated through modification performed by the user on the current timbre according to the candidate timbre in the audio edit region, wherein, the target voice clip is a voice clip generated by reading the text clip identified by the first target track clip by using the second target timbre.
10 . The method according to claim 1 , wherein, the second interface further comprises:
a fourth edit track, configured to identify background audio data;
displaying a current background sound used by the fourth edit track in a preset background sound edit region, in response to a trigger operation for the fourth edit track, and displaying an alternative candidate background sound; and
updating to identify a target background sound on the fourth edit track, based on the target background sound generated through modification by the user on the current background sound according to the candidate background sound in the background sound edit region.
11 . An electronic device, comprising:
a processor; and a memory, configured to store executable instructions, wherein the processor is configured to read the executable instructions from the memory, and execute the executable instructions to implement operations comprising:
receiving text information input by a user via a first interface, wherein the first interface comprises a control for generating a video;
splitting the text information into a plurality of text clips;
in response to receiving a selection of the control, generating a plurality of multimedia clips of the video based on the text information, wherein each of the plurality of multimedia clips corresponds to a different one of the plurality of text clips, and wherein each of the plurality of multimedia clips comprises a voice clip and a video image clip associated with the corresponding text clip;
displaying a second interface configured to perform editing operations on at least a portion of the generated plurality of multimedia clips, wherein the second interface comprises a first region and a second region;
displaying a first edit track, a second edit track, and a third edit track in the first region of the second interface, and simultaneously displaying a plurality of first track clips in the first edit track, a plurality of second track clips in the second edit track, and a plurality of third track clips in the third edit track, wherein the plurality of first track clips are configured to display the plurality of text clips, wherein the plurality of second track clips are configured to display video image clips corresponding to the plurality of text clips, wherein the plurality of third track clips are configured to display voice clips corresponding to the plurality of text clips, and wherein a timeline of the plurality of first track clips in the first edit track is aligned with a timeline of the plurality of second track clips in the second edit track and a timeline of the plurality of third track clips in the third edit track to respectively identify a text clip, a video image clip corresponding to the text clip, and a voice clip corresponding to the text clip;
determining a type of a clip selected from the plurality of text clips, the video image clips, and the voice clips displayed in the first region of the second interface; and
displaying a set of editing controls corresponding to the type of the selected clip in the second region of the second interface, wherein the second region of the second interface is configured to display one of a plurality sets of editing controls, and the plurality of sets of editing controls correspond to a plurality of clip types.
12 . The electronic device according to claim 11 , wherein, the receiving text information input by a user, comprises:
receiving text information input by the user in a text region; and/or,
receiving link information input by the user in a link region, identifying the link information to acquire text information on a corresponding page, and displaying the same in the text region for edition by the user.
13 . The electronic device to claim 11 , the operations further comprising:
displaying a timbre selection entry control;
displaying a candidate timbre menu, in response to a trigger operation of the user for the timbre selection entry control, wherein, the candidate timbre menu comprises candidate timbres, as well as an audition control corresponding to the candidate timbre;
determining a first target timbre according to a selection operation of the user for the candidate timbre menu; and
acquiring a plurality of voice clips generated through speeches of the plurality of text clips split from the text information based on the first target timbre.
14 . The electronic device according to claim 11 , the operations further comprising:
displaying a currently identified text clip on a first target track clip in a text edit region, in response to the first target track clip selected by the user on the first edit track;
updating to identify a target text clip on the first target track clip, based on the target text clip generated through modification performed by the user on the currently displayed text clip in the text edit region.
15 . The electronic device according to claim 14 , the operations further comprising:
determining a third target track clip corresponding to the first target track clip on the third edit track clip, in response to a text update operation performed on the target text clip on the first target track clip; and
acquiring a target voice clip corresponding to the target text clip, and updating to identify the target voice clip on the third target track clip.
16 . The electronic device according to claim 15 , the operations further comprising:
keeping the second edit track unchanged, in a case where it is detected that a first update time length corresponding to the target text clip on the first edit track is inconsistent with a time length corresponding to the text clip before modification, and displaying a first update track clip corresponding to the first update time length in a preset first candidate region, wherein, the target text clip is identified on the first update track clip;
keeping the second edit track unchanged, in a case where it is detected that a third update time length corresponding to the target voice clip on the third edit track is inconsistent with a time length corresponding to the voice clip before modification, and displaying a third update track clip corresponding to the third update time length in a preset second candidate region, wherein, the target voice clip is identified on the third update track clip.
17 . The electronic device according to claim 15 , the operations further comprising:
adjusting a length of the first target track clip according to a first update time, in a case where it is detected that a first update time length corresponding to the target text clip on the first edit track is inconsistent with a time length corresponding to the text clip before modification;
adjusting a length of the third target track clip according to a third update time, in a case where it is detected that a third update time length corresponding to the target voice clip on the third edit track is inconsistent with a time length corresponding to the voice clip before modification;
correspondingly adjusting a length of a second target track clip corresponding to the first target track clip and the third target track clip on the second edit track, so that timelines of the adjusted first target track clip, the adjusted second target track clip, and the adjusted third target track clip are aligned.
18 . The electronic device according to claim 14 , the operations further comprising:
determining a second target track clip corresponding to the first target track clip is on the second edit track, in response to a text update operation performed on the target text clip on the first target track clip;
acquiring a target video image clip matched with the target text clip, and updating to identify the target video image clip on the second target track clip.
19 . A non-transitory computer readable storage medium, having a computer program stored therein, wherein the computer program, when executed by a processor, causes the processor to implement operations comprising:
receiving text information input by a user via a first interface, wherein the first interface comprises a control for generating a video;
splitting the text information into a plurality of text clips;
in response to receiving a selection of the control, generating a plurality of multimedia clips of the video based on the text information, wherein each of the plurality of multimedia clips corresponds to a different one of the plurality of text clips, and wherein each of the plurality of multimedia clips comprises a voice clip and a video image clip associated with the corresponding text clip;
displaying a second interface configured to perform editing operations on at least a portion of the generated plurality of multimedia clips, wherein the second interface comprises a first region and a second region;
displaying a first edit track, a second edit track, and a third edit track in the first region of the second interface, and simultaneously displaying a plurality of first track clips in the first edit track, a plurality of second track clips in the second edit track, and a plurality of third track clips in the third edit track, wherein the plurality of first track clips are configured to display the plurality of text clips, wherein the plurality of second track clips are configured to display video image clips corresponding to the plurality of text clips, wherein the plurality of third track clips are configured to display voice clips corresponding to the plurality of text clips, and wherein a timeline of the plurality of first track clips in the first edit track is aligned with a timeline of the plurality of second track clips in the second edit track and a timeline of the plurality of third track clips in the third edit track to respectively identify a text clip, a video image clip corresponding to the text clip, and a voice clip corresponding to the text clip;
determining a type of a clip selected from the plurality of text clips, the video image clips, and the voice clips displayed in the first region of the second interface; and
displaying a set of editing controls corresponding to the type of the selected clip in the second region of the second interface, wherein the second region of the second interface is configured to display one of a plurality sets of editing controls, and the plurality of sets of editing controls correspond to a plurality of clip types.