IP Library › Granted Patent US 12,518,110
Granted Patent B2
US 12,518,110 · App. 17/588,025 · Granted Jan 6, 2026

Text generation method and apparatus

Inventors: Xuming Lin (Hangzhou, CN); Zhongzhou Zhao (Hangzhou, CN); Ji Zhang (Hangzhou, CN); Liming Pu (Hangzhou, CN); Jiashuo Zhang (Hangzhou, CN)
Assignee: Alibaba Group Holding Limited
G06F40/40G06F40/166G06F40/205G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,110
App. No.
17/588,025
Granted
Jan 6, 2026
Kind
B2
Abstract

A method including acquiring source data related to an object; acquiring one or more pieces of source data related to the object; analyzing the source data to obtain one or more pieces of material information; parsing the material information to obtain one or more pieces of corresponding text paragraph information; and generating the text describing the object using the text paragraph information. Using the techniques described herein, users comprehensively understand the object according to the generated text directly without having to conduct a large number of searches.

Claims (61)

1 . A method used in live broadcasting, the method comprising:

acquiring one or more pieces of source data related to an object to be presented in a live broadcast;

analyzing the source data to obtain one or more pieces of material information;

dividing the one or more pieces of material information into two parts according to functions, wherein a first part of the one or more pieces of material information is used for providing text paragraph information, and a second part of the one or more pieces of material information is used for creating a text output frame, the text output frame being used for determining an output sequence of the text paragraph information;

parsing the first part of the one or more pieces of material information to obtain one or more pieces of corresponding text paragraph information;

creating the text output frame according to at least one of respective occurrence frequencies or respective search popularities of attributes included in the second part of the one or more pieces of material information;

generating a text describing the object based at least in part on the text output frame and the second part of the one or more pieces of corresponding text paragraph information using an encoder-decoder model, priorities of the attributes in the text output frame being positively correlated with the at least one of the respective occurrence frequencies or the respective search popularities of the attributes, wherein the attributes include an appearance characteristic, a function characteristic, a user characteristic, and an application scenario characteristic of the object, and generating the text describing the object comprises:

inputting a name of the object, attribute information of the object, and the one or more pieces of corresponding text paragraph information into the encoder-decoder model to obtain the text; and

enabling the text to be broadcast to a user through a voice playback on a user interface of an electronic device of the user when the object is presented in the live broadcast and a playback button presented on the user interface is clicked by the user.

2 . The method according to claim 1 , further comprising:

selecting at least one piece of the one or more pieces of material information and at least one piece of the one or more pieces of corresponding text paragraph information;

converting the at least one piece of text paragraph information into voice information;

combining the at least one piece of material information and the voice information into demo data; and

presenting the demo data on the user interface.

3 . The method according to claim 1 , wherein generating the text describing the object using the second part of the one or more pieces of corresponding text paragraph information further comprises:

determining, based on the text output frame, an output sequence of the second part of the one or more pieces of corresponding text paragraph information to generate the text.

4 . The method according to claim 3 , further comprising:

outputting the text paragraph information according to the output sequence.

5 . The method according to claim 3 , further comprising:

adjusting the text output frame based on an external input.

6 . The method according to claim 1 , wherein analyzing the source data to obtain the one or more pieces of material information comprises:

analyzing the attributes of the source data.

7 . The method according to claim 6 , further comprising:

calculating a similarity level between materials having different attributes; and

combining attributes corresponding to materials having a similarity level greater than a threshold into one attribute.

8 . The method according to claim 1 , further comprising:

adjusting the text output frame based on at least one of an audio material or a video material.

9 . The method according to claim 1 , wherein the source data comprises data in multiple modalities.

10 . The method according to claim 9 , wherein the source data comprises pictures, texts, images, sound, and a combination thereof.

11 . The method according to claim 1 , wherein analyzing the source data to obtain the one or more pieces of material information comprises at least one of:

identifying text information in a picture to obtain first material information comprising the picture and the text information;

obtaining, based on a first text, second material information comprising a second text used for describing the object; or

analyzing audiovisual data to obtain third material information comprising at least one of video data, a speech recognition result of the audiovisual data, or audiovisual analysis data.

12 . The method according to claim 1 , wherein the material information comprises at least one of a text material, an audio material, a picture material, or a video material.

13 . The method according to claim 12 , wherein parsing the part part of the one or more pieces of material information to obtain the one or more pieces of corresponding text paragraph information comprises at least one of:

performing speech recognition on the audio material to obtain text paragraph information corresponding to the audio material; or

performing semantic understanding on at least one of the picture material or the video material to obtain text paragraph information corresponding to the at least one of the picture material or the video material.

14 . The method according to claim 1 , wherein the material information comprises:

unprocessed source data; and

processed data having undergone preset analysis processing.

15 . An apparatus used in live broadcasting, the apparatus comprising:

one or more processors; and

one or more memories storing thereon computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform acts comprising:

acquiring one or more pieces of source data related to an object to be presented in a live broadcast;

analyzing the source data to obtain one or more pieces of material information;

dividing the one or more pieces of material information into two parts according to functions, wherein a first part of the one or more pieces of material information is used for providing text paragraph information, and a second part of the one or more pieces of material information is used for creating a text output frame, the text output frame being used for determining an output sequence of the text paragraph information;

parsing the first part of the one or more pieces of material information to obtain one or more pieces of corresponding text paragraph information;

creating the text output frame according to at least one of respective occurrence frequencies or respective search popularities of attributes included in the second part of the one or more pieces of material information;

generating a text describing the object based at least in part on the text output frame and the second part of the one or more pieces of corresponding text paragraph information using an encoder-decoder model, priorities of the attributes in the text output frame being positively correlated with the at least one of the respective occurrence frequencies or the respective search popularities of the attributes, wherein the attributes include an appearance characteristic, a function characteristic, a user characteristic, and an application scenario characteristic of the object, and generating the text describing the object comprises:

inputting a name of the object, attribute information of the object, and the one or more pieces of corresponding text paragraph information into the encoder-decoder model to obtain the text; and

enabling the text to be broadcast through a voice playback on a user interface when the object is presented in the live broadcast.

16 . One or more memories storing thereon computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform acts comprising:

acquiring one or more pieces of source data related to an object to be presented in a live broadcast;

analyzing the source data to obtain one or more pieces of material information;

dividing the one or more pieces of material information into two parts according to functions, wherein a first part of the one or more pieces of material information is used for providing text paragraph information, and a second part of the one or more pieces of material information is used for creating a text output frame, the text output frame being used for determining an output sequence of the text paragraph information;

parsing the first part of the one or more pieces of material information to obtain one or more pieces of corresponding text paragraph information;

creating the text output frame according to at least one of respective occurrence frequencies or respective search popularities of attributes included in the second part of the one or more pieces of material information;

generating a text describing the object based at least in part on the text output frame and the second part of the one or more pieces of corresponding text paragraph information using an encoder-decoder model, priorities of the attributes in the text output frame being positively correlated with the at least one of the respective occurrence frequencies or the respective search popularities of the attributes, wherein the attributes include an appearance characteristic, a function characteristic, a user characteristic, and an application scenario characteristic of the object, and generating the text describing the object comprises:

inputting a name of the object, attribute information of the object, and the one or more pieces of corresponding text paragraph information into the encoder-decoder model to obtain the text; and

enabling the text to be broadcast through a voice playback on a user interface when the object is presented in the live broadcast.

17 . The one or more memories according to claim 16 , wherein the source data comprises association relationships between contents of the live broadcast at different times and respective indicator data at the different times.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2025
From: LIN, XUMING; ZHAO, ZHONGZHOU; ZHANG, JI; PU, LIMING; ZHANG, JIASHUO
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 070280/0856 →
Continuity (1)
Related Publication 20220253613A1 · Aug 11, 2022
References Cited (28)
US 11250947B2 · Divine et al. · 2022 [cited by applicant]
US 11756093B2 · Aher · 2023 [cited by examiner]
US 11861674B1 · Huang · 2024 [cited by examiner]
US 20070078889A1 · Hoskinson · 2007 [cited by examiner]
US 20120215661A1 · Baran · 2012 [cited by applicant]
US 20190180733A1 · Saito · 2019 [cited by examiner]
US 20200042646A1 · Nagaraja · 2020 [cited by examiner]
US 20200074013A1 · Chen · 2020 [cited by examiner]
US 20200311158A1 · Huet · 2020 [cited by examiner]
US 20210118035A1 · Misawa · 2021 [cited by examiner]
US 20210150546A1 · Zhu · 2021 [cited by examiner]
US 20210174031A1 · Wang · 2021 [cited by examiner]
US 20210174784A1 · Min · 2021 [cited by examiner]
US 20210406993A1 · Sethi · 2021 [cited by examiner]
US 20220114349A1 · Sollami · 2022 [cited by examiner]
CN 103914151 · 2014 [cited by applicant]
CN 105528452 · 2016 [cited by applicant]
CN 109584013 · 2019 [cited by applicant]
CN 110909130A · 2020 [cited by examiner]
CN 111368562A · 2020 [cited by applicant]
CN 112199970 · 2021 [cited by applicant]
CN 112232052A · 2021 [cited by applicant]
CN 112287168A · 2021 [cited by applicant]
WO WO2021142999A1 · 2021 [cited by examiner]
Cho et al. “Describing Multimedia Content using Attention-based Encoder-Decoder Networks”. arXiv:1507.01053v1 [cs.NE] Jul. 4, 2015 (Year: 2015). [cited by examiner]
First Examination Report for CN Application CN111368562A, dated Nov. 20, 2024. [cited by applicant]
First Search Report for CN Application CN111368562A, dated Nov. 20, 2024. [cited by applicant]
Translation of CN Office Action mailed Jun. 12, 2025, 11 pages. [cited by applicant]