IP Library Granted Patent US 12,657,940
Granted Patent B2
US 12,657,940 · App. 17/853,907 · Granted Jun 16, 2026

Image paragraph generator

Inventors: Yujia Xie (Redmond, WA); Lu Yuan (Redmond, WA); Nguyen Hung Bach (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC.
G06V20/70G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,940
App. No.
17/853,907
Granted
Jun 16, 2026
Kind
B2
Abstract

Example solutions for image paragraph captioning use a first vision language model to generate visual information (comprising text) for an image. The visual information may include tags, an initial image caption, and information on objects within the image (e.g., further tags and captions, and object attributes and locations within the image). In some examples, the visual information further includes visual clues. A generative language model generates a plurality of image story caption candidates (e.g., descriptive paragraphs) from the visual information. A second vision language model evaluates the plurality of image story caption candidates and selects a caption as the final output caption.

Claims (69)

1 . A system comprising:

a processor; and

a computer storage medium storing instructions that are operative upon execution by the processor to:

receive, by a first vision language model, a request for a story caption of an image, the request comprising a caption focus for the image, the caption focus comprising instructions for a format of the story caption, the format being based on a platform the requested story caption is presented in;

generate, by the first vision language model, for the image, visual information comprising an image tag and a visual clue constructed from the image tag;

based on at least the caption focus, generate, by a generative language model, from the visual information, a plurality of image story caption candidates, wherein each of the plurality of image story caption candidates comprise a coherent paragraph of natural-language text describing the image;

for each image story caption candidate of the plurality of image story caption candidates, perform, by a second vision language model, the following:

parse the image story caption candidate into a plurality of sentences;

for each sentence of the plurality of sentences, measure a similarity between the image and each sentence of the plurality of sentences;

remove any sentence from the plurality of sentences with a similarity score below a similarity threshold; and

update the image storage caption candidate based on the removing, wherein the updated image story caption comprises an updated coherent paragraph of natural-language text describing the image without any of the sentences having a similarity score below the similarity threshold; and

based on at least an evaluation of the image and the plurality of image story caption candidates by the second vision language model, select a caption from among the plurality of image story caption candidates.

2 . The system of claim 1 , wherein the platform comprises one of the following: visual storytelling platform, advertisement generation platform, and social media platform.

3 . The system of claim 1 , wherein the format defines a style the story caption is to be applied to.

4 . The system of claim 1 , wherein the instructions are further operative to: based on at least an initial image caption of the visual information and object information of the visual information, generate the visual clue.

5 . The system of claim 1 , wherein the instructions are further operative to:

detect, by an object detector, a plurality of objects in the image; and

determine, for each object in the plurality of objects in the image, object information, wherein the object information comprises at least one item selected from the list consisting of: an object tag, an object caption, an object attribute, and an object location.

6 . The system of claim 1 , wherein the instructions are further operative to:

generate, by a first captioner, for the image, captions of the visual information, and wherein the visual information comprises at least one item selected from the list consisting of:

an image tag, an initial image caption, and object information.

7 . The system of claim 1 , wherein any sentence from the plurality of sentences with a similarity score below the similarity threshold is identified as sentence that includes a hallucination.

8 . A computerized method comprising:

receiving, by a first vision language model, a request for a story caption of an image, the request comprising a caption focus for the image, the caption focus comprising instructions for a format of the story caption, the format being based on a platform the requested story caption is presented in;

generating, by first vision language model, for the image, visual information comprising an image tag and a visual clue constructed from the image tag;

based on at least the caption focus, generating, by a generative language model, from the visual information, a plurality of image story caption candidates, wherein each of the plurality of image story caption candidates comprise a coherent paragraph of natural-language text describing the image;

for each image story caption candidate of the plurality of image story caption candidates, perform, by a second vision language model, the following:

parse the image story caption candidate into a plurality of sentences;

for each sentence of the plurality of sentences, measure a similarity between the image and each sentence of the plurality of sentences;

remove any sentence from the plurality of sentences with a similarity score below a similarity threshold; and

update the image storage caption candidate based on the removing, wherein the updated image story caption comprises an updated coherent paragraph of natural-language text describing the image without any of the sentences having a similarity score below the similarity threshold; and

based on at least an evaluation of the image and the plurality of image story caption candidates by the second vision language model, selecting a caption from among the plurality of image story caption candidates.

9 . The method of claim 8 , further comprising:

pairing the image with the selected caption.

10 . The method of claim 8 , further comprising:

based on at least the image tag of the visual information, generating the visual clue.

11 . The method of claim 8 , further comprising:

based on at least an initial image caption of the visual information and object information of the visual information, generating the visual clue.

12 . The method of claim 8 , further comprising:

detecting, by an object detector, a plurality of objects in the image; and

determining, for each object in the plurality of objects in the image, object information, wherein the object information comprises at least one item selected from the list consisting of:

an object tag, an object caption, an object attribute, and an object location.

13 . The method of claim 8 , further comprising:

generating, by a first captioner, for the image, captions of the visual information, and wherein the visual information comprises at least one item selected from the list consisting of:

an image tag, an initial image caption, and object information.

14 . The method of claim 8 , wherein the first vision language model and the second vision language model both comprise a common vision language model.

15 . One or more computer storage media having computer-executable instructions stored thereon, which, on execution by a computer, cause the computer to perform operations comprising:

receiving a request for a story caption of an image, the request comprising a caption focus for the image, the caption focus comprising instructions for a format of the story caption, the format being based on a platform the requested story caption is presented in;

generating, for the image, visual information comprising an image tag and a visual clue constructed from the image tag;

based on at least the caption focus, generating, from the visual information, a plurality of image story caption candidates, wherein each of the plurality of image story caption candidates comprise a coherent paragraph of natural-language text describing the image;

for each image story caption candidate of the plurality of image story caption candidates, perform, by a second vision language model, the following:

parse the image story caption candidate into a plurality of sentences;

for each sentence of the plurality of sentences, measure a similarity between the image and each sentence of the plurality of sentences;

remove any sentence from the plurality of sentences with a similarity score below a similarity threshold; and

update the image storage caption candidate based on the removing, wherein the updated image story caption comprises an updated coherent paragraph of natural-language text describing the image without any of the sentences having a similarity score below the similarity threshold; and

based on at least an evaluation of the image and the plurality of image story caption candidates by the second vision language model, selecting a caption from among the plurality of image story caption candidates.

16 . The one or more computer storage media of claim 15 , wherein the operations further comprise:

pairing the image with the selected caption.

17 . The one or more computer storage media of claim 15 , wherein the operations further comprise:

based on at least the image tag of the visual information, generating the visual clue.

18 . The one or more computer storage media of claim 15 , wherein the operations further comprise:

based on at least an initial image caption of the visual information and object information of the visual information, generating the visual clue.

19 . The one or more computer storage media of claim 15 , wherein the operations further comprise:

detecting, by an object detector, a plurality of objects in the image; and

determining, for each object in the plurality of objects in the image, object information, wherein the object information comprises at least one item selected from the list consisting of:

an object tag, an object caption, an object attribute, and an object location.

20 . The one or more computer storage media of claim 15 , wherein the operations further comprise:

generating, by a first captioner, for the image, captions of the visual information, and wherein the visual information comprises at least one item selected from the list consisting of:

an image tag, an initial image caption, and object information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2022
From: XIE, YUJIA; YUAN, LU; BACH, NGUYEN HUNG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061219/0572 →
Continuity (2)
Provisional Application 63347997 · Jun 1, 2022
Related Publication 20230394855A1 · Dec 7, 2023
References Cited (81)
US 10467274B1 · Ren · 2019 [cited by examiner]
US 10503738B2 · Jhamtani · 2019 [cited by examiner]
US 10592751B2 · Chen · 2020 [cited by examiner]
US 11120268B2 · Gupta et al. · 2021 [cited by applicant]
US 11847424B1 · Harkous · 2023 [cited by examiner]
US 20170061250A1 · Gao · 2017 [cited by examiner]
US 20170115853A1 · Allekotte · 2017 [cited by examiner]
US 20170132821A1 · Valliani · 2017 [cited by examiner]
US 20170200066A1 · Wang · 2017 [cited by examiner]
US 20180150444A1 · Kasina · 2018 [cited by examiner]
US 20180173681A1 · Dedhia et al. · 2018 [cited by applicant]
US 20190200050A1 · Zhang · 2019 [cited by examiner]
US 20210064879A1 · Gupta · 2021 [cited by examiner]
US 20210117723A1 · Kim et al. · 2021 [cited by applicant]
US 20210224310A1 · Zhelezniakov · 2021 [cited by examiner]
US 20220114361A1 · Kale et al. · 2022 [cited by applicant]
US 20230103340A1 · Gao · 2023 [cited by examiner]
US 20230229288A1 · Sicora · 2023 [cited by examiner]
US 20230316803A1 · Kelkar · 2023 [cited by examiner]
US 20230377031A1 · Khalaf · 2023 [cited by examiner]
Natasha Lomas, Hypotenuse AI wants to take the strain out of copywriting for e-commerce, Aug. 7, 2020 (Year: 2020). [cited by examiner]
Zou, et al., “Object Detection in 20 Years: A Survey”, In Repository of arXiv:1905.05055v1, May 13, 2019, 40 Pages. [cited by applicant]
Xu, et al., “Interactive Key-Value Memory-augmented Attention for Image Paragraph Captioning”, In the Proceedings of the 28th International Conference on Computational Linguistics, Dec. 8, 2020, pp. 3132-3142. [cited by applicant]
Yang, et al., “An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA”, In Repository of arXiv:2109.05014v1, Sep. 10, 2021, 10 Pages. [cited by applicant]
Yu, et al., “CoCa: Contrastive Captioners are Image-Text Foundation Models”, In Repository of arXiv:2205.01917v1, May 4, 2022, 19 Pages. [cited by applicant]
Yuan, et al., “Florence: A New Foundation Model for Computer Vision”, In Repository of arXiv:2111.11432v1, Nov. 22, 2021, 17 Pages. [cited by applicant]
Zeng, et al., “Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language”, In Repository of arXiv:2204.00598v1, Apr. 1, 2022, 20 Pages. [cited by applicant]
Zhou, et al., “Unified Vision-Language Pre-Training for Image Captioning and VQA”, In the Proceedings of The Thirty-Fourth AAAI Conference on Artificial Intelligence, Feb. 7, 2020, pp. 13041-13049. [cited by applicant]
Zhu, et al., “Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks”, In Repository of arXiv:2112.01522v1, Dec. 2, 2021, 15 Pages. [cited by applicant]
Alayrac, et al., “Flamingo: a Visual Language Model for Few-Shot Learning”, In Repository of arXiv:2204.14198v1, Apr. 29, 2022, 66 Pages. [cited by applicant]
Anderson, et al., “SPICE: Semantic Propositional Image Caption Evaluation”, In Proceedings of European conference on computer vision, Oct. 11, 2016, pp. 382-398. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners”, In the Proceedings of the 34th International Conference on Neural Information Processing Systems, Dec. 6, 2020, 25 Pages. [cited by applicant]
Chatterjee, et al., “Diverse and Coherent Paragraph Generation from Images”, In Proceedings of the European Conference on Computer Vision, Sep. 8, 2018, pp. 747-763. [cited by applicant]
Chen, et al., “Microsoft COCO Captions: Data Collection and Evaluation Server”, In Repository of arXiv:1504.00325v1, Apr. 1, 2015, 7 Pages. [cited by applicant]
Cho, et al., “Unifying Vision-and-Language Tasks via Text Generation”, In the Proceedings of the 38th International Conference on Machine Learning, vol. 139, Jul. 18, 2021, 12 Pages. [cited by applicant]
Dai, et al., “Dynamic Head: Unifying Object Detection Heads with Attentions”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2021, pp. 7369-7378. [cited by applicant]
Dai, et al., “Towards Diverse and Natural Image Descriptions via a Conditional GAN”, In Proceedings of the IEEE International conference on computer vision, Oct. 22, 2017, pp. 2989-2998. [cited by applicant]
Denkowski, et al., “Meteor Universal: Language Specific Translation Evaluation for Any Target Language”, In Proceedings of the ninth workshop on statistical machine translation, Jun. 26, 2014, pp. 376-380. [cited by applicant]
Eichenberg, et al., “MAGMA—Multimodal Augmentation of Generative Models through Adapter-based Finetuning”, In Repository of arXiv:2112.05253v1, Dec. 9, 2021, pp. 1-11. [cited by applicant]
Fei-Fei, et al., “Searching for Computer Vision North Stars”, In Dædalus, the Journal of the American Academy of Arts & Sciences, vol. 151, Issue 2, May 1, 2022, pp. 85-99. [cited by applicant]
Guo, et al., “Matching Visual Features to Hierarchical Semantic Topics for Image Paragraph Captioning”, In Repository of arXiv:2105.04143v1, May 10, 2021, 20 Pages. [cited by applicant]
Huang, et al., “Visual Storytelling”, In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 2016, pp. 1233-1239. [cited by applicant]
Jia, et al., “Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision”, In the Proceedings of the 38th International Conference on Machine Learning, vol. 139, Jul. 18, 2021, 13 Pages. [cited by applicant]
Jocher, et al., “ultralytics/yolov5: v3.1—Bug Fixes and Performance Improvements”, Retrieved From: https://zenodo.org/record/4154370#.Yr2RwRVByUk, Oct. 29, 2020, 6 Pages. [cited by applicant]
Johnson, et al., “Image Retrieval using Scene Graphs”, In the Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 7, 2015, pp. 3668-3678. [cited by applicant]
Jyoshna, et al., “Image Paragraph Captioning Using Deep Learning and NPL Techniques”, A Project report submitted in partial fulfillment of the requirements for the award of the degree of Bachelor of Technology in Comput… [cited by applicant]
Klein, et al., “Accurate Unlexicalized Parsing”, In the Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics, Jul. 2003, 8 Pages. [cited by applicant]
Krause, et al., “A Hierarchical Approach for Generating Descriptive Image Paragraphs”, In the Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition, Jul. 21, 2017, pp. 3337-3345. [cited by applicant]
Krishna, et al., “Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations”, In the International Journal of Computer Vision, vol. 123, Issue 1, May 2017, pp. 32-73. [cited by applicant]
Li, et al., “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation”, In Repository of arXiv:2201.12086v1, Jan. 28, 2022, 12 Pages. [cited by applicant]
Li, et al., “DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset”, In Repository of arXiv:1710.03957v1, Oct. 11, 2017, 10 Pages. [cited by applicant]
Liang, et al., “Recurrent Topic-Transition GAN for Visual Paragraph Generation”, In the Proceedings of the IEEE International conference on computer vision, Oct. 22, 2017, pp. 3382-3391. [cited by applicant]
Lin, Chin-Yew, “ROUGE: A Package for Automatic Evaluation of Summaries”, In the Journal of Association for Computational Linguistics, Jul. 2004, 8 Pages. [cited by applicant]
Liu, et al., “Beyond Narrative Description: Generating Poetry from Images by Multi-Adversarial Training”, In the Proceedings of the 26th ACM international conference on Multimedia, Oct. 22, 2018, pp. 783-791. [cited by applicant]
Luo, et al., “A thorough review of models, evaluation metrics, and datasets on image captioning”, Retrieved From: https://ietresearch.onlinelibrary.wiley.com/doi/full/10.1049/ipr2.12367, Nov. 22, 2021, pp. 311-332. [cited by applicant]
Luo, et al., “Curiosity-driven Reinforcement Learning for Diverse Visual Paragraph Generation”, In the Proceedings of the 27th ACM International Conference on Multimedia, Oct. 21, 2019, pp. 2341-2350. [cited by applicant]
Mao, et al., “Show and Tell More: Topic-Oriented Multi-Sentence Image Captioning”, In the Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, Jul. 13, 2018, pp. 4258-4264. [cited by applicant]
Marr, David, “Vision: A Computational Investigation into the Human Representation and Processing of Visual Information”, Published in MIT Press, 2010, 41 Pages. [cited by applicant]
Melas-Kyriazi, et al., “Training for Diversity in Image Paragraph Captioning”, In the Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Oct. 31, 2018, pp. 757-761. [cited by applicant]
Miller, Georgea. , “WordNet: a lexical database for English”, In the Journal of Communications of the ACM, vol. 38, Issue 11, Nov. 1995, pp. 39-41. [cited by applicant]
Nikolaus, et al., “Compositional Generalization in Image Captioning”, In Repository of arXiv:1909.04402v1, Sep. 10, 2019, 16 Pages. [cited by applicant]
Papineni, et al., “Bleu: a Method for Automatic Evaluation of Machine Translation”, In the Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Jul. 2002, 8 Pages. [cited by applicant]
Radford, et al., “Learning Transferable Visual Models From Natural Language Supervision”, In the Proceedings of the 38th International Conference on Machine Learning, vol. 139, Jul. 18, 2021, 16 Pages. [cited by applicant]
Rajpurkar, et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text”, In the Repository of arXiv:1606.05250v1, Jun. 16, 2016, 10 Pages. [cited by applicant]
Salvador, et al., “Inverse Cooking: Recipe Generation from Food Images”, In the Proceedings of 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 10445-10454. [cited by applicant]
Schuster, et al., “Generating Semantically Precise Scene Graphs from Textual Descriptions for Improved Image Retrieval”, In the Proceedings of the Fourth Workshop on Vision and Language, Sep. 18, 2015, pp. 70-80. [cited by applicant]
Shah, Chintan, “Image Captioning: Generating Stories from Unstructured Data Using Applied NLG”, Retrieved From: https://www.dataversity.net/image-captioning-generating-stories-from-unstructured-data-using-applied-nlg/, … [cited by applicant]
Shi, et al., “S2TD: A Tree-Structured Decoder for Image Paragraph Captioning”, In the Proceedings of ACM Multimedia Asia, Dec. 1, 2021, 7 Pages. [cited by applicant]
Singh, “FLAVA: A Foundational Language And Vision Alignment Model”, In Repository of arXiv:2112.04482v1, Dec. 8, 2021, 17 Pages. [cited by applicant]
Su, et al., “Language Models Can See: Plugging Visual Controls in Text Generation”, In Repository of arXiv:2205.02655v1, May 5, 2022, 20 Pages. [cited by applicant]
Sutskever, et al., “CLIP: Connecting Text and Images”, Retrieved From: https://openai.com/blog/clip/, Jan. 5, 2021, 16 Pages. [cited by applicant]
Thrush, et al., “Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality”, In the Proceedings of The IEEE / CVF Computer Vision and Pattern Recognition Conference, Jun. 19, 2022, pp. 5238-52… [cited by applicant]
Vedantam, et al., “CIDEr: Consensus-based image description evaluation”, In the Proceedings of 2015 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 7, 2015, pp. 4566-4575. [cited by applicant]
Wang, et al., “Convolutional Auto-encoding of Sentence Topics for Image Paragraph Generation”, In Repository of arXiv:1908.00249v1, Aug. 1, 2019, 7 Pages. [cited by applicant]
Xie, et al., “Visual Clues: Bridging Vision and Language Foundations for Image Paragraph Captioning”, In Repository of arXiv:2206.01843v1, Jun. 3, 2022, 20 Pages. [cited by applicant]
Wang, Zirui, “SimVLM: Simple Visual Language Model Pre-training with Weak Supervision”, Retrieved From: https://ai.googleblog.com/2021/10/simvlm-simple-visual-language-model-pre.html, Oct. 15, 2021, 5 Pages. [cited by applicant]
Wang, et al., “Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework”, In Repository of arXiv:2202.03052v1, Feb. 7, 2022, 23 Pages. [cited by applicant]
Wu, et al., “Image Captioning with an Intermediate Attributes Layer”, In Repository of arXiv:1506.01144v2, Jun. 6, 2015, 11 Pages. [cited by applicant]
Xiao, et al., “Azure AI milestone: New foundation model Florence v1.0 advances state of the art, topping popular computer vision leaderboards”, Retrieved From: https://www.microsoft.com/en-us/research/blog/azure-ai-mile… [cited by applicant]
Cho, et al., “Unifying Vision-and-Language Tasks via Text Generation”, In Repository of arXiv:2102.02779v1, Feb. 4, 2021, 16 Pages. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/018767”, Mailed Date: Jul. 25, 2023, 10 Pages. [cited by applicant]