IP Library Granted Patent US 12,198,430
Granted Patent B1
US 12,198,430 · App. 17/009,542 · Granted Jan 14, 2025

Multimodal state tracking via scene graphs for assistant systems

Inventor: Satwik Kottur (Menlo Park, CA)
Assignee: Meta Platforms, Inc.
G06V20/35G06F16/9024G06F16/9536G06V20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,430
App. No.
17/009,542
Granted
Jan 14, 2025
Kind
B1
Abstract

In one embodiment, a method includes receiving, from a client system associated with a user, a first user request that includes a reference to a target object and one or more of an attribute or a relationship of the target object. Visual data including one or more images portraying the target object may then be accessed, and the reference may be resolved to the target object portrayed in the one or more images. Object information of the target object that corresponds to the referenced attribute or relationship of the first user request may be determined based on a visual analysis of the one or more images. Finally, responsive to receiving the first user request, the object information of the target object may be stored in a multimodal dialog state.

Claims (80)

1. A method comprising:

receiving, from a client system associated with a user, a first user request comprising a reference to a target object and one or more of an attribute or a relationship of the target object;

accessing, from the client system, visual data captured by one or more cameras of the client system, wherein the visual data comprises images portraying the target object;

resolving, based on a multimodal dialog state, the reference to the target object portrayed in the images captured by the one or more cameras of the client system;

determining that first object information, corresponding to the attribute or the relationship, is not already stored in the multimodal dialog state;

responsive to resolving the reference to the target object and the first object information not being stored in the multimodal dialog state,

executing a first visual analysis by retrieving one or more of the images, and identifying the first object information from the one or more of the images based on an identification of the target object in the one or more of the images and by analyzing the one or more of the images, and

halting provisioning of other images of the images in response to the first object information being identified from the one or more of the images; and

storing, responsive to receiving the first user request, the first object information of the target object as an incremental update in the multimodal dialog state.

2. The method of claim 1 , further comprising:

sending, to the client system, instructions for presenting a response to the first user request, wherein the response comprises the first object information.

3. The method of claim 1 , wherein the first object information is stored as a scene graph in the multimodal dialog state.

4. The method of claim 1 , further comprising:

receiving, from the client system, the visual data;

analyzing, by a computer vision module, the visual data to identify one or more objects portrayed in the one or more images captured by the one or more cameras of the client system; and

assigning respective object identifiers to one or more of the identified one or more objects.

5. The method of claim 4 , wherein resolving the reference to the target object comprises:

selecting, from among the respective object identifiers of the one or more identified objects, a target object identifier corresponding to the target object.

6. The method of claim 5 , wherein determining the first object information of the target object comprises:

providing, in response to receiving the first user request, the visual data and the target object identifier of the target object to a scene understanding engine; and

generating, by the scene understanding engine, the first object information based on the target object identifier.

7. The method of claim 6 , further comprising:

generating, by a dialog engine, a partial scene graph consisting of the target object identifier and the first object information of the target object;

receiving, from the client system, a second user request comprising another reference to the target object;

resolving, based on the multimodal dialog state, the another reference to the target object, wherein a second object information about the target object is not already stored in the multimodal dialog state; and

responsive to resolving the another reference to the target object, executing a second visual analysis of the one or more of the images to determine the second object information, and storing the second object information as an incremental update to the partial scene graph,

wherein storing the first object information of the target object comprises storing the partial scene graph in the multimodal dialog state.

8. The method of claim 6 , further comprising:

in response to the generation of the first object information, halting provision of the visual data to the scene understanding engine.

9. The method of claim 1 , wherein resolving the reference to the target object comprises:

generating, by a dialog engine, an initial scene graph comprising the target object and one or more additional objects portrayed in the one or more images captured by the one or more cameras of the client system; and

resolving the reference to the target object based on the initial scene graph.

10. The method of claim 9 , further comprising:

selecting a portion of the initial scene graph comprising the target object and the first object information of the target object from among a plurality of portions of the initial scene graph, wherein remaining portions of the initial scene graph do not comprise the target object and the first object information.

11. The method of claim 10 , wherein storing the first object information of the target object comprises:

storing the selected portion of the initial scene graph in the multimodal dialog state; and

deleting the remaining portions of the initial scene graph.

12. The method of claim 1 , further comprising:

determining, by a scene understanding engine, one or more properties of the target object based on the first visual analysis of the one or more images captured by the one or more cameras of the client system; and

resolving the target object to a specific entity based on the one or more properties.

13. The method of claim 12 , further comprising:

accessing a knowledge graph based on the specific entity; and

retrieving the first object information from the knowledge graph.

14. The method of claim 1 , further comprising:

receiving, from the client system, a second user request comprising a reference to the target object and one or more of an additional attribute or an additional relationship of the target object;

determining a second object information of the target object corresponding to the referenced additional attribute or relationship of the second user request based on a subsequent visual analysis of the one or more images captured by the one or more cameras of the client system; and

storing, responsive to receiving the second user request, the second object information of the target object in the multimodal dialog state.

15. The method of claim 1 , further comprising:

receiving additional visual data, the additional visual data comprising one or more additional images captured by the one or more cameras of the client system portraying a second object sharing an attribute or relationship with the target object.

16. The method of claim 15 , further comprising:

determining whether a second user request referencing the second object has been received; and

responsive to determining whether the second user request referencing the second object has been received:

if the second user request has been received, storing, responsive to receiving the second user request, a second object information of the second object in the multimodal dialog state; and

if the second user request has not been received, storing the additional visual data without storing the second object information of the second object in the multimodal dialog state.

17. The method of claim 15 , further comprising:

receiving, from the client system, a second user request, wherein the second user request comprises a reference to the shared attribute or relationship; and

determining, based on second object information of the second object stored in the multimodal dialog state, that the second user request is associated with the target object.

18. The method of claim 17 , further comprising:

sending, to the client system and in response to the second user request, instructions for presenting a response to the second user request, wherein the response comprises the second object information of the target object.

19. The method of claim 1 , wherein the first user request is received during a current dialog session between the user and an assistant system associated with the client system, and wherein the first object information of the target object is stored during the current dialog session.

20. One or more computer-readable non-transitory storage media embodying software that is operable when executed to:

receive, from a client system associated with a user, a first user request comprising a reference to a target object and one or more of an attribute or a relationship of the target object;

access, from the client system, visual data captured by one or more cameras of the client system, wherein the visual data comprises images portraying the target object;

resolve, based on a multimodal dialog state, the reference to the target object portrayed in the images captured by the one or more cameras of the client system;

determining that first object information, corresponding to the attribute or the relationship, is not already stored in the multimodal dialog state;

responsive to resolving the reference to the target object and the first object information not being stored in the multimodal dialog state,

executing a visual analysis by retrieving one or more of the images, and identifying the first object information from the one or more of the images based on an identification of the target object in the one or more of the images and by analyzing the one or more of the images, and

halting provisioning of other images of the images in response to the first object information being identified from the one or more of the images; and

store, responsive to receiving the first user request, the first object information of the target object as an incremental update in the multimodal dialog state.

21. A system comprising:

one or more processors; and

a non-transitory memory coupled to the processors comprising instructions executable by the processors, the processors operable when executing the instructions to:

receive, from a client system associated with a user, a first user request comprising a reference to a target object and one or more of an attribute or a relationship of the target object;

access, from the client system, visual data captured by one or more cameras of the client system, wherein the visual data comprises images portraying the target object;

resolve, based on a multimodal dialog state, the reference to the target object portrayed in the images captured by the one or more cameras of the client system;

determining that first object information, corresponding to the attribute or the relationship, is not already stored in the multimodal dialog state;

responsive to resolving the reference to the target object and the first object information not being stored in the multimodal dialog state,

executing a visual analysis by retrieving one or more of the images, and identifying the first object information from the one or more of the images based on an identification of the target object in the one or more of the images and by analyzing the one or more of the images, and

halting provisioning of other images of the images in response to the first object information being identified from the one or more of the images; and

store, responsive to receiving the first user request, the first object information of the target object as an incremental update in the multimodal dialog state.

Assignments (2)
CHANGE OF NAME Recorded Dec 20, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058553/0802 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2020
From: KOTTUR, SATWIK
To: FACEBOOK, INC.
Reel/Frame 053792/0181 →
References Cited (210)
US 7124123B1 · Roskind · 2006 [cited by applicant]
US 7158678B2 · Nagel · 2007 [cited by applicant]
US 7397912B2 · Aasman · 2008 [cited by applicant]
US 8027451B2 · Arendsen · 2011 [cited by applicant]
US 8560564B1 · Hoelzle · 2013 [cited by applicant]
US 8677377B2 · Cheyer · 2014 [cited by applicant]
US 8935192B1 · Ventilla · 2015 [cited by applicant]
US 8983383B1 · Haskin · 2015 [cited by applicant]
US 9154739B1 · Nicolaou · 2015 [cited by applicant]
US 9299059B1 · Marra · 2016 [cited by applicant]
US 9304736B1 · Whiteley · 2016 [cited by applicant]
US 9338242B1 · Suchland · 2016 [cited by applicant]
US 9338493B2 · Van Os · 2016 [cited by applicant]
US 9390724B2 · List · 2016 [cited by applicant]
US 9418658B1 · David · 2016 [cited by applicant]
US 9472206B2 · Ady · 2016 [cited by applicant]
US 9479931B2 · Ortiz · 2016 [cited by applicant]
US 9576574B2 · van Os · 2017 [cited by applicant]
US 9659577B1 · Langhammer · 2017 [cited by applicant]
US 9720955B1 · Cao · 2017 [cited by applicant]
US 9747895B1 · Jansche · 2017 [cited by applicant]
US 9792281B2 · Sarikaya · 2017 [cited by applicant]
US 9858925B2 · Gruber · 2018 [cited by applicant]
US 9865260B1 · Vuskovic · 2018 [cited by applicant]
US 9875233B1 · Tomkins · 2018 [cited by applicant]
US 9875741B2 · Gelfenbeyn · 2018 [cited by applicant]
US 9881077B1 · Alfonseca · 2018 [cited by applicant]
US 9886953B2 · Lemay · 2018 [cited by applicant]
US 9990591B2 · Gelfenbeyn · 2018 [cited by applicant]
US 10042032B2 · Scott · 2018 [cited by applicant]
US 10127220B2 · Bellegarda · 2018 [cited by applicant]
US 10134395B2 · Typrin · 2018 [cited by applicant]
US 10199051B2 · Binder · 2019 [cited by applicant]
US 10241752B2 · Lemay · 2019 [cited by applicant]
US 10276170B2 · Gruber · 2019 [cited by applicant]
US 10462422B1 · Harrison · 2019 [cited by applicant]
US 10511808B2 · Harrison · 2019 [cited by applicant]
US 10719786B1 · Treseler · 2020 [cited by applicant]
US 10782986B2 · Martin · 2020 [cited by applicant]
US 20080240379A1 · Maislos · 2008 [cited by applicant]
US 20080300884A1 · Smith · 2008 [cited by applicant]
US 20090282033A1 · Alshawi · 2009 [cited by applicant]
US 20110246383A1 · Gibson · 2011 [cited by applicant]
US 20120245944A1 · Gruber · 2012 [cited by applicant]
US 20120246191A1 · Xiong · 2012 [cited by applicant]
US 20120265528A1 · Gruber · 2012 [cited by applicant]
US 20120311126A1 · Jadallah · 2012 [cited by applicant]
US 20130035930A1 · Ferrucci · 2013 [cited by applicant]
US 20130268839A1 · Lefebvre · 2013 [cited by applicant]
US 20130275138A1 · Gruber · 2013 [cited by applicant]
US 20130275164A1 · Gruber · 2013 [cited by applicant]
US 20140074483A1 · van Os · 2014 [cited by applicant]
US 20140164506A1 · Tesch · 2014 [cited by applicant]
US 20140244712A1 · Walters · 2014 [cited by applicant]
US 20140280017A1 · Indarapu · 2014 [cited by applicant]
US 20140297284A1 · Gruber · 2014 [cited by applicant]
US 20150081674A1 · Ali · 2015 [cited by applicant]
US 20150142420A1 · Sarikaya · 2015 [cited by applicant]
US 20150142704A1 · London · 2015 [cited by applicant]
US 20150169284A1 · Quast · 2015 [cited by applicant]
US 20150169744A1 · Walkingshaw · 2015 [cited by applicant]
US 20150179168A1 · Hakkani-Tur · 2015 [cited by applicant]
US 20150186156A1 · Brown · 2015 [cited by applicant]
US 20150207765A1 · Brantingham · 2015 [cited by applicant]
US 20150347375A1 · Tremblay · 2015 [cited by applicant]
US 20160019290A1 · Ratnaparkhi · 2016 [cited by applicant]
US 20160037311A1 · Cho · 2016 [cited by applicant]
US 20160063118A1 · Campbell · 2016 [cited by applicant]
US 20160196491A1 · Chandrasekaran · 2016 [cited by applicant]
US 20160225370A1 · Kannan · 2016 [cited by applicant]
US 20160255082A1 · Rathod · 2016 [cited by applicant]
US 20160306505A1 · Vigneras · 2016 [cited by applicant]
US 20160308799A1 · Schubert · 2016 [cited by applicant]
US 20160328096A1 · Tran · 2016 [cited by applicant]
US 20160378849A1 · Myslinski · 2016 [cited by applicant]
US 20160378861A1 · Eledath · 2016 [cited by examiner]
US 20170026318A1 · Daniel · 2017 [cited by applicant]
US 20170091168A1 · Bellegarda · 2017 [cited by applicant]
US 20170092264A1 · Hakkani-Tur · 2017 [cited by applicant]
US 20170132019A1 · Karashchuk · 2017 [cited by applicant]
US 20170193390A1 · Weston · 2017 [cited by applicant]
US 20170353469A1 · Selekman · 2017 [cited by applicant]
US 20170358304A1 · Castillo Sanchez · 2017 [cited by applicant]
US 20170359707A1 · Diaconu · 2017 [cited by applicant]
US 20180013699A1 · Sapoznik · 2018 [cited by applicant]
US 20180018562A1 · Jung · 2018 [cited by applicant]
US 20180018987A1 · Zass · 2018 [cited by applicant]
US 20180040020A1 · Kurian · 2018 [cited by applicant]
US 20180054523A1 · Zhang · 2018 [cited by applicant]
US 20180096071A1 · Green · 2018 [cited by applicant]
US 20180096072A1 · He · 2018 [cited by applicant]
US 20180107917A1 · Hewavitharana · 2018 [cited by applicant]
US 20180121508A1 · Halstvedt · 2018 [cited by applicant]
US 20180189629A1 · Yatziv · 2018 [cited by applicant]
US 20180210874A1 · Fuxman · 2018 [cited by applicant]
US 20180293484A1 · Wang · 2018 [cited by applicant]
US 20190080698A1 · Miller · 2019 [cited by applicant]
US 20190087491A1 · Bax · 2019 [cited by applicant]
US 20190139150A1 · Brownhill · 2019 [cited by applicant]
US 20190213490A1 · White · 2019 [cited by applicant]
US 20190324527A1 · Presant · 2019 [cited by applicant]
US 20190324553A1 · Liu · 2019 [cited by applicant]
US 20190324780A1 · Zhu · 2019 [cited by applicant]
US 20190325042A1 · Yu · 2019 [cited by applicant]
US 20190325080A1 · Natarajan · 2019 [cited by applicant]
US 20190325081A1 · Liu · 2019 [cited by applicant]
US 20190325084A1 · Peng · 2019 [cited by applicant]
US 20190327330A1 · Natarajan · 2019 [cited by applicant]
US 20190327331A1 · Natarajan · 2019 [cited by applicant]
US 20190348033A1 · Chen · 2019 [cited by applicant]
US 20190361408A1 · Tokuchi · 2019 [cited by applicant]
US 20210248375A1 · Geng · 2021 [cited by examiner]
WO WO2012116241 · 2012 [cited by applicant]
Niu, Yulei, et al. “Recursive visual attention in visual dialog.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. (Year: 2019). [cited by examiner]
Kottur, Satwik, et al. “Clevr-dialog: A diagnostic dataset for multi-round reasoning in visual dialog.” arXiv preprint arXiv:1903.03166 (2019). (Year: 2019). [cited by examiner]
Guo, Dan, et al. “Iterative context-aware graph inference for visual dialog.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. (Year: 2020). [cited by examiner]
U.S. Appl. No. 15/953,957, filed Apr. 16, 2018, Kemal El Moujahid. [cited by applicant]
U.S. Appl. No. 15/967,193, filed Apr. 30, 2018, Davide Testuggine. [cited by applicant]
U.S. Appl. No. 15/967,279, filed Apr. 30, 2018, Fuchun Peng. [cited by applicant]
U.S. Appl. No. 16/025,317, filed Jul. 2, 2018, Sonal Gupta. [cited by applicant]
U.S. Appl. No. 16/036,827, filed Jul. 16, 2018, Emmanouil Koukoumidis. [cited by applicant]
U.S. Appl. No. 16/038,120, filed Jul. 17, 2018, Jason Schissel. [cited by applicant]
U.S. Appl. No. 16/048,049, filed Jul. 27, 2018, Markku Salkola. [cited by applicant]
U.S. Appl. No. 16/048,072, filed Jul. 27, 2018, Markku Salkola. [cited by applicant]
U.S. Appl. No. 16/048,101, filed Jul. 27, 2018, Markku Salkola. [cited by applicant]
U.S. Appl. No. 16/057,414, filed Aug. 7, 2018, Jeremy Gillmor Kahn. [cited by applicant]
U.S. Appl. No. 16/103,775, filed Aug. 14, 2018, Zheng Zhou. [cited by applicant]
U.S. Appl. No. 16/107,601, filed Aug. 21, 2018, Rajesh Krishna Shenoy. [cited by applicant]
U.S. Appl. No. 16/107,847, filed Aug. 21, 2018, Rajesh Krishna Shenoy. [cited by applicant]
U.S. Appl. No. 16/121,393, filed Sep. 4, 2018, Zheng Zhou. [cited by applicant]
U.S. Appl. No. 16/127,173, filed Sep. 10, 2018, Zheng Zhou. [cited by applicant]
U.S. Appl. No. 16/129,638, filed Sep. 12, 2018, Vivek Natarajan. [cited by applicant]
U.S. Appl. No. 16/135,752, filed Sep. 19, 2018, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/150,184, filed Oct. 2, 2018, Francislav P. Penov. [cited by applicant]
U.S. Appl. No. 16/151,040, filed Oct. 3, 2018, Brian Nelson. [cited by applicant]
U.S. Appl. No. 16/168,536, filed Oct. 23, 2018, Benoit F. Dumoulin. [cited by applicant]
U.S. Appl. No. 16/176,081, filed Oct. 31, 2018, Anusha Balakrishnan. [cited by applicant]
U.S. Appl. No. 16/176,312, filed Oct. 31, 2018, Emmanouil Koukoumidis. [cited by applicant]
U.S. Appl. No. 16/182,542, filed Nov. 6, 2018, Michael Robert Hanson. [cited by applicant]
U.S. Appl. No. 16/183,650, filed Nov. 7, 2018, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/192,538, filed Nov. 15, 2018, Emmanouil Koukoumidis. [cited by applicant]
U.S. Appl. No. 16/222,923, filed Dec. 17, 2018, Jason Schissel. [cited by applicant]
U.S. Appl. No. 16/222,957, filed Dec. 17, 2018, Emmanouil Koukoumidis. [cited by applicant]
U.S. Appl. No. 16/229,828, filed Dec. 21, 2018, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/247,439, filed Jan. 14, 2019, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/264,173, filed Jan. 31, 2019, Ashwini Challa. [cited by applicant]
U.S. Appl. No. 16/376,832, filed Apr. 5, 2019, Honglei Liu. [cited by applicant]
U.S. Appl. No. 16/389,769, filed Apr. 19, 2019, Honglei Liu. [cited by applicant]
U.S. Appl. No. 16/389,634, filed Apr. 19, 2019, Paul Anthony Crook. [cited by applicant]
U.S. Appl. No. 16/389,738, filed Apr. 19, 2019, Fuchun Peng. [cited by applicant]
U.S. Appl. No. 16/389,728, filed Apr. 19, 2019, William Crosby Presant. [cited by applicant]
U.S. Appl. No. 16/434,010, filed Jun. 6, 2019, Sergiu Dogaru. [cited by applicant]
U.S. Appl. No. 16/552,559, filed Aug. 27, 2019, Seungwhan Moon. [cited by applicant]
U.S. Appl. No. 16/557,055, filed Aug. 30, 2019, Seungwhan Moon. [cited by applicant]
U.S. Appl. No. 16/659,070, filed Oct. 21, 2019, Lisa Xiaoyi Huang. [cited by applicant]
U.S. Appl. No. 16/659,203, filed Oct. 21, 2019, Lisa Xiaoyi Huang. [cited by applicant]
U.S. Appl. No. 16/659,363, filed Oct. 21, 2019, Lisa Xiaoyi Huang. [cited by applicant]
U.S. Appl. No. 16/659,419, filed Oct. 21, 2019, Lisa Xiaoyi Huang. [cited by applicant]
U.S. Appl. No. 16/703,700, filed Dec. 4, 2019, Ahmed Aly. [cited by applicant]
U.S. Appl. No. 16/733,044, filed Jan. 2, 2020, Francislav P. Penov. [cited by applicant]
U.S. Appl. No. 16/741,630, filed Jan. 13, 2020, Paul Anthony Crook. [cited by applicant]
U.S. Appl. No. 16/741,642, filed Jan. 13, 2020, Fuchun Peng. [cited by applicant]
U.S. Appl. No. 16/742,769, filed Jan. 14, 2020, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/742,668, filed Jan. 14, 2020, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/790,497, filed Feb. 13, 2020, Yang Gao. [cited by applicant]
U.S. Appl. No. 16/815,960, filed Mar. 11, 2020, Malik. [cited by applicant]
U.S. Appl. No. 16/815,990, filed Mar. 11, 2020, Malik. [cited by applicant]
U.S. Appl. No. 16/842,366, filed Apr. 7, 2020, Kamisetty. [cited by applicant]
U.S. Appl. No. 16/847,155, filed Apr. 13, 2020, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/914,966, filed Jun. 29, 2020, Noam Yakob Behar. [cited by applicant]
U.S. Appl. No. 16/917,664, filed Jun. 30, 2020, Xiaohu Liu. [cited by applicant]
U.S. Appl. No. 16/921,665, filed Jul. 6, 2020, Honglei Liu. [cited by applicant]
U.S. Appl. No. 16/998,423, filed Aug. 20, 2020, Armen Aghajanyan. [cited by applicant]
U.S. Appl. No. 17/006,377, filed Aug. 28, 2020, Shivani Poddar. [cited by applicant]
U.S. Appl. No. 17/006,339, filed Aug. 28, 2020, Shivani Poddar. [cited by applicant]
U.S. Appl. No. 17/006,260, filed Aug. 28, 2020, William Crosby Presant. [cited by applicant]
U.S. Appl. No. 17/035,253, filed Sep. 28, 2020, Piyush Khemka. [cited by applicant]
U.S. Appl. No. 62/660,876, filed Apr. 20, 2018, Anuj Kumar. [cited by applicant]
U.S. Appl. No. 62/675,090, filed May 22, 2018, Michael Robert Hanson. [cited by applicant]
U.S. Appl. No. 62/747,628, filed Oct. 18, 2018, Honglei Liu. [cited by applicant]
U.S. Appl. No. 62/749,608, filed Oct. 23, 2018, Ashwini Challa. [cited by applicant]
U.S. Appl. No. 62/750,746, filed Oct. 25, 2018, Honglei Liu. [cited by applicant]
U.S. Appl. No. 62/923,342, filed Oct. 18, 2019, Michael Robert Hanson. [cited by applicant]
Tepper, Naama, Anat Hashavit, Maya Barnea, Inbal Ronen, and Lior Leiba. “Collabot: Personalized Group Chat Summarization.” In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, pp. 7… [cited by applicant]
Honglei Liu, et al.: Explore-Exploit: A Framework for Interactive and Online Learning, arXiv:1812.00116, Dec. 1, 2018. [cited by applicant]
Chat Extensions, https://developers.facebook.com/docs/messenger-platform/guides/chat-extensions, Apr. 18, 2017. [cited by applicant]
Kottur, Satwik, et al. “Visual coreference resolution in visual dialog using neural module networks.” Proceedings of the European Conference on Computer Vision (ECCV). 2018, Sep. 8-14, 2018. [cited by applicant]
Kumar, Ankit, et al. “Ask me anything: Dynamic memory networks for natural language processing.” International conference on machine learning. 2016, Jan. 6, 2016. [cited by applicant]
Moon, Seungwhan, Suyoun Kim, and Haohan Wang. “Multimodal transfer deep leaming with applications in audio-visual recognition.” arXiv preprint arXiv:1412.3121 (2014), Dec. 9, 2014. [cited by applicant]
Moon, Seungwhan, Leonardo Neves, and Vitor Carvalho. “Multimodal named entity recognition for short social media posts.” arXiv preprint arXiv:1802.07862 (2018), Feb. 22, 2018. [cited by applicant]
Moon, Seungwhan, Leonardo Neves, and Vitor Carvalho. “Zeroshot Multimodal Named Entity Disambiguation for Noisy Social Media Posts.” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistic… [cited by applicant]
Shah, Pararth, et al. “Bootstrapping a neural conversational agent with dialogue self-play, crowdsourcing and on-line reinforcement learning.” Proceedings of the 2018 Conference of the North American Chapter of the Asso… [cited by applicant]
Dinan, Emily, et al. “Advances in Conversational AI” https://ai.facebook.com/blog/advances-in-conversational-ai/?_xts_%5b0%5d=68.ARDgZpslcbW2Y4dGWBF1BBfrsZkeNMXeTFXLveffyaOCRJ0iNA80NQfAJ9Y6urka2DI6EQcbA0JoTxUuSGUFT-BkfY… [cited by applicant]
Ott, Myle, et al. “New advances in natural language processing to better connect people” https://ai.facebook.com/blog/new-advances-in-natural-language-processing-to-better-connect-people/?_xts_%5b0%5d=68.ARBpsX-0s8sV0sN… [cited by applicant]
Das, Abhishek, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. “Visual dialog.” In [cited by applicant]
Seo, Paul Hongsuck, Andreas Lehrmann, Bohyung Han, and Leonid Sigal. “Visual reference resolution using attention memory for visual dialog.” In [cited by applicant]
Kottur, Satwik, José MF Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. “Visual coreference resolution in visual dialog using neural module networks.” In [cited by applicant]
Lu, Jiasen, Anitha Kannan, Jianwei Yang, Devi Parikh, and Dhruv Batra. “Best of both worlds: Transferring knowledge from discriminative learning to a generative visual dialog model.” In [cited by applicant]
Das, Abhishek, Satwik Kottur, José MF Moura, Stefan Lee, and Dhruv Batra. “Learning cooperative visual dialog agents with deep reinforcement learning.” In [cited by applicant]
Kottur, Satwik, Josć MF Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. “Clevr-dialog: A diagnostic dataset for multi-round reasoning in visual dialog.” arXiv preprint arXiv:1903.03166. pp. 1-13, Sep. 18, 2019. [cited by applicant]
Johnson, Justin, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. “Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.” In [cited by applicant]
Niu, Yulei, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen. “Recursive visual attention in visual dialog.” In [cited by applicant]
Massiceti, Daniela, N. Siddharth, Puneet K. Dokania, and Philip HS Torr. “Flipdial: A generative model for two-way visual dialogue.” In [cited by applicant]
Wu, Qi, Peng Wang, Chunhua Shen, Ian Reid, and Anton Van Den Hengel. “Are you talking to me? reasoned visual dialog generation through adversarial learning.” In [cited by applicant]
Schwartz, Idan, Seunghak Yu, Tamir Hazan, and Alexander G. Schwing. “Factor graph attention.” In [cited by applicant]
Yang, Jianwei, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. “Graph r-cnn for scene graph generation.” In [cited by applicant]
Hudson, Drew A., and Christopher D. Manning. “Gqa: A new dataset for real-world visual reasoning and compositional question answering.” In [cited by applicant]
De Vries, Harm, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville. “Guesswhat?! visual object discovery through multi-modal dialogue.” In [cited by applicant]
Strub, Florian, Harm De Vries, Jeremie Mary, Bilal Piot, Aaron Courville, and Olivier Pietquin. “End-to-end optimization of goal-driven and visually grounded dialogue systems.” arXiv preprint arXiv:1703.05423. pp 1-7, M… [cited by applicant]
Yang, Jianwei, Jiasen Lu, Stefan Lee, Dhruv Batra, and Devi Parikh. “Visual curiosity: Learning to ask questions to learn visual recognition.” arXiv preprint arXiv:1810.00912. pp. 1-18, Oct. 1, 2018. [cited by applicant]
Cited By (1)
US 12,400,289