IP Library › Granted Patent US 12,277,126
Granted Patent B2
US 12,277,126 · App. 18/217,490 · Granted Apr 15, 2025

Methods and systems for search and ranking of code snippets using machine learning models

Inventors: Ashok Balasubramanian (Chennai, IN); Arul Reagan S (Chengalpattu District, IN); Balaji Munusamy (Chennai, IN); Karthikeyan Krishnaswamy Raja (Chennai, IN)
Assignee: OPEN WEAVER INC.
G06F16/24578G06F16/248
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,126
App. No.
18/217,490
Granted
Apr 15, 2025
Kind
B2
Abstract

Systems and methods for automatically generating search list rankings of code snippets are provided. An exemplary method includes searching source code repositories to identify code snippets in response to a search query, assigning weight values to ranking parameters, and processing the code snippets using machine learning models to generate rating scores for each of the code snippets, where each rating score applies to a corresponding ranking parameter. The method includes generating a combined score for each of the code snippets by combining the rating scores for the code snippet according to the weight values assigned to the corresponding ranking parameters and generating and presenting a user interface including an ordered list of the code snippets based on the combined scores for the code snippets.

Claims (63)

1. A system for automatically generating search list rankings of code snippets using one or more machine learning models, the system comprising:

one or more processors and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

searching one or more source code repositories to identify a plurality of code snippets in response to a search query;

assigning a plurality of weight values to a plurality of ranking parameters, each of the plurality of ranking parameters corresponding to a different type of rating score for plurality of code snippets;

processing the plurality of code snippets using the one or more machine learning models to generate a plurality of rating scores for each of the plurality of code snippets, wherein each rating score of the plurality of rating scores applies to a corresponding ranking parameter of the plurality of ranking parameters;

generating a combined score for each of the plurality of code snippets, wherein the combined score for a code snippet is generated by combining the plurality of rating scores for the code snippet according to the plurality of weight values assigned to the corresponding ranking parameters; and

generating and presenting a user interface comprising an ordered list of the plurality of code snippets based on the combined score for each of the plurality of code snippets,

wherein the ranking parameters comprise a popularity score, a relevancy score, a recency score, an engagement score, and a document type score, wherein:

the popularity score is a standardized measurement relative to a popularity of a code snippet;

the relevancy score is a standardized measurement relative to a relevance of a code snippet to the search query;

the recency score is a standardized measurement relative to how recently a code snippet was used;

the engagement score is a standardized measurement relative to an amount of engagement a code snippet receives; and

the document type score is a standardized measurement relative to a type of document from which the code snippet was derived.

2. The system of claim 1 , wherein the popularity score, the relevancy score, the recency score, the engagement score, and the document type score are normalized across provider data and third party data by a machine learning model.

3. The system of claim 1 , wherein the popularity score is determined based on a number of views and/or user downloads of a code snippet across different providers of the code snippet, the different providers of the code snippet comprising a plurality of code repositories.

4. The system of claim 1 , wherein the relevancy score is determined by computing a similarity between the code snippet and the search query based on a L2 Euclidian distance.

5. The system of claim 1 , wherein the engagement score is determined based on a number of clicks, comments, likes, and shares of a code snippet across different providers of the code snippet, the different providers of the code snippet comprising a plurality of code repositories.

6. The system of claim 1 , wherein the operations further comprise:

scanning, by a data-crawler, at least one of public repositories, cloud providers, Q&A sites, or review sites to retrieve information on popularity, relevancy, recency, engagement, and document type information regarding a code snippet; and

storing, by the data-crawler, the information to the one or more machine learning models.

7. The system of claim 1 , the operations further comprising:

receiving a large volume of technical documents from a variety of sources, wherein the large volume of technical documents comprises at least 30 million documents;

pre-processing the large volume of technical documents to extract relevant data;

semantically vectorizing the technical documents using a pre-trained machine learning model; and

storing the vectorized technical documents in a vector database.

8. A method for automatically generating search list rankings of code snippets, the method comprising:

searching one or more source code repositories to identify a plurality of code snippets in response to a search query;

assigning a plurality of weight values to a plurality of ranking parameters, each of the plurality of ranking parameters corresponding to a different type of rating score for plurality of code snippets;

processing the plurality of code snippets using one or more machine learning models to generate a plurality of rating scores for each of the plurality of code snippets, wherein each rating score of the plurality of rating scores applies to a corresponding ranking parameter of the plurality of ranking parameters;

generating a combined score for each of the plurality of code snippets, wherein the combined score for a code snippet is generated by combining the plurality of rating scores for the code snippet according to the plurality of weight values assigned to the corresponding ranking parameters; and

generating and presenting a user interface comprising an ordered list of the plurality of code snippets based on the combined score for each of the plurality of code snippets,

wherein the ranking parameters comprise a popularity score, a relevancy score, a recency score, an engagement score, and a document type score, wherein:

the popularity score is a standardized measurement relative to a popularity of a code snippet;

the relevancy score is a standardized measurement relative to a relevance of a code snippet to the search query;

the recency score is a standardized measurement relative to how recently a code snippet was used;

the engagement score is a standardized measurement relative to an amount of engagement a code snippet receives; and

the document type score is a standardized measurement relative to a type of document from which the code snippet was derived.

9. The method of claim 8 , wherein the popularity score is determined based on a number of views and/or user downloads of a code snippet across different providers of the code snippet, the different providers of the code snippet comprising a plurality of code repositories.

10. The method of claim 8 , wherein the relevancy score is determined by computing a similarity between the code snippet and the search query based on a L2 Euclidian distance.

11. The method of claim 8 , wherein the engagement score is determined based on a number of clicks, comments, likes, and shares of a code snippet across different providers of the code snippet, the different providers of the code snippet comprising a plurality of code repositories.

12. The method of claim 8 , wherein the method further comprises:

scanning, by a data-crawler, at least one of public repositories, cloud providers, Q&A sites, or review sites to retrieve information on popularity, relevancy, recency, engagement, and document type information regarding a code snippet; and

storing, by the data-crawler, the information to the one or more machine learning models.

13. The method of claim 8 , further comprising:

receiving a large volume of technical documents from a variety of sources, wherein the large volume of technical documents comprises at least 30 million documents;

pre-processing the large volume of technical documents to extract relevant data;

semantically vectorizing the technical documents using a pre-trained machine learning model; and

storing the vectorized technical documents in a vector database.

14. One or more non-transitory computer-readable media storing instructions thereon, wherein the instructions, when executed by one or more processors, cause the one or more processors to:

search one or more source code repositories to identify a plurality of code snippets in response to a search query;

assign a plurality of weight values to a plurality of ranking parameters, each of the plurality of ranking parameters corresponding to a different type of rating score for plurality of code snippets;

process the plurality of code snippets using one or more machine learning models to generate a plurality of rating scores for each of the plurality of code snippets, wherein each rating score of the plurality of rating scores applies to a corresponding ranking parameter of the plurality of ranking parameters;

generate a combined score for each of the plurality of code snippets, wherein the combined score for a code snippet is generated by combining the plurality of rating scores for the code snippet according to the plurality of weight values assigned to the corresponding ranking parameters; and

generate and present a user interface comprising an ordered list of the plurality of code snippets based on the combined score for each of the plurality of code snippets,

wherein the ranking parameters comprise a popularity score, a relevancy score, a recency score, an engagement score, and a document type score, wherein:

the popularity score is a standardized measurement relative to a popularity of a code snippet;

the relevancy score is a standardized measurement relative to a relevance of a code snippet to the search query;

the recency score is a standardized measurement relative to how recently a code snippet was used;

the engagement score is a standardized measurement relative to an amount of engagement a code snippet receives; and

the document type score is a standardized measurement relative to a type of document from which the code snippet was derived.

15. The one or more non-transitory computer-readable media of claim 14 , wherein the popularity score is determined based on a number of views and/or user downloads of a code snippet across different providers of the code snippet, the different providers of the code snippet comprising a plurality of code repositories.

16. The one or more non-transitory computer-readable media of claim 14 , wherein the relevancy score is determined by computing a similarity between the code snippet and the search query based on a L2 Euclidian distance.

17. The one or more non-transitory computer-readable media of claim 14 , wherein the engagement score is determined based on a number of clicks, comments, likes, and shares of a code snippet across different providers of the code snippet, the different providers of the code snippet comprising a plurality of code repositories.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2023
From: BALASUBRAMANIAN, ASHOK; S, ARUL REAGAN; MUNUSAMY, BALAJI; RAJA, KARTHIKEYAN KRISHNASWAMY
To: OPEN WEAVER INC.
Reel/Frame 064138/0717 →
Continuity (1)
Related Publication 20250005026A1 · Jan 2, 2025
References Cited (185)
US 5953526A · Day et al. · 1999 [cited by applicant]
US 7322024B2 · Carlson et al. · 2008 [cited by applicant]
US 7703070B2 · Bisceglia · 2010 [cited by applicant]
US 7774288B2 · Acharya et al. · 2010 [cited by applicant]
US 7958493B2 · Lindsey et al. · 2011 [cited by applicant]
US 8010539B2 · Blair-Goldensohn et al. · 2011 [cited by applicant]
US 8051332B2 · Zakonov et al. · 2011 [cited by applicant]
US 8112738B2 · Pohl et al. · 2012 [cited by applicant]
US 8112744B2 · Geisinger · 2012 [cited by applicant]
US 8219557B2 · Grefenstette et al. · 2012 [cited by applicant]
US 8296311B2 · Rapp et al. · 2012 [cited by applicant]
US 8412813B2 · Carlson et al. · 2013 [cited by applicant]
US 8417713B1 · Blair-Goldensohn et al. · 2013 [cited by applicant]
US 8452742B2 · Hashimoto et al. · 2013 [cited by applicant]
US 8463595B1 · Rehling et al. · 2013 [cited by applicant]
US 8498974B1 · Kim et al. · 2013 [cited by applicant]
US 8627270B2 · Fox et al. · 2014 [cited by applicant]
US 8677320B2 · Wilson et al. · 2014 [cited by applicant]
US 8688676B2 · Rush et al. · 2014 [cited by applicant]
US 8838606B1 · Cormack et al. · 2014 [cited by applicant]
US 8838633B2 · Dhillon et al. · 2014 [cited by applicant]
US 8935192B1 · Ventilla et al. · 2015 [cited by applicant]
US 8943039B1 · Grieselhuber et al. · 2015 [cited by applicant]
US 9015730B1 · Allen et al. · 2015 [cited by applicant]
US 9043753B2 · Fox et al. · 2015 [cited by applicant]
US 9047283B1 · Zhang et al. · 2015 [cited by applicant]
US 9135665B2 · England et al. · 2015 [cited by applicant]
US 9176729B2 · Mockus et al. · 2015 [cited by applicant]
US 9201931B2 · Lightner et al. · 2015 [cited by applicant]
US 9268805B2 · Crossley et al. · 2016 [cited by applicant]
US 9330174B1 · Zhang · 2016 [cited by applicant]
US 9361294B2 · Smith · 2016 [cited by applicant]
US 9390268B1 · Martini et al. · 2016 [cited by applicant]
US 9471559B2 · Castelli et al. · 2016 [cited by applicant]
US 9558098B1 · Alshayeb et al. · 2017 [cited by applicant]
US 9589250B2 · Palanisamy et al. · 2017 [cited by applicant]
US 9626164B1 · Fuchs · 2017 [cited by applicant]
US 9672554B2 · Dumon et al. · 2017 [cited by applicant]
US 9977656B1 · Mannopantar et al. · 2018 [cited by applicant]
US 10305758B1 · Bhide et al. · 2019 [cited by applicant]
US 10474509B1 · Dube et al. · 2019 [cited by applicant]
US 10484429B1 · Fawcett et al. · 2019 [cited by applicant]
US 10761839B1 · Migoya et al. · 2020 [cited by applicant]
US 10922740B2 · Gupta et al. · 2021 [cited by applicant]
US 10983760B2 · Guan · 2021 [cited by examiner]
US 11023210B2 · Li et al. · 2021 [cited by applicant]
US 11238027B2 · Frost et al. · 2022 [cited by applicant]
US 11256484B2 · Nikumb et al. · 2022 [cited by applicant]
US 11288167B2 · Vaughan · 2022 [cited by applicant]
US 11294984B2 · Kittur et al. · 2022 [cited by applicant]
US 11295375B1 · Chitrapura et al. · 2022 [cited by applicant]
US 11301631B1 · Atallah et al. · 2022 [cited by applicant]
US 11334351B1 · Pandurangarao et al. · 2022 [cited by applicant]
US 11461093B1 · Edminster et al. · 2022 [cited by applicant]
US 11474817B2 · Sousa et al. · 2022 [cited by applicant]
US 11704406B2 · Lee et al. · 2023 [cited by applicant]
US 11893117B2 · Segal et al. · 2024 [cited by applicant]
US 11966446B2 · Socher · 2024 [cited by examiner]
US 12034754B2 · O'Hearn et al. · 2024 [cited by applicant]
US 20010054054A1 · Olson · 2001 [cited by applicant]
US 20020059204A1 · Harris · 2002 [cited by applicant]
US 20020099694A1 · Diamond · 2002 [cited by examiner]
US 20020150966A1 · Muraca · 2002 [cited by applicant]
US 20020194578A1 · Irie et al. · 2002 [cited by applicant]
US 20040243568A1 · Wang et al. · 2004 [cited by applicant]
US 20060090077A1 · Little et al. · 2006 [cited by applicant]
US 20060104515A1 · King et al. · 2006 [cited by applicant]
US 20060200741A1 · Demesa et al. · 2006 [cited by applicant]
US 20060265232A1 · Katariya et al. · 2006 [cited by applicant]
US 20070050343A1 · Siddaramappa et al. · 2007 [cited by applicant]
US 20070168946A1 · Drissi · 2007 [cited by examiner]
US 20070185860A1 · Lissack · 2007 [cited by applicant]
US 20070234291A1 · Ronen et al. · 2007 [cited by applicant]
US 20070299825A1 · Rush et al. · 2007 [cited by applicant]
US 20090043612A1 · Szela et al. · 2009 [cited by applicant]
US 20090319342A1 · Shilman et al. · 2009 [cited by applicant]
US 20100106705A1 · Rush et al. · 2010 [cited by applicant]
US 20100121857A1 · Elmore et al. · 2010 [cited by applicant]
US 20100122233A1 · Rath et al. · 2010 [cited by applicant]
US 20100174670A1 · Malik et al. · 2010 [cited by applicant]
US 20100205198A1 · Mishne et al. · 2010 [cited by applicant]
US 20100205663A1 · Ward et al. · 2010 [cited by applicant]
US 20100262454A1 · Sommer et al. · 2010 [cited by applicant]
US 20110231817A1 · Hadar et al. · 2011 [cited by applicant]
US 20120143879A1 · Stoitsev · 2012 [cited by applicant]
US 20120259882A1 · Thakur et al. · 2012 [cited by applicant]
US 20120278064A1 · Leary et al. · 2012 [cited by applicant]
US 20130103662A1 · Epstein · 2013 [cited by applicant]
US 20130117254A1 · Manuel-Devadoss et al. · 2013 [cited by applicant]
US 20130254744A1 · Sahoo et al. · 2013 [cited by applicant]
US 20130326469A1 · Fox et al. · 2013 [cited by applicant]
US 20140040238A1 · Scott et al. · 2014 [cited by applicant]
US 20140075414A1 · Fox et al. · 2014 [cited by applicant]
US 20140122182A1 · Cherusseri et al. · 2014 [cited by applicant]
US 20140149894A1 · Watanabe et al. · 2014 [cited by applicant]
US 20140163959A1 · Hebert et al. · 2014 [cited by applicant]
US 20140188746A1 · Li · 2014 [cited by applicant]
US 20140297476A1 · Wang · 2014 [cited by examiner]
US 20140331200A1 · Wadhwani et al. · 2014 [cited by applicant]
US 20140337355A1 · Heinze · 2014 [cited by applicant]
US 20150127567A1 · Menon et al. · 2015 [cited by applicant]
US 20150220608A1 · Crestani Campos et al. · 2015 [cited by applicant]
US 20150331866A1 · Shen et al. · 2015 [cited by applicant]
US 20150378692A1 · Dang · 2015 [cited by examiner]
US 20160253688A1 · Nielsen et al. · 2016 [cited by applicant]
US 20160350105A1 · Kumar et al. · 2016 [cited by applicant]
US 20160378618A1 · Cmielowski et al. · 2016 [cited by applicant]
US 20170034023A1 · Nickolov et al. · 2017 [cited by applicant]
US 20170063776A1 · Nigul · 2017 [cited by applicant]
US 20170154543A1 · King et al. · 2017 [cited by applicant]
US 20170177318A1 · Mark · 2017 [cited by examiner]
US 20170220633A1 · Porath et al. · 2017 [cited by applicant]
US 20170242892A1 · Ali et al. · 2017 [cited by applicant]
US 20170286541A1 · Mosley et al. · 2017 [cited by applicant]
US 20170286548A1 · De et al. · 2017 [cited by applicant]
US 20170344556A1 · Wu et al. · 2017 [cited by applicant]
US 20180046609A1 · Agarwal et al. · 2018 [cited by applicant]
US 20180067836A1 · Apkon et al. · 2018 [cited by applicant]
US 20180107983A1 · Mir Ghaderi · 2018 [cited by examiner]
US 20180114000A1 · Taylor · 2018 [cited by applicant]
US 20180189055A1 · Dasgupta et al. · 2018 [cited by applicant]
US 20180191599A1 · Balasubramanian et al. · 2018 [cited by applicant]
US 20180329883A1 · Leidner et al. · 2018 [cited by applicant]
US 20180349388A1 · Skiles et al. · 2018 [cited by applicant]
US 20190026106A1 · Burton et al. · 2019 [cited by applicant]
US 20190229998A1 · Cattoni · 2019 [cited by applicant]
US 20190278933A1 · Bendory et al. · 2019 [cited by applicant]
US 20190286683A1 · Kittur et al. · 2019 [cited by applicant]
US 20190294703A1 · Bolin et al. · 2019 [cited by applicant]
US 20190303141A1 · Li · 2019 [cited by examiner]
US 20190311044A1 · Xu et al. · 2019 [cited by applicant]
US 20190324981A1 · Counts et al. · 2019 [cited by applicant]
US 20200097261A1 · Smith · 2020 [cited by examiner]
US 20200110839A1 · Wang et al. · 2020 [cited by applicant]
US 20200125482A1 · Smith et al. · 2020 [cited by applicant]
US 20200133830A1 · Sharma et al. · 2020 [cited by applicant]
US 20200293354A1 · Song et al. · 2020 [cited by applicant]
US 20200301672A1 · Li et al. · 2020 [cited by applicant]
US 20200301908A1 · Frost et al. · 2020 [cited by applicant]
US 20200348929A1 · Sousa et al. · 2020 [cited by applicant]
US 20200356363A1 · Dewitt et al. · 2020 [cited by applicant]
US 20210049091A1 · Hikawa et al. · 2021 [cited by applicant]
US 20210065045A1 · Kummamuru et al. · 2021 [cited by applicant]
US 20210073293A1 · Fenton et al. · 2021 [cited by applicant]
US 20210081189A1 · Nucci et al. · 2021 [cited by applicant]
US 20210081418A1 · Silveira et al. · 2021 [cited by applicant]
US 20210141863A1 · Wu et al. · 2021 [cited by applicant]
US 20210149658A1 · Cannon et al. · 2021 [cited by applicant]
US 20210149668A1 · Gupta et al. · 2021 [cited by applicant]
US 20210256367A1 · Mor et al. · 2021 [cited by applicant]
US 20210303989A1 · Bird · 2021 [cited by examiner]
US 20210349801A1 · Rafey · 2021 [cited by applicant]
US 20210357210A1 · Clement et al. · 2021 [cited by applicant]
US 20210382712A1 · Richman et al. · 2021 [cited by applicant]
US 20210397418A1 · Nikumb et al. · 2021 [cited by applicant]
US 20210397546A1 · Cser et al. · 2021 [cited by applicant]
US 20220012297A1 · Basu et al. · 2022 [cited by applicant]
US 20220083577A1 · Yoshida et al. · 2022 [cited by applicant]
US 20220107802A1 · Rao et al. · 2022 [cited by applicant]
US 20220197916A1 · Sarkar et al. · 2022 [cited by applicant]
US 20220215068A1 · Kittur et al. · 2022 [cited by applicant]
US 20220261241A1 · Balasubramanian et al. · 2022 [cited by applicant]
US 20220269580A1 · Balasubramanian et al. · 2022 [cited by applicant]
US 20220269687A1 · Balasubramanian et al. · 2022 [cited by applicant]
US 20220269743A1 · Balasubramanian et al. · 2022 [cited by applicant]
US 20220269744A1 · Balasubramanian et al. · 2022 [cited by applicant]
US 20230308700A1 · Perez · 2023 [cited by applicant]
CN 108052442A · 2018 [cited by applicant]
KR 1020200062917 · 2020 [cited by applicant]
WO WO2007013418A1 · 2007 [cited by applicant]
WO WO2020086773A1 · 2020 [cited by applicant]
M. Squire, “Should We Move to Stack Overflow?” Measuring the Utility of Social Media for Developer Support, 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering, Florence, Italy, 2015, pp. 219-228, d… [cited by applicant]
S. Bayati, D. Parson, T. Sujsnjak and M. Heidary, “Big data analytics on large-scale socio-technical software engineering archives,” 2015 3rd International Conference on Information and Communication Technology (ICoICT)… [cited by applicant]
Lampropoulos et al, “React—A Process for Improving Open-Source Software Reuse”, IEEE, pp. 251-254 (Year: 2018). [cited by applicant]
Leclair et al., “A Neural Model for Generating Natural Language Summaries of Program Subroutines,” Collin McMillan, Dept. of Computer Science and Engineering, University of Notre Dame Notre Dame, IN, USA, Feb. 5, 2019. [cited by applicant]
Stanciulescu et al, “Forked and Integrated Variants in an Open-Source Firmware Project”, IEEE, pp. 151-160 (Year: 2015). [cited by applicant]
Zaimi et al, “:An Empirical Study on the Reuse of Third-Party Libraries in Open-Source Software Development”, ACM, pp. 1-8 (Year: 2015). [cited by applicant]
Andreas DAutovic, “Automatic Assessment of Software Documentation Quality”, published by IEEE, ASE 2011, Lawrence, KS, USA, pp. 665-669, (Year: 2011). [cited by applicant]
Iderli Souza, An Analysis of Automated Code Inspection Tools for PHP Available on Github Marketplace, Sep. 2021, pp. 10-17 (Year: 2021). [cited by applicant]
Khatri et al, “Validation of Patient Headache Care Education System (PHCES) Using a Software Reuse Reference Model”, Journal of System Architecture, pp. 157-162 (Year: 2001). [cited by applicant]
Lotter et al, “Code Reuse in Stack Overflow and Popular Open Source Java Projects”, IEEE, pp. 141-150 (Year: 2018). [cited by applicant]
Rothenberger et al, “Strategies for Software Reuse: A Principal Component Analysis of Reuse Practices”, IEEE, pp. 825-837 (Year:2003). [cited by applicant]
Tung et al, “A Framework of Code Reuse in Open Source Software”, ACM, pp. 1-6 (Year: 2014). [cited by applicant]
S. Bayati, D. Parson, T. Susnjakand M. Heidary, “Big data analytics on large-scale socio-technical software engineering archives,” 2015 3rd International Conference on Information and Communication Technology (ICoICT), … [cited by applicant]
Chung-Yang et al. “Toward Since-Source of Software Project Documented Contents: A Preliminary Study”, [Online], [Retrieve from Internet on Sep. 28, 2024], https://www.proquest.com/openview/c15dc8b34c7da061fd3ea39f1875d8… [cited by applicant]