IP Library › Granted Patent US 12,657,010
Granted Patent B2
US 12,657,010 · App. 18/202,756 · Granted Jun 16, 2026

Doc4code—an AI-driven documentation recommender system to aid programmers

Inventors: Arno Schneuwly (Effretikon, CH); Saeid Allahdadian (Vancouver, CA); Pritam Dash (Vancouver, CA); Matteo Casserini (Zurich, CH); Felix Schmidt (Baden-Dattwil, CH); Eric Sedlar (Portola Valley, CA)
Assignee: Oracle International Corporation
G06F8/36G06F16/955G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,010
App. No.
18/202,756
Granted
Jun 16, 2026
Kind
B2
Abstract

Herein for each source logic in a corpus, a computer stores an identifier of the source logic and operates a logic encoder that infers a distinct fixed-size encoded logic that represents the variable-size source logic. At build time, a multidimensional index is generated and populated based on the encoded logics that represent the source logics in the corpus. At runtime, a user may edit and select a new source logic such as in a text editor or an integrated development environment (IDE). The logic encoder infers a new encoded logic that represents the new source logic. The multidimensional index accepts the new encoded logic as a lookup key and automatically selects and returns a result subset of encoded logics that represent similar source logics in the corpus. For display, the multidimensional index may select and return only encoded logics that are the few nearest neighbors to the new encoded logic.

Claims (51)

1 . A method comprising:

for each source logic in a plurality of source logics that include a particular source logic:

storing an identifier of the source logic, and

first inferring a distinct fixed-size encoded logic that represents the source logic;

generating a multidimensional index based on the distinct fixed-size encoded logics that represent the plurality of source logics;

second inferring a new fixed-size encoded logic that represents a new source logic that contains a syntax error;

selecting, based on the multidimensional index, a result subset of the distinct fixed-size encoded logics that represent the plurality of source logics that are nearest neighbors to the new fixed-size encoded logic, wherein the result subset of the distinct fixed-size encoded logics contains the distinct fixed-size encoded logic that represents the particular source logic;

first displaying, on a display of a computer, the identifiers for the result subset of the distinct fixed-size encoded logics that represent the plurality of source logics;

receiving, from an input device of the computer, an interactive selection, on the display of the computer, of the identifier of the particular source logic; and

second displaying, on the display of the computer in response to said receiving, the particular source logic;

wherein the method is performed by one or more computers that includes said computer.

2 . The method of claim 1 wherein:

the identifier of the particular source logic comprises a uniform resource locator (URL) of a webpage;

said second displaying comprises displaying the webpage;

said storing the identifier of the particular source logic comprises extracting the particular source logic from the webpage.

3 . The method of claim 2 wherein the webpage contains a question and an answer that contains the particular source logic.

4 . The method of claim 3 wherein said first displaying comprises displaying at least one selected from a group consisting of the question and a title of the webpage.

5 . The method of claim 2 wherein the webpage contains application program interface (API) reference documentation.

6 . The method of claim 1 wherein the particular source logic comprises a declaration of a subroutine without a definition of the subroutine.

7 . The method of claim 1 wherein said selecting the result subset of the distinct fixed-size encoded logics is performed in an integrated development environment (IDE).

8 . The method of claim 7 wherein said second displaying is performed inside or outside of the IDE.

9 . The method of claim 1 further comprising at least one of:

interactively selecting the particular source logic that is less than an entire subroutine,

automatically selecting a lexical scope that is less than a source file, or

selecting the particular source logic that is multiple files.

10 . The method of claim 1 wherein the plurality of source logics contains at least a million source logics.

11 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

for each source logic in a plurality of source logics that include a particular source logic:

storing an identifier of the source logic, and

first inferring a distinct fixed-size encoded logic that represents the source logic;

generating a multidimensional index based on the distinct fixed-size encoded logics that represent the plurality of source logics;

second inferring a new fixed-size encoded logic that represents a new source logic that contains a syntax error;

selecting, based on the multidimensional index, a result subset of the distinct fixed-size encoded logics that represent the plurality of source logics that are nearest neighbors to the new fixed-size encoded logic, wherein the result subset of the distinct fixed-size encoded logics contains the distinct fixed-size encoded logic that represents the particular source logic;

first displaying, on a display of a computer, the identifiers for the result subset of the distinct fixed-size encoded logics that represent the plurality of source logics;

receiving, from an input device of the computer, an interactive selection, on the display of the computer, of the identifier of the particular source logic; and

second displaying, on the display of the computer in response to said receiving, the particular source logic.

12 . The one or more non-transitory computer-readable media of claim 11 wherein:

the identifier of the particular source logic comprises a uniform resource locator (URL) of a webpage;

said second displaying comprises displaying the webpage;

said storing the identifier of the particular source logic comprises extracting the particular source logic from the webpage.

13 . The one or more non-transitory computer-readable media of claim 12 wherein the webpage contains a question and an answer that contains the particular source logic.

14 . The one or more non-transitory computer-readable media of claim 13 wherein said first displaying comprises displaying at least one selected from a group consisting of the question and a title of the webpage.

15 . The one or more non-transitory computer-readable media of claim 12 wherein the webpage contains application program interface (API) reference documentation.

16 . The one or more non-transitory computer-readable media of claim 11 wherein the particular source logic comprises a declaration of a subroutine without a definition of the subroutine.

17 . The one or more non-transitory computer-readable media of claim 11 wherein said selecting the result subset of the distinct fixed-size encoded logics is performed in an integrated development environment (IDE).

18 . The one or more non-transitory computer-readable media of claim 17 wherein said second displaying is performed inside or outside of the IDE.

19 . The one or more non-transitory computer-readable media of claim 11 wherein the instructions further cause at least one of:

interactively selecting the particular source logic that is less than an entire subroutine,

automatically selecting a lexical scope that is less than a source file, or

selecting the particular source logic that is multiple files.

20 . The one or more non-transitory computer-readable media of claim 11 wherein the plurality of source logics contains at least a million source logics.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2023
From: SCHNEUWLY, ARNO; ALLAHDADIAN, SAEID; DASH, PRITAM; CASSERINI, MATTEO; SCHMIDT, FELIX; SEDLAR, ERIC
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 063778/0493 →
Continuity (2)
Provisional Application 63459421 · Apr 14, 2023
Related Publication 20240345811A1 · Oct 17, 2024
References Cited (52)
US 7293008B2 · Potter · 2007 [cited by applicant]
US 9495398B2 · Parkkinen · 2016 [cited by applicant]
US 11656851B2 · Clement · 2023 [cited by examiner]
US 11740879B2 · Huang · 2023 [cited by examiner]
US 12073195B2 · Duan · 2024 [cited by examiner]
US 12130809B2 · Miller · 2024 [cited by examiner]
US 20130074036A1 · Brandt · 2013 [cited by applicant]
US 20160253500A1 · Alme · 2016 [cited by applicant]
US 20180075349A1 · Zhan · 2018 [cited by applicant]
US 20190073404A1 · Klouche · 2019 [cited by applicant]
US 20200097261A1 · Smith · 2020 [cited by applicant]
US 20200249918A1 · Svyatkovskiy · 2020 [cited by applicant]
US 20200349052A1 · Wehr · 2020 [cited by applicant]
US 20210132913A1 · Singh · 2021 [cited by applicant]
US 20230236811A1 · Sundaresan · 2023 [cited by applicant]
US 20230273776A1 · Wang · 2023 [cited by applicant]
US 20230376743A1 · Nikolic · 2023 [cited by applicant]
US 20240028740A1 · Chan · 2024 [cited by applicant]
US 20240086164A1 · Kramer · 2024 [cited by examiner]
US 20240127112A1 · Ziegler · 2024 [cited by applicant]
US 20240134614A1 · Bakshi · 2024 [cited by examiner]
US 20240289606A1 · Wang · 2024 [cited by applicant]
US 20240345815A1 · Dash · 2024 [cited by examiner]
Rose, Stuart et al., “Automatic Keyword Extraction from Individual Documents”, Apr. 2010, pp. 1-20. [cited by applicant]
Campos, Ricardo et al., “YAKE! Keyword Extraction from Single Documents using Multiple Local Features”, Information Sciences, 509, 2020, 257-289. [cited by applicant]
Zheng, Qinkai, et al., “A Multilingual Code Generation Tool—CodeGeeX”, https://codegeex.ai/, 2023, 8pgs. [cited by applicant]
Johnson, Jeff, et al., “Billion-scale similarity search with GPUs”, https://arxiv.org/pdf/1702.08734.pdf, Feb. 28, 2017, 12pgs. [cited by applicant]
Feng, Zhangyin, et al., “CodeBERT: A Pre-Trained Model for Programming and Natural Languages”, Findings of the Assoctn for Computnl Linguistics: EMNLP 2020, pp. 1536-1547, https://doi.org/10.18653/v1/2020.findings-emnlp… [cited by applicant]
Chen, Mark, et al., “Evaluating Large Language Models Trained on Code”, https://arxiv.org/abs/2107.03374, Jul. 14, 2021, 35pgs. [cited by applicant]
“Tabnine is an AI assistant that speeds up delivery and keeps your code safe,” Tabnine Ltd., https://www.tabnine.com/, 2023, 9pgs. [cited by applicant]
“Stack Overflow—Where Developers Learn, Share, & Build Careers”, Stack Exchange Inc., https://stackoverflow.com/, 2023, 13pgs. [cited by applicant]
“Stack Exchange Data Dump: Stack Exchange, Inc.”, https://archive.org/details/stackexchange, 2023, 8pgs. [cited by applicant]
“Sourcery—Automatically Improve Code Quality”, https://sourcery.ai/, 2023, 14pgs. [cited by applicant]
“Github copilot”, GitHub Inc. https://github.com/features/copilot, 2023, 16pgs. [cited by applicant]
“GitHub Copilot litigation”, Joseph Saveri Law Firm, https://githubcopilotlitigation.com/, 2023, 4pgs. [cited by applicant]
“ChatGPT: Optimizing Language Models for Dialogue”, GOpenAI, LLC, https://chat.openai.com/, 2023, 1pg. [cited by applicant]
Zugner, Daniel, et al., “Language-Agnostic Representation Learning of Source Code from Structure and Context”, ICLR 2021, https://arxiv.org/abs/2103.11318, Mar. 21, 2021, 22pgs. [cited by applicant]
Zeng, Jie, et al., “Fast Code Clone Detection Based on Weighted Recursive Autoencoders”, IEEE Access, vol. 7; pp. 125062-125078, 2019. doi: 10.1109/ACCESS.2019.2938825, published Sep. 2, 2019, 17pgs. [cited by applicant]
White, Martin, et al., “Deep Learning Code Fragments for Code Clone Detection”, 31st IEEE/ACM ICASE 2016, dx.doi.org/10.1145/2970276.2970326, pp. 87-98, publ Aug. 25, 2016, 12pgs. [cited by applicant]
Wang, Xin, et al., “SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation”, https://arxiv.org/abs/2108.04556, Sep. 9, 2021, 9pgs. [cited by applicant]
Jiang, Lingxiao, et al., “Deckard: Scalable and Accurate Tree-based Detection of Code Clones”, 29th ICSE '07, doi: 10.1109/ICSE.2007.30, pp. 96-105, May 24, 2007, 10pgs. [cited by applicant]
Guo, Daya, et al., “UniXcoder: Unified Cross-Modal Pre-training for Code Representation”, https://arxiv.org/abs/2203.03850, Mar. 8, 2022, 14pgs. [cited by applicant]
Gao, Yi, et al., “TECCD: A Tree Embedding Approach for Code Clone Detection”, 2019 IEEE CSME, doi: 10.1109/ICSME.2019.00025, pp. 145-156, 2019, 12pgs. [cited by applicant]
Feng, Zhangyin, et al., “CodeBERT: A Pre-Trained Model for Programming and Natural Languages”, https://arxiv.org/abs/2002.08155, Feb. 19, 2020, 10pgs. [cited by applicant]
Fang, Chunrong, et al., “Functional code clone detection with syntax and semantics fusion learning”, Proc of the 29th ACM SIGSOFT ISSTA '20, pp. 516-527, doi: 10.1145/3395363.3397362, Jul. 18, 2020, 12pgs. [cited by applicant]
Devlin, Jacob, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, https://arxiv.org/abs/1810.04805, Oct. 11, 2018, 14pgs. [cited by applicant]
Bojanowski, Piotr, et al., “Enriching Word Vectors with Subword Information”, https://arxiv.org/abs/1607.04606v1, Jul. 15, 2016, 7pgs. [cited by applicant]
“Hugging face: Build, train and deploy state of the art models powered by the reference open source in machine learning”, https://huggingface.co, downloaded May 5, 2023, 11pgs. [cited by applicant]
“Fasttext: Library for efficient text classification and representation learning”, https://fasttext.cc/, downloaded May 5, 2023, 5pgs. [cited by applicant]
Prompting Large Language Model for Machine Translation-A Case Study, Biao et al., (Year: 2023). [cited by applicant]
Prompting a Large Language Model to Generate Diverse Motivational Messages, Samuel et al., (Year: 2023). [cited by applicant]
Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models—A Case Study on ChatGPT, Qingyu Lu (Year: 2023). [cited by applicant]