IP Library › Granted Patent US 11,900,083
Granted Patent B2
US 11,900,083 · App. 16/878,079 · Granted Feb 13, 2024

Systems and methods for indexing source code in a search engine

Inventors: Charles Olivier (Sydney, AU); Stefan Saasen (Sydney, AU); Robin Stocker (Sydney, AU)
Assignee: ATLASSIAN PTY LTD.
G06F8/36G06F8/71G06F9/54G06F16/1873
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,900,083
App. No.
16/878,079
Granted
Feb 13, 2024
Kind
B2
Abstract

Method, system and computer readable storage medium for transmitting content from an SCM version of a repository maintained by an SCM system to a corresponding search engine (SE) version of the repository maintained by a search engine system. The method includes generating a content request, the content request comprising information defining a start state of the SCM version of the repository and a filter field; identifying one or more files in the SCM version of the repository that have changed between the start state and an end state; filtering the identified files based on the filter field in the content request to form a filtered set of files and a removed set of files; extracting content and metadata for one or more files from the filtered set of files; and transmitting the extracted content to the search system for storage as part of the search system version of the repository.

Claims (68)

1. A computer implemented method for transmitting content from a source code management (SCM) version of a repository maintained by a source code management system to a separate search engine (SE) version of the repository maintained by a search engine system, the method comprising:

determining a state of the SCM version of the repository, comprising a first set of files maintained by the SCM system;

determining an indexed state of the separate SE version of the repository using an index state query to generate an index state descriptor for a second set of files that are maintained by the SE system, the second set of files corresponding to the first set of files and including a subset of information from the first set of files;

in response to determining that the indexed state of the SE version of the repository is not the same as the state of the SCM version of the repository:

generating a content request, the content request comprising information defining a start state of the SCM version of the repository and a filter field;

identifying one or more files of the first set of files that have changed between the start state and an end state;

filtering the identified one or more files of the first set of files based on the filter field in the content request to form a filtered set of files;

extracting content and metadata for one or more files of the filtered set of files to generate a set of indexed files; and

updating the SE version of the repository to include the set of indexed files.

2. The method of claim 1 , wherein the content request further includes the end state of the SCM version of the repository, wherein, the start state defines a state of the SE version of the repository and the end state defines a state of the SCM version of the repository.

3. The method of claim 1 , further comprising transforming the extracted content for the one or more files from the filtered set of files.

4. The method of claim 1 , wherein filtering the identified files in the SCM version of the repository comprises, for a given identified file:

determining a size of the identified file;

comparing the size of the identified file with a threshold file size; and

adding the identified to the removed set of files if the determined file size exceeds the threshold file size.

5. The method of claim 1 , wherein filtering the identified files in the SCM version of the repository comprises, for a given identified file:

identifying a file type of the identified file;

comparing the file type with an invalid file type;

adding the identified file to the removed set of files if the identified file type matches the invalid file type.

6. The method of claim 1 , wherein filtering the identified files in the SCM version of the repository comprises filtering a given identified file based one or more of a file status or a file permission.

7. The method of claim 3 , wherein transforming the extracted data comprises adding line numbers to the extracted data.

8. The method of claim 7 , wherein adding line numbers to the extracted data comprising:

scanning the extracted content to identify line endings;

calculating incrementing line numbers for the identified line endings; and

prefixing the incrementing line numbers for identified line endings.

9. The method of claim 1 , further comprising creating a file descriptor for each file from the filtered set of files, wherein the file descriptor comprising a metadata field and a content field.

10. The method of claim 9 , further comprising creating a batch file comprising a plurality of file descriptors and transmitting the batch file to the search engine system.

11. A system for transmitting content from a source code management (SCM) version of a repository maintained by a source code management system to a separate search engine (SE) version of the repository maintained by a search engine system, the system comprising:

a processor,

a communication interface, and

a non-transitory computer-readable storage medium storing sequences of instructions, which when executed by the processor, cause the processor to:

determine a state of the SCM version of the repository comprising a first set of files maintained by the SCM system;

determine an indexed state of the separate SE version of the repository using an index state query to generate an index state descriptor for a second set of files that are maintained by the SE system, the second set of files corresponding to the first set of files and including a subset of information from the first set of files;

in response to determining that the indexed state of the SE version of the repository is not the same as the state of the SCM version of the repository:

generate a content request, the content request comprising information defining a start state and a filter field;

identify one or more files of the first set of files that have changed between the start state and an end state;

filter the identified one or more files of the first set of files based on the filter field in the content request to form a filtered set of files;

extract content and metadata for one or more files of the filtered set of files to generate a set of indexed files; and

update the SE version of the repository to include the set of indexed files.

12. The system of claim 11 , wherein:

the content request further includes the end state; and

the start state defines a state of the SE version of the repository and the end state defines a state of the SCM version of the repository.

13. The system of claim 11 , wherein the processor is configured to execute instructions which cause the processor to: transform the extracted content for the one or more files from the filtered set of files.

14. The system of claim 11 , wherein to filter the identified files in the SCM version of the repository, the processor is configured to execute instructions which cause the processor to, for a given identified file:

determine a size of the identified file,

compare the size of the identified file with a threshold file size, and

add the identified file to the removed set of files if the determined file size exceeds the threshold file size.

15. The system of claim 11 , wherein to filter the identified files in the SCM version of the repository, the processor is configured to execute instructions which cause the processor to, for a given identified file:

identify a file type of the identified file;

compare the file type with an invalid file type; and

add the identified file to the removed set of files if the identified file type matches an invalid file type.

16. The system of claim 11 , wherein the processor is configured to execute instructions which cause the processor to filter the identified files in the SCM version of the repository based one or more of a file status or a file permission.

17. The system of claim 13 , wherein transforming the extracted data comprises adding line numbers to the extracted data.

18. The system of claim 15 , wherein to add line numbers to the extracted data, the processor is configured to execute instructions which cause the processor to:

scan the extracted content for line endings;

calculate incrementing line numbers of the identified line endings; and

prefix the incrementing line numbers for the identified line endings.

19. The system of claim 11 , wherein the processor is configured to execute instructions which cause the processor to create a file descriptor for each file from the filtered set of files, wherein the file descriptor comprises a metadata field and a content field.

20. The system of claim 17 , wherein the processor is configured to execute instructions which cause the processor to create a batch file comprising a plurality of file descriptors and transmit the batch file to the search engine system.

21. Non-transient computer readable storage comprising instructions which, when executed by a processor, cause the processor to:

determine a state of a source code management (SCM) version of the repository comprising a first set of files maintained by the SCM system;

determine an indexed state of a separate search engine (SE) version of the repository using an index state query to generate an index state descriptor for a second set of files that are maintained by the SE system, the second set of files corresponding to the first set of files and including a subset of information from the first set of files;

in response to determining that the indexed state of the SE version of the repository is not the same as the state of the SCM version of the repository:

generate a content request, the content request comprising information defining a start state and an end state of the SCM version of the repository and one or more filter fields, the start state defining a state of the search SE version of the repository and the end state defining a state of the SCM version of the repository;

identify one or more files of the first set of files that have changed between the start state and the end state;

filter the identified one or more files of the first set of files based on the one or more filter fields in the content request to form a filtered set of files;

extract content and metadata for one or more files of the filtered set of files to generate a set of indexed files; and

update the SE version of the repository to include the set of indexed files.

Continuity (2)
Continuation 15362683 · Nov 28, 2016
Related Publication 20200278843A1 · Sep 3, 2020