IP Library Granted Patent US 11,551,784
Granted Patent B2
US 11,551,784 · App. 16/575,278 · Granted Jan 10, 2023

Ordinal position-specific and hash-based efficient comparison of sequencing results

Inventors: Geert Trooskens (Ghent, BE); Wim Maria R. Van Criekinge (Waarloos, BE)
Assignee: SHARECARE AI, INC.
G16B30/00C12Q1/6827C12Q1/6869G06F16/2255G06F17/18G16B5/00G16B10/00G16B20/20G16B20/40G16B40/00G16B40/10G16B45/00G16B50/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,551,784
App. No.
16/575,278
Granted
Jan 10, 2023
Kind
B2
Abstract

The technology disclosed generates a reference array of variant data for locations that are shared between read results which are to be compared, and generates hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the read results. It then selects for comparison window hashes that occur less than a ceiling number of times and compares the selected window hashes to identify common window hashes between the read results. It then determines a similarity measure for the read results based on the common window hashes.

Claims (79)

1. A computer-implemented method of efficiently comparing sequenced outputs, the method including:

accessing a first sequenced output and a second sequenced output, wherein the first and second sequenced outputs contain variants occurring at different carriers and at different carrier positions;

generating a reference array for those carrier positions that are shared between the first and second sequenced outputs;

based on the reference array, generating a first sequence from the first sequenced output and a second sequence from the second sequenced output;

generating hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the first and second sequences;

selecting for comparison window hashes that occur less than a ceiling number of times;

comparing the selected window hashes between the first and second sequences on a starting position basis such that selected window hashes for base patterns having same start positions in the read results are compared;

identifying common window hashes between the first and second sequences based on the comparing; and

determining a similarity measure between the first and second sequences based on the common window hashes.

2. The computer-implemented method of claim 1 , further including:

based on starting position-wise similarity measures, determining a percentage of shared bases between the first and second sequenced outputs.

3. The computer-implemented method of claim 2 , wherein the percentage of shared bases are determined on a carrier-by-carrier basis.

4. The computer-implemented method of claim 3 , further including:

determining inherited traits based on the percentage of shared bases.

5. The computer-implemented method of claim 3 , further including:

identifying common ancestors and close and distant relatives based on the percentage of shared bases.

6. The computer-implemented method of claim 1 , further including:

based on the starting position-wise similarity measures, determining a percentage of shared bases between a given individual's sequence and an ethnicity-specific sequence; and

identifying ethnic ancestry of the given individual based on the percentage of shared bases.

7. The computer-implemented method of claim 1 , further including:

based on the starting position-wise similarity measures, generating a distance tree between the first and second sequenced outputs.

8. The computer-implemented method of claim 1 , wherein the variants are those variants that have highest observation frequency.

9. The computer-implemented method of claim 1 , wherein the variants are identified by sixteen phased pairings.

10. The computer-implemented method of claim 1 , wherein the variants are identified by ten unphased pairings.

11. The computer-implemented method of claim 1 , wherein the selected pattern length of positions ranges from fifteen to forty bases.

12. The computer-implemented method of claim 1 , wherein length of the reference array ranges from one hundred thousand to one million base positions.

13. The computer-implemented method of claim 12 , wherein the reference array is ordered by carriers and by carrier positions.

14. The computer-implemented method of claim 12 , wherein the pattern length of positions is selected based on the length of the reference array.

15. The computer-implemented method of claim 1 , wherein the ceiling number of times ranges from one to ten.

16. The computer-implemented method of claim 1 , wherein the similarity measure is determined by a distance formula defined as

1

-

number

of

common

window

hashes

number

of

unique

window

hashes

.

17. The computer-implemented method of claim 1 , further including:

comparing the selected window hashes between the first and second sequences on a starting position basis such that a first selected window hash produced for a base pattern having a given start position in the first sequence is compared only to a second selected window hash produced for a base pattern having the given start position in the second sequence.

18. A non-transitory computer readable storage medium impressed with computer program instructions to efficiently compare sequenced outputs, the instructions, when executed on a processor, implement a method comprising:

accessing a first sequenced output and a second sequenced output, wherein the first and second sequenced outputs contain variants occurring at different carriers and at different carrier positions;

generating a reference array for those carrier positions that are shared between the first and second sequenced outputs;

based on the reference array, generating a first sequence from the first sequenced output and a second sequence from the second sequenced output;

generating hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the first and second sequences;

selecting for comparison window hashes that occur less than a ceiling number of times;

comparing the selected window hashes between the first and second sequences on a starting position basis such that selected window hashes for base patterns having same start positions in the read results are compared;

identifying common window hashes between the first and second sequences based on the comparing; and

determining a similarity measure between the first and second sequences based on the common window hashes.

19. A system including one or more processors coupled to memory, the memory loaded with computer instructions to efficiently compare sequenced outputs, the instructions, when executed on the processors, implement actions comprising:

accessing a first sequenced output and a second sequenced output, wherein the first and second sequenced outputs contain variants occurring at different carriers and at different carrier positions;

generating a reference array for those carrier positions that are shared between the first and second sequenced outputs;

based on the reference array, generating a first sequence from the first sequenced output and a second sequence from the second sequenced output;

generating hashes over a selected pattern length of positions in the reference array to independently produce non-unique window hashes for base patterns in the first and second sequences;

selecting for comparison window hashes that occur less than a ceiling number of times;

comparing the selected window hashes between the first and second sequences on a starting position basis such that selected window hashes for base patterns having same start positions in the read results are compared;

identifying common window hashes between the first and second sequences based on the comparing; and

determining a similarity measure between the first and second sequences based on the common window hashes.

Assignments (3)
SECURITY INTEREST Recorded Oct 22, 2024
From: HEALTHWAYS SC, LLC; SHARECARE AI, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 068977/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: DOC.AI, INC.
To: SHARECARE AI, INC.
Reel/Frame 061248/0436 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2021
From: TROOSKENS, GEERT; VAN CRIEKINGE, WIM MARIA R.
To: DOC.AI, INC.
Reel/Frame 058157/0906 →
Continuity (4)
Provisional Application 62734872 · Sep 21, 2018
Provisional Application 62734840 · Sep 21, 2018
Provisional Application 62734895 · Sep 21, 2018
Related Publication 20200095628A1 · Mar 26, 2020