IP Library Granted Patent US 11,984,198
Granted Patent B2
US 11,984,198 · App. 18/152,118 · Granted May 14, 2024

Hash-based efficient comparison of sequencing results

Inventors: Geert Trooskens (Meise, BE); Wim Maria R. Van Criekinge (Waarloos, BE)
Assignee: SHARECARE AI, INC.
G16B30/00C12Q1/6827C12Q1/6869G06F16/2255G06F17/18G16B5/00G16B10/00G16B20/20G16B20/40G16B40/00G16B40/10G16B45/00G16B50/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,984,198
App. No.
18/152,118
Granted
May 14, 2024
Kind
B2
Abstract

First and second sequenced outputs are accessed. The sequenced outputs contain variants occurring at different carriers and at different carrier positions. Hashes are generated over a selected pattern length of positions for those carrier positions that are shared between the sequenced outputs to produce window hashes for base patterns in first and second sequences. Each sequence is based on the shared carrier positions and the respective sequenced output. The window hashes are non-unique. Window hashes that occur less than a ceiling number times are selected. The selected window hashes are compared between the sequences on a starting position basis such that selected window hashes for base patterns having same start positions in the sequenced outputs are compared. Common window hashes are identified between the sequences based on the comparing. A similarity measure is determined between the sequences based on the common window hashes.

Claims (72)

1. A computer-implemented method comprising:

accessing a first sequenced output and a second sequenced output, wherein the first sequenced output and the second sequenced output contain variants occurring at different carriers and at different carrier positions;

generating hashes over a selected pattern length of positions for those carrier positions that are shared between the first sequenced output and the second sequenced output to produce window hashes for base patterns in a first sequence and a second sequence, and wherein the first sequence is based on the shared carrier positions and the first sequenced output, the second sequence is based on the shared carrier positions and the second sequenced output, and the window hashes are non-unique;

selecting those of the window hashes that occur less than a ceiling number of times;

comparing the selected window hashes between the first sequence and the second sequence on a starting position basis such that selected window hashes for base patterns having same start positions in the first sequenced output and the second sequenced output are compared;

identifying common window hashes between the first sequence and the second sequence based on the comparing; and

determining a similarity measure between the first sequence and the second sequence based on the common window hashes.

2. The computer-implemented method of claim 1 , further comprising:

based on starting position-wise similarity measures, determining a percentage of shared bases between the first sequenced output and the second sequenced output.

3. The computer-implemented method of claim 2 , wherein the percentage of shared bases are determined on a carrier-by-carrier basis.

4. The computer-implemented method of claim 3 , further comprising:

determining inherited traits based on the percentage of shared bases.

5. The computer-implemented method of claim 3 , further comprising:

identifying common ancestors and close and distant relatives based on the percentage of shared bases.

6. The computer-implemented method of claim 2 , further comprising:

based on the starting position-wise similarity measures, determining a percentage of shared bases between a given individual's sequence and an ethnicity-specific sequence; and

identifying ethnic ancestry of the given individual based on the percentage of shared bases.

7. The computer-implemented method of claim 2 , further comprising:

based on the starting position-wise similarity measures, generating a distance tree between the first sequenced output and the second sequenced output.

8. The computer-implemented method of claim 1 , wherein the variants are those variants that have highest observation frequency.

9. The computer-implemented method of claim 1 , wherein the variants are identified by sixteen phased pairings.

10. The computer-implemented method of claim 1 , wherein the variants are identified by ten unphased pairings.

11. The computer-implemented method of claim 1 , wherein the selected pattern length of positions ranges from fifteen to forty bases.

12. The computer-implemented method of claim 1 , wherein the window hashes are produced independently.

13. The computer-implemented method of claim 1 , further comprising providing a zoom-in option and a zoom-out option of a genomic browser, and wherein the zoom-in option and the zoom-out option change size based on size of the first sequence and the second sequence.

14. The computer-implemented method of claim 1 , wherein the similarity measure is determined by a distance formula defined as

1

-

n

umber

of

common

window

hashes

n

umber

of

unique

window

hashes

.

15. The computer-implemented method of claim 1 , further comprising:

comparing the selected window hashes between the first sequence and the second sequence on a starting position basis such that a first selected window hash produced for a base pattern having a given start position in the first sequence is compared only to a second selected window hash produced for a base pattern having the given start position in the second sequence.

16. A non-transitory computer readable storage medium impressed with computer program instructions, the instructions, when executed on a processor, implement a method comprising:

accessing a first sequenced output and a second sequenced output, wherein the first sequenced output and the second sequenced output contain variants occurring at different carriers and at different carrier positions;

generating hashes over a selected pattern length of positions for those carrier positions that are shared between the first sequenced output and the second sequenced output to produce window hashes for base patterns in a first sequence and a second sequence, and wherein the first sequence is based on the shared carrier positions and the first sequenced output, the second sequence is based on the shared carrier positions and the second sequenced output, and the window hashes are non-unique;

selecting those of the window hashes that occur less than a ceiling number of times;

comparing the selected window hashes between the first sequence and the second sequence on a starting position basis such that selected window hashes for base patterns having same start positions in the first sequenced output and the second sequenced output are compared;

identifying common window hashes between the first sequence and the second sequence based on the comparing; and

determining a similarity measure between the first sequence and the second sequence based on the common window hashes.

17. The non-transitory computer readable storage medium of claim 16 , wherein the method further comprises:

based on starting position-wise similarity measures, determining a percentage of shared bases between the first sequenced output and the second sequenced output.

18. The non-transitory computer readable storage medium of claim 17 , wherein the percentage of shared bases are determined on a carrier-by-carrier basis.

19. A system comprising one or more processors coupled to memory, the memory loaded with computer instructions, the instructions, when executed on the processors, implement actions comprising:

accessing a first sequenced output and a second sequenced output, wherein the first sequenced output and the second sequenced output contain variants occurring at different carriers and at different carrier positions;

generating hashes over a selected pattern length of positions for those carrier positions that are shared between the first sequenced output and the second sequenced output to produce window hashes for base patterns in a first sequence and a second sequence, and wherein the first sequence is based on the shared carrier positions and the first sequenced output, the second sequence is based on the shared carrier positions and the second sequenced output, and the window hashes are non-unique;

selecting those of the window hashes that occur less than a ceiling number of times;

comparing the selected window hashes between the first sequence and the second sequence on a starting position basis such that selected window hashes for base patterns having same start positions in the first sequenced output and the second sequenced output are compared;

identifying common window hashes between the first sequence and the second sequence based on the comparing; and

determining a similarity measure between the first sequence and the second sequence based on the common window hashes.

20. The system of claim 18 , wherein the actions further comprise:

based on starting position-wise similarity measures, determining a percentage of shared bases between the first sequenced output and the second sequenced output.

Assignments (3)
SECURITY INTEREST Recorded Oct 22, 2024
From: HEALTHWAYS SC, LLC; SHARECARE AI, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION
Reel/Frame 068977/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2023
From: TROOSKENS, GEERT; VAN CRIEKINGE, WIM MARIA R.
To: DOC.AI, INC.
Reel/Frame 062414/0919 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2023
From: DOC.AI, INC.
To: SHARECARE AI, INC.
Reel/Frame 062414/0970 →
Continuity (5)
Continuation 16575278 · Sep 18, 2019
Provisional Application 62734872 · Sep 21, 2018
Provisional Application 62734895 · Sep 21, 2018
Provisional Application 62734840 · Sep 21, 2018
Related Publication 20230162817A1 · May 25, 2023