IP Library Granted Patent US 7,657,383
Granted Patent B2
US 7,657,383 · App. 10/856,600 · Granted Feb 2, 2010

Method, system, and apparatus for compactly storing a subject genome

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,657,383
App. No.
10/856,600
Granted
Feb 2, 2010
Kind
B2
Abstract

A method of representing a subject genome can include comparing the subject genome to a base reference genome and identifying a difference between the subject genome and the base reference genome. The method can further include assigning one or more items of descriptive information to the difference, and compiling the items into a data set, where the data set represents the subject genome.

Claims (61)

1. A method of compactly representing and storing a complete subject genome comprising:

programming a computer to carry out the following steps:

constructing a base reference genome for a species, wherein said base reference genome is a benchmark or model genome for the species;

providing the subject genome, said subject genome being a genome of a member of the species;

retrieving the base reference genome;

comparing the subject genome to the base reference genome;

identifying all differences between the subject genome and the base reference genome;

assigning one or more items of descriptive information to each of the differences, wherein each of said items of descriptive information describes a particular deviation of the subject genome from the base reference genome resulting in each of said differences;

compiling the items of descriptive information describing all the differences between the subject genome and the base reference genome into a data set, wherein the data set represents the complete subject genome, and wherein the data set does not include data other than the items of descriptive information describing each difference between the subject genome and the base reference genome; and

storing the data set representing the complete subject genome on a machine readable storage; and

using the stored data set for genetic research or medical treatment.

2. The method of claim 1 , wherein the difference is a nucleotide in the subject genome that is not in the base reference genome or is a nucleotide that is not in the subject genome that is in the base reference genome.

3. The method of claim 1 , wherein the difference is a nucleotide sequence contained in both the subject genome and the base reference genome but which is situated in different locations.

4. The method of claim 1 , wherein the difference is a single nucleotide polymorphism.

5. The method of claim 1 , wherein the difference is a discrepancy in the repeat count of a micro-satellite between the subject genome and the base reference genome.

6. The method of claim 1 , said step of assigning an item of descriptive information further comprising:

identifying a position reference point in the base reference genome;

locating a corresponding position reference point in the subject genome; and

assigning a description to the difference.

7. The method of claim 6 , wherein the position reference point is a gene marker.

8. The method of claims 6 , wherein the position reference point is a micro-satellite marker.

9. The method of claim 6 , further comprising:

comparing portions of DNA around the position reference point;

locating a starting point of the difference;

determining an ending point of the difference; and

calculating one or more offset values based on the starting and ending points.

10. A system for compactly representing and storing a complete subject genome comprising:

a computer programmed to:

construct a base reference genome for a species, wherein said base reference genome is a benchmark or model genome for the species;

obtain the subject genome, said subject genome being a genome of a member of the species;

retrieve the base reference genome;

compare the subject genome to the base reference genome;

identify all differences between the subject genome and the base reference genome;

assign one or more items of descriptive information to each of the differences, wherein each of said items of descriptive information describes a particular deviation of the subject genome from the base reference genome resulting in each of said differences;

compile the items of descriptive information describing all the differences between the subject genome and the base reference genome into a data set,

wherein the data set represents the complete subject genome and wherein the data set does not include data other than the items of descriptive information describing each difference between the subject genome and the base reference genome; and

store the data set representing the complete subject genome on a machine readable storage, wherein the stored data set is used for genetic research or medical treatment.

11. A machine readable storage, having stored thereon a computer program having a plurality of code sections executable by a machine for causing the machine to perform the steps of:

constructing a base reference genome for a species, wherein said base reference genome is a benchmark or model genome for the species;

providing the subject genome, said subject genome being a genome of a member of the species;

retrieving the base reference genome;

comparing the subject genome to the base reference genome;

identifying all differences between the subject genome and the base reference genome;

assigning one or more items of descriptive information to each of the differences, wherein each of said items of descriptive information describes a particular deviation of the subject genome from the base reference genome resulting in each of said differences;

compiling the items of descriptive information describing all the differences between the subject genome and the base reference genome into a data set, wherein the data set represents the complete subject genome, and wherein the data set does not include data other than the items of descriptive information describing each difference between the subject genome and the base reference genome; and

storing the data set representing the complete subject genome on a machine readable storage, wherein the stored data set is used for genetic research or medical treatment.

12. The machine readable storage of claim 11 , wherein the difference is a nucleotide in the subject genome that is not in the base reference genome or a nucleotide that is not in the subject genome that is in the base reference genome.

13. The machine readable storage of claim 11 , wherein the difference is a nucleotide sequence contained in both the subject genome and the base reference genome but which is situated in different locations.

14. The machine readable storage of claim 11 , wherein the difference is a single nucleotide polymorphism.

15. The machine readable storage of claim 11 , wherein the difference is a discrepancy in the repeat count of a micro-satellite between the subject genome and the base reference genome.

16. The machine readable storage of claim 11 , said step of assigning one or more items of descriptive information further comprising:

identifying a position reference point in the base reference genome;

locating a corresponding position reference point in the subject genome; and

assigning a text description to the difference.

17. The machine readable storage of claim 16 , wherein the position reference point is a gene marker.

18. The machine readable storage of claim 16 , wherein the position reference point is a micro-satellite marker.

19. The machine readable storage of claim 16 , further comprising:

comparing portions of DNA around the position reference point;

locating a starting point of the difference;

determining an ending point of the difference; and

calculating one or more offset values based on the starting and ending points.

Assignments (8)
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS (REEL 062079, FRAME 0677) Recorded Mar 3, 2026
From: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
To: X CORP. (F/K/A TWITTER, INC.)
Reel/Frame 075015/0574 →
RELEASE OF SECURITY INTEREST Recorded Apr 30, 2025
From: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
To: X CORP. (F/K/A TWITTER, INC.)
Reel/Frame 071127/0240 →
RELEASE OF SECURITY INTEREST Recorded Mar 27, 2025
From: MORGAN STANLEY SENIOR FUNDING, INC.
To: X CORP. (F/K/A TWITTER, INC.)
Reel/Frame 070670/0857 →
SECURITY INTEREST Recorded Oct 28, 2022
From: TWITTER, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 062079/0677 →
SECURITY INTEREST Recorded Oct 28, 2022
From: TWITTER, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 061804/0001 →
SECURITY INTEREST Recorded Oct 28, 2022
From: TWITTER, INC.
To: MORGAN STANLEY SENIOR FUNDING, INC.
Reel/Frame 061804/0086 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2014
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: TWITTER, INC.
Reel/Frame 032075/0404 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2004
From: ALLARD, DAVID J.; SZABO, ROBERT M.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 014762/0844 →
Continuity (1)
Related Publication 20050267693A1 · Dec 1, 2005