US 5,582,979 AGrant
Length Polymorphisms in (dC-dA).sub.n.(dG-dT).sub.n Sequences and Method of Using the Same
Issue Date:1996-12-10
•10 Claims
•8 Drawing Sheets
Abstract
Abundant interspersed repetitive DNA sequences of the form (dC-dA).sub.n.(dG-dT).sub.n have been shown to exhibit length polymorphisms. The polymorphisms can be used to identify individuals as in paternity and forensic testing, and can be used to map genes which are involved in genetic diseases or in other economically important traits. The polynucleotide provided consists of a DNA fragment, preferably .ltoreq.300 base pairs in length containing one or more blocks of tandem dinucleotide repeats (dC-dA).sub.n.(dG-dT).sub.n wherein n.gtoreq.6 and preferably .gtoreq.10.
Metadata
Assignee
- Marshfield Clinic
Inventor
- James L. Weber
Application Information
Application Number:US 2221772
Filing Date:1994-04-04
Priority Date:1991-04-21
Art Unit:185
Classifications
IPC:
C12Q 168C12P 1934C07H 2104
Field of Search:
53643523.1;24.31;24.336;91.2
Patent Drawings (8 sheets)
Description
FIELD OF THE INVENTION
The present invention relates to polynucleotides which comprise an abundant new class of DNA polymorphisms and to methods for analyzing and using these polymorphisms. The polymorphisms can be used to identify individuals such as in paternity and forensic testing cases, and can also be used to map genes which are involved in genetic diseases or in other economically important traits.
BACKGROUND OF THE INVENTION
The vast majority of DNA in higher organisms is identical in sequence among different individuals (or more accurately among the chromosomes of those individuals). A small fraction of DNA, however, is variable or polymorphic in sequence among individuals, with the formal definition of polymorphism being that the most frequent variant (or allele) has a population frequency which does not exceed 99% (Gusella, J. F. (1986), Ann. Rev. Biochem. 55:831-854). In the past, polymorphisms were usually detected as variations in gene products or phenotypes such as human blood types. Currently, almost all polymorphisms are detected directly as variations in genomic DNA.
Analysis of DNA polymorphisms has relied on variations in the lengths of DNA fragments produced by restriction enzyme digestion. Most of these restriction fragment length polymorphisms (RFLPs) involve sequence variations in one of the recognition sites for the specific restriction enzyme used. This type of RFLP contains only two alleles, and hence is relatively uninformative.
A second type of RFLP is more informative and involves variable numbers of tandemly repeated DNA sequences between the restriction enzyme sites. These polymorphisms called minisatellites or VNTRs (for variable numbers of tandem repeats) were developed first by Jeffreys (Jeffreys et al. (1985), Nature 314:67-73). Jeffreys has filed two European patent applications, 186,271 and 238,329, dealing with the minisatellites. The first Jeffreys' application ('271) identified the existence of DNA regions containing hypervariable tandem repeats of DNA. Although the tandem repeat sequences generally varied between minisatellite regions, Jeffreys noted that many minisatellites had repeats which contain core regions of highly similar sequences. Jeffreys isolated or cloned, from genomic DNA, polynucleotide probes comprised essentially of this core sequence (i.e., wherein the probe had at least 70% homology with one of his defined cores). These probes were found to hybridize with multiple minisatellite regions (or loci). The probes were found to be useful in forensic or paternity testing by the identification of unique or characteristic minisatellite profiles. The later Jeffreys' European patent application proposed the use of probes which were specific for individual minisatellites located at specific loci in the genome. One problem with the Jeffreys' approach is that some of the most highly variable and hence useful minisatellites are susceptible to significant frequencies of random mutation (Jeffreys et al., 1988, Nature 332:278-281).
Other tandemly repeated DNA families, different in sequence from the Jeffreys minisatellites, are known to exist. In particular, (dC-dA).sub.n.(dG-dT).sub.n sequences have been found in all eukaryotes that have been examined. In humans there are 50,000-100,000 blocks of (dC-dA).sub.n.(dG-dT).sub.n sequences, with n ranging from about 15-30 (Miesfeld et al. (1981), Nucleic Acids Res. 9:5931-5947; Hamada and Kakunaga (1982), Nature 298:396-398; Tautz and Renz (1984), Nucleic Acids Res. 12:4127-4138).
Prior to the work of this invention, a number of different human blocks of (dC-dA).sub.n.(dG-dT).sub.n repeats had been cloned and sequenced, mostly unintentionally along with other sequences of interest. Several of these characterized sequences were analyzed independently from two or more alleles. In arriving at this invention, sequences from these different alleles were compared. Variations in the number of repeats per block of repeats were found in several cases (Weber and May, 1989, Am. J. Hum. Genet. 44:388-396, incorporated herein by reference in its entirety).
Although three isolated research groups produced published notations of site specific differences in sequence length (Das et al., 1987, J. Biol Chem. 262:4787-4793; Slightom et al., 1980, Cell 21:627-638; Shen and Rutter, 1984, Science 244:168-171), none of the groups recognized nor appreciated the extent of this variability or its usefulness and none generalized the observation. The other groups also did not consider the use of (dC-dA).sub.n.(dG-dT).sub.n sequences as genetic markers and did not offer a method by which such polymorphisms might be analyzed.
SUMMARY OF THE INVENTION
It has been discovered that (dC-dA).sub.n.(dG-dT).sub.n sequences exhibit length polymorphisms and therefore serve as an abundant pool of potential genetic markers (Weber and May, 1988, Am. J. Hum. Genet. 43:A161 Abstract); Weber and May, 1989, Am J. Hum. Genet. 44:388-396) (both incorporated herein by reference). Accordingly, as a first feature of the present invention, polynucleotides are provided consisting of a DNA fragment, preferably .ltoreq.300 base pairs (bp) in length, containing one or more blocks of tandem dinucleotide repeats (dC-dA).sub.n.(dG-dT).sub.n where n is preferably .gtoreq.6, and more preferably .gtoreq.10.
A further aspect of the invention is the provision of a method for analyzing one or more specific (dC-dA).sub.n. (dG-dT).sub.n polymorphisms individually or in combination, which involves amplification of a small segment(s) of DNA containing the block of repeats and some non-repeated flanking DNA, starting with a DNA template using the polymerase chain reaction, and sizing the resulting amplified DNA, preferably by electrophoresis on polyacrylamide gels.
In a preferred embodiment, the amplified DNA is labeled during the amplification reactions by incorporation of radioactive nucleotides or nucleotides modified with a non-radioactive reporter group.
A further aspect of the invention is the provision of primers for the amplification of the polymorphic tandemly repeated fragments. The primers are cloned, genomic or preferably synthesized, and contain at least a portion of the non-repeated, non-polymorphic flanking region sequence.
A further aspect of the invention is the provision of a method for determining the sequence information necessary for primer production through the isolation of DNA fragments, preferably as clones, containing the (dC-dA).sub.n.(dG-dT).sub.n repeats, by hybridization of a synthetic, cloned, amplified or genomic probe, which contains a sequence that is substantially homologous to the tandemly repeated sequence (dC-dA).sub.n.(dG-dT).sub.n, to the DNA fragment. In a preferred embodiment the probe would be labeled, e.g., end labeling, internal labeling or nick translation.
A further aspect of this invention is to define the sequence in terms of numbers of repeats and repeat sequence. Using a set of precise classification rules, the sequences were divided into three categories: perfect repeat sequences without interruptions in the runs of CA or GT dinucleotides, imperfect repeat sequences with one or more interruptions in the run of repeats, and compound repeat sequences with adjacent tandem simple repeats of a different sequence. Informativeness of (dC-dA).sub.n.(dG-dT).sub.n markers in the perfect sequence category was found to increase with increasing average numbers of repeats.
The present invention is also directed to a method for detecting the presence in genomic DNA of a specific trait in a subject, such as a human or other animal. The method includes isolating the genomic DNA from the subject and analyzing the genomic DNA with a polymorphic amplified DNA marker containing one or more sequences in the form (dC-dA).sub.n.(dG-dT).sub.n, wherein n is .gtoreq.6. The analysis comprises amplification using the polymerase chain reaction of one or more short DNA fragments containing the tandem repeats followed by measurement of the sizes of the amplified fragments using gel electrophoresis. This method has specific use for disease traits such as Huntington's disease and cystic fibrosis.
The present invention is also directed to a method for determining the paternity of an individual comprising amplification of polymorphic DNA fragments from the mother of the individual, the suspected father of the individual and the individual, analyzing the sizes of the fragments by gel electrophoresis, and comparing the electrophoretic patterns to determine correspondence between the individual's pattern and the mother's pattern and thereby determining whether the suspected father is the actual father of the individual.
The present invention is also directed to a kit for the rapid analysis of one or more specific DNA polymorphisms of the type described in this application which through proximity to a DNA abnormality causing a genetic disease permit determination of the presence or absence of the abnormality. The kit includes oligodeoxynucleotide primers for the amplification of fragments containing one or more sequences of the form (dC-dA).sub.n.(dG-dT).sub.n, where n.gtoreq.6.
In addition to the inherent useful properties of the (dC-dA).sub.n.(dG-dT).sub.n markers, the use of the polymerase chain reaction (PCR) to analyze the markers offers substantial advantages over the conventional blotting and hybridization used to type RFLPs. One of these advantages is sensitivity. Whereas microgram amounts of DNA are generally used to type RFLPs, nanogram amounts of genomic DNA are sufficient for routine genotyping of the block markers, and the polymerase chain reaction has recently been described as capable of amplifying DNA from a single template molecule (Saiki et al., 1988), Science 239:487-491). Enough DNA can be isolated from a single modest blood sample to type tens of thousands of (dC-dA).sub.n.(dG-dT).sub.n block markers.
Another advantage of the polymerase chain reaction is that the technique can be partially automated. For example, several commercial heating blocks are available which can automatically complete the temperature cycles used for the polymerase reaction. Automatic amplification reactions and the capability to analyze hundreds of markers on each polyacrylamide gel mean that the (dC-dA).sub.n.(dG-dT).sub.n markers can be analyzed faster than RFLPs and are more readily usable in practical applications such as identity testing.
Specific applications for the present invention include the identification of individuals such as in paternity and maternity testing, immigration and inheritance disputes, zygosity testing in twins, tests for inbreeding in man, evaluation of the success of bone marrow transplantation, quality control of human cultured cells, identification of human and animal remains, and testing of semen samples, blood stains, and other material in forensic medicine. In this application, the ability to run numerous markers in a single amplification reaction and gel lane gives this procedure the possibility of extreme efficiency and high throughput.
Another specific application is in human genetic analysis, particularly in the mapping through linkage analysis of genetic disease genes and genes affecting other human traits, and in the diagnosis of genetic disease through coinheritance of the disease gene with one or more of the polymorphic (dC-dA).sub.n.(dG-dT).sub.n markers.
A third specific application contemplated for the present invention is in commercial animal breeding and pedigree analysis. All mammals tested for (dC-dA).sub.n.(dG-dT).sub.n sequences have been found to contain them (Gross and Garrard, 1986, Mol. Cell. Biol. 6:3010-3013). Also, as a byproduct of efforts to develop (dC-dA).sub.n.(dG-dT).sub.n markers specific for human chromosome 19 from a library development from a hamster-human somatic cell hybrid, several hamster (dC-dA).sub.n.(dG-dT).sub.n markers have been developed (Weber and May, 1988, Am. J. Hum. Genet. 44(3):388-396).
A fourth specific application is in commercial plant breeding. Traits of major economic importance in plant crops can be identified through linkage analysis using polymorphic DNA markers. The present invention offers an efficient new approach to developing such markers for various plant species.
It is also contemplated that the present invention and method of characterization could be easily extended to include other tandemly repeated simple sequences which may be polymorphic. Examples include (dG-dA).sub.n.(dC-dT).sub.n, (dT-dA).sub.n.(dA-dT).sub.n, and even (dA).sub.n.(dT).sub.n.
BRIEF DESCRIPTION OF THE FIGURES
FIG. 1 is an example of a human (dC-dA).sub.n.(dG-dT).sub.n polymorphism showing the sequence of the amplified DNA, the primers used in the amplification, and an autoradiograph of a polyacrylamide gel loaded with amplified DNA from this marker.
FIG. 2 is an additional example of length polymorphisms in amplified fragments containing (dC-dA).sub.n.(dG-DT).sub.n sequences. Shown is an autoradiograph of a polyacrylamide gel.
FIG. 3 is an autoradiograph of a polyacrylamide gel loaded with DNA amplified from the human Mfd3, ApoAII locus using DNA from three unrelated individuals as template and labeled through three different approaches: labeling the interiors of both strands with .varies..sup.32 P-dATP, end-labeling the GT-strand primer with .sup.32 P phosphate, or end-labeling the CA-strand primer with .sup.32 P phosphate.
FIG. 4 is an example of the Mendelian inheritance of three different human (dC-dA).sub.n.(dG-dT).sub.n markers through three generations. Shown are the pedigree of this family, an autoradiograph of a gel loaded with the amplified DNA and a list of the genotypes of the individual family members.
FIG. 5 is an autoradiograph showing treatment of amplified DNA containing (dC-dA).sub.n.(dG-dT).sub.n sequences with either the Klenow fragment of DNA polymerase I or with T4 DNA polymerase.
FIG. 6 is an autoradiograph showing the effect of additional polymerase chain reaction cycles on amplified DNA for the Mfd3 marker from a single individual.
FIG. 7 is a graph illustrating the length distributions of (dC-dA).sub.n.(dG-dT).sub.n sequences. Abscissa values were taken as the number of repeats within the longest run of uninterrupted repeats, even for imperfect and compound repeat sequences. Repeats were counted beginning with either purines or pyrimidines, and half repeats were counted. Numbers of repeats were not averages but were rather taken from individual sequences. The open bars depict results for the GenBank sequences and the hatched bars results for the cloned sequences.
FIG. 8A illustrates a plot showing the informativeness of perfect repeat sequence (dC-dA).sub.n.(dG-dT).sub.n polymorphisms as a function of repeat length. FIG. 8A was obtained using weighted average numbers of repeats for each polymorphism.
FIG. 8B illustrates a plot showing the informativeness of perfect repeat sequence (dC-dA).sub.n.(dG-dT).sub.n polymorphisms as a function of repeat length similar to FIG. 8A. FIG. 8A was obtained using the number of repeats found within the original sequences used for PCR primer synthesis. The curves in the plots of FIGS. 8A and 8B are identical.
FIG. 9 is a chart illustrating the informativeness of imperfect repeat sequence (dC-dA).sub.n.(dG-dT).sub.n polymorphisms as a function of repeat length. Data for the open squares were obtained by counting the number of repeats within the entire imperfect repeat sequence. Data for the filled squares were obtained by counting only the repeats within the longest run of uninterrupted repeats. All variation in numbers of repeats for each marker was assumed to be confined to the longest run of uninterrupted repeats. The curve is identical to the one shown in FIG. 8.
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS DEFINITIONS:
All of the terms used herein should be known to one skilled in the art. However, the following definitions are provided to assist in providing a clear and consistent understanding scope and detail of the terms:
(CA).sub.n :
The subscript ".sub.n " refers to the number of repeats of the nucleotides in the parentheses immediately preceding the number. For example, (CA).sub.10 is a shorthand version of "CACACACACACACACACACA." (i.e., nucleotides 1-20 of SEQ. ID NO:59)
Heterozygosity:
The fraction of individuals that have different alleles at a particular locus (as opposed to two copies of the same allele). Heterozygosity is usually expressed in percent. Heterozygosity values range from 0 to 100%.
Hybridization:
The process in which a strand of nucleic acid joins with a complementary strand through base pairing.
Informativeness:
The informativeness of a human DNA polymorphism is a measure of the utility of the polymorphism. In general, higher informativeness means greater utility. Informativeness is usually defined in terms of either heterozygosity or Polymorphism Information Content (PIC).
Polymorphism information content (PIC): Because matings between two individuals who are both heterozygous but have identical genotypes are often not useful in genetic analysis, PIC was defined to more accurately reflect true informativeness (Botstein, D., et al., 1980, Am. J. Hum. Genet., 32, 314-331). PIC values range from to 1.0, and are smaller in value than heterozygosities. For markers that are highly informative (heterozygosities exceeding about 70%), the difference between heterozygosity and PIC is slight.
Polymorphism:
A condition in DNA in which the most frequent variant (or allele) has a population frequency which does not exceed 99%.
Primers: A substrate that is required for a polymerization reaction and that is structurally similar to the product of the reaction.
Development of Polymorphic DNA Marker.
Development of a polymorphic DNA marker based on length variations in blocks of (dC-dA).sub.n.(dG-dT).sub.n repeats involves a series of steps.
First, the sequence of a segment of DNA containing the (dC-dA).sub.n.(dG-dT).sub.n repeats must be determined. This is accomplished most commonly by selecting a genomic DNA clone through hybridization to synthetic poly(dC-dA).poly(dG-dT) and then subsequently sequencing that clone. This same step can also be accomplished simply by selecting a suitable sequence from the literature or from one of the DNA sequence databases such as GenBank. The latter approach is severely limited by the relatively small number of (dC-dA).sub.n.(dG-dT).sub.n sequences that have been published.
Second, once the sequence to be used is in hand, a pair of appropriate primers can be synthesized which are at least partially complementary to non-repeated, non-polymorphic sequences which flank the block of dinucleotide repeats on either side.
Third, these primer pairs are used in conjunction with a genomic DNA (or occasionally cloned DNA) template to amplify a small segment of DNA containing the repeats using the polymerase chain reaction (Saiki et al., 1985, Science 230:1350-1354 and U.S. Pat. No. 4,683,202 to Mullis et al, the substance of which is incorporated herein in its entirety). The process of polymerase chain reaction (PCR) uses an exponential process of replication. PCR permits a target sequence of DNA to be multiplied as quickly as a millionfold within hours. In PCR, a target sequence of double-stranded DNA is denatured into single-stranded form by a process of heating. Two small pieces of synthetic DNA, each complementing a sequence at one end of the target sequence, serve as primers and bind with their complementary sequences on the single strand. Polymerases start at each primer, copying the sequence of that strand and ultimately producing exact replicas of the target sequence. The product of each cycle then serves as a template for succeeding cycles, resulting in an exponential process of replication. After repeated cycles, the pool of pieces of DNA with the target sequence has been greatly multiplied. This amplified genetic material is then available for further analysis and use. The DNA is preferably labeled during the amplification process by incorporating radioactive nucleotides.
Fourth, the amplified DNA is resolved by polyacrylamide gel electrophoresis in order to determine the sizes of these fragments and hence the genotypes of the genomic DNA donor.
In a more particular aspect of the present invention, some or all of the polynucleotide primers are .sup.32 P or .sup.35 P labeled in any conventional manner, such as end labeling, interior labeling, or post reaction labeling. Alternative methods of labeling are fully within the contemplation of the invention such as biotin labeling or enzyme labeling (Matthews and Kricka 1988, Anal. Biochem. 169:1-25).
The practical outer limits of the length of the amplified DNA fragment is generally limited only by the resolving power of the particular separation system employed. The thin denaturing gels used in the work leading to this application are capable of resolving fragments differing by as little as 2 bases up to a total fragment length of about 300 bp. Use of longer gels and longer electrophoresis times could extend the resolving power up to perhaps 600 bp or even more. However the longer the fragment the lower the proportion of its length will be made up of the (dC-dA).sub.n.(dG-dT).sub.n sequences, and hence the more difficult the resolution.
Categorizing the Sequences.
Examination of over 100 human (dC-dA).sub.n.(dG-dT).sub.n sequences both from direct sequencing of clones selected by hybridization to poly(dC-dA).poly(dG-dT) and from computer searches of GenBank, show that the sequences differ from each other in two major respects. First, they contain different total numbers of repeats, and second, they contain variable numbers of sequence imperfections and or tandem repeats of other simple sequence families.
Using a set of precise classification rules, the sequences can be divided into three categories. The rules for categorizing the (dC-dA).sub.n.(dG-dT).sub.n sequences follow:
1. Perfect repeat sequences are defined as alternating, tandem CA repeats, i.e., CACACA . . . , without interruption and without adjacent repeats of another sequence.
2. Imperfect repeat sequences are defined as two or more runs of uninterrupted CA repeats separated by no more than 3 consecutive non-repeat bases. Terminal runs of uninterrupted CA repeats (outside of non-repeat bases) must each be at least 3 full repeats in length. The sequence (CA).sub.22 GACACAC (nucleotides 109-160 of SEQ. ID. NO:8) would therefore be classified as an imperfect repeat sequence with 25.5 total repeats, while the sequence (CA).sub.22 GACACA (nucleotides 109-159 of SEQ. ID NO:8) would be classified as a perfect repeat sequence with 22 repeats. Internal runs of uninterrupted CA repeats must be at least 1.5 repeats in length. The sequence (AC).sub.12 GTACATAA(AC).sub.10 (SEQ. ID. NO:457) would be scored as a single imperfect repeat sequence, but the sequence A(CA).sub.15 TACG(CA).sub.6 (SEQ. ID. NO:458) would be scored as two separate perfect repeat sequences.
3. Compound repeat sequences are defined as runs of CA repeats separated by no more than 3 consecutive non-repeat bases from a run of .gtoreq.5 uninterrupted dinucleotide or longer repeat length repeats of a sequence other than (dC-dA).sub.n.(dG-dT).sub.n or from .gtoreq.10 uninterrupted mononucleotides. Compound repeats are subclassified as perfect or imperfect depending on the status of the (dC-dA).sub.n.(dG-dT).sub.n block, perfect repeat sequences without interruptions in the runs of CA or GT dinucleotides (64% of total), imperfect repeat sequences with one or more interruptions in the run of repeats (25%), and compound repeat sequences with adjacent tandem simple repeats of a different sequence (11%).
The rules are written for the (CA).sub.n strands, but apply equally well to the (GT).sub.n strands.
Informativeness of the (dC-dA).sub.n.(dG-dT).sub.n Repeats.
To exemplify the informativeness of the (dC-dA).sub.n.(dG-dT).sub.n repeats, sequences from over 100 different polymorphic markers of the type which can be used within the present invention are listed in Table 1 and in Table 40. Each sequence represents only one allele at each specific locus. The first five sequences in Table 1 were taken from a computer search of GenBank; the remaining sequences were determined in the laboratory (see Example I). As can be seen in this compilation, the sequences exhibit substantial variation in the form of the tandem repeats. Some sequences, for example markers Mfd3, Mfd17 and Mfd23, contain only CA-GT repeats with no imperfections. Other sequences, for example Mfd2, Mfd7, and Mfd19 contain in addition to long runs of perfect CA-GT repeats, one or more imperfections in the run of repeats. These imperfections can be additional bases as in Mfd2 or more frequently GA-TC, AT-TA or CG-GC dinucleotide repeats as in Mfd7, Mfd13 and Mfd19. Homogenous runs of other dinucleotide repeats are often found in association with the CA-GT repeats like for example in Mfd5 and Mfd21. All of these repeat sequences can be used in this application.
Every human (dC-dA).sub.n.(dG-dT).sub.n sequence with 11 or more repeats that has been tested by the invention has been found to be polymorphic (over 100 sequences to date). Since there are an estimated 50,000-100,000 (dC-dA).sub.n.(dG-dT).sub.n blocks in the human genome, blocks are separated by an average spacing of 30,000-60,000 bp which is extremely tight in genetic terms. Two polymorphic markers spaced so that there is only 1% recombination between them are generally thought to be about 10.sup.6 bp apart; markers spaced only 50,000 bp apart on the average would be coinherited 99.95% of the time. This means that (dC-dA).sub.n.(dG-dT).sub.n markers should find significant usage in the genetic mapping and clinical diagnosis of human genetic diseases, much as RFLPs have been used in the mapping and diagnosis of diseases such as cystic fibrosis (White and Lalouel, 1988, Ann. Rev. Genet. 22:259-279).
The correspondence between polymorphisms which are relatively rare in the genome and the (dC-dA).sub.n.(dG-dT).sub.n sequences is very strong evidence that the repeats are mainly, if not entirely, responsible for the sequence length variations. Further evidence comes from the fact that amplified polymorphic fragments containing the (dC-dA).sub.n.(dG-dT).sub.n sequences always differ in size by multiples of 2 bp. Direct sequencing (see Example VI below) of allelic DNA also confirms this interpretation.
The informativeness of the (dC-dA).sub.n.(dG-dT).sub.n polymorphisms is good to very good, with heterozygosities ranging from 34-91%. Most of the (dC-dA).sub.n.(dG-dT).sub.n markers are therefore more informative than the two-allele RFLPs (Donis-Keller et al. (1987), Cell 51:319-339; Schumm et al. (1988), Am. J. Hum. Genet. 42:143-159). The number of alleles counted for the (dC-dA).sub.n.(dG-dT).sub.n markers tested to date has ranged from 4-11. Relatively high numbers of alleles also improve the usefulness of these markers. Alleles tend to differ by relatively few numbers of repeats, with the result that all alleles for a single marker may span a range in size of 20 bp or less. This means that amplified fragments from several different markers can be analyzed simultaneously on the same polyacrylamide gel lanes, greatly improving the efficiency of the amplification process and the ability to identify individuals using the test.
As of the date of this application, over 100 polymorphic repeat sequences have been discovered by the inventor. A summary of these repeat sequences is presented on Table 1. Tables 2-39 describe some of the repeat sequences in more detail, including the name of the locus, the source, the primer sequences, the frequency, the chromosomal location, and the Mendelian inheritance, and other identifying comments. Table 40 lists the complete nucleic acid sequence of some of the claimed polymorphic repeat sequences. - Marker Repeat Sequence Primer Sequence Mfd1 CATA(CA).sub.19 (SEQ. ID. NO: 53) GCTAGCCAGCTGGTGTTATT (SEQ. ID. NO: 54) ACCACTCTGGGAGAAGGGTA (SEQ. ID. NO: 55) Mfd2 (AC).sub.13 A(AC).sub.17 A (SEQ. ID. NO: 56) CATTAGGATGCATTCTTCTG (SEQ. ID. NO: 57) GTCAGGATTGAACTGGGAAC (SEQ. ID. NO: 58) Mfd3 (CA).sub.16 C (SEQ. ID. NO: 59) GGTCTGGAAGTACTGAGAAA (SEQ. ID. NO: 60) GATTCACTGCTGTGGACCCA (SEQ. ID. NO: 61) Mfd4 (AC).sub.12 GCACAA(AC).sub.13 A (SEQ. ID. NO: 62) GCTCAAATGTTTCTGCAACC (SEQ. ID. NO: 63) CTTTGTAGCTCGT GATGTGA (SEQ. ID. NO: 64) Mfd5 (CT).sub.7 (CA).sub.23 (SEQ. ID. NO: 65) CATAGCGAGACTCCATCTCC (SEQ. ID. NO: 66) GGGAGAGGGCAAAGATCGAT (SEQ. ID. NO: 67) Mfd6 (CA).sub.5 AA(CA).sub.13 (SEQ. ID. NO: 68) TCCTACCTTAATTTCTGCCT (SEQ. ID. NO: 69) GCAGGTTGTTTAATTTCGGC (SEQ. ID. NO: 70) Mfd7 (CA).sub.20 TA(CA).sub.2 (SEQ. ID. NO: 71) GTTAGCATAATGCCCTCAAG (SEQ. ID. NO: 72) CGATGGAGTTTATGTTGAGA (SEQ. ID. NO: 73) Mfd8 (AC).sub.20 A (SEQ. ID. NO: 74) CGAAAGTTCAGAGATTTGCA (SEQ. ID. NO: 75) ACATTAGGATTAGCTGTGGA (SEQ. ID. NO: 76) Mfd9 (CA).sub.17 G (SEQ. ID. NO: 77) ATGTCTCCTTGGTAAGTTA (SEQ. ID. NO: 78) AATACCTAGGAAGGGGAGGG (SEQ. ID. NO: 79) Mfd10 (AC).sub.14 A (SEQ. ID. NO: 80) CATGCCTGGCCTTACTTGC (SEQ. ID. NO: 81) AGTTTGAGACCAGCCTGCG (SEQ. ID. NO: 82) Mfd11 (AC).sub.23 A (SEQ. ID. NO: 83) ACTCATGAAGGTGACAGTTC (SEQ. ID. NO: 84) GTGTTGTTGACCTATTGCAT (SEQ. ID. NO: 85) Mfd12 (AC).sub.11 AT(AC).sub.8 A (SEQ. ID. NO: 86) GGTTGAGATGCTGACATGC (SEQ. ID. NO: 87) CAGGGTGGCTGTTATAATG (SEQ. ID. NO: 88) Mfd13 (CA).sub.4 CGCG(CA).sub.19 C (SEQ. ID. NO: 89) TTCCCTTTGCTCCCCAAACG (SEQ. ID. NO: 90) ATTAATCCATCTA AAAGCGAA (SEQ. ID. NO: 91) Mfd14 (AC).sub.23 A (SEQ. ID. NO: 92) AAGGATATTGTCCTGAGGA (SEQ. ID. NO: 93) TTCTGATATCAAAACCTGGC (SEQ. ID. NO: 94) Mfd15 (AC).sub.25 (SEQ. ID. NO: 95) GGAAGAATCAAATAGACAAT (SEQ. ID. NO: 96) GCTGGCCATATATATATTTAAACC (SEQ. ID. NO: 97) Mfd16 G(CG).sub.4 (CA).sub.5 TA(CA).sub.3 (TA).sub.2 (CA).sub.6 CCAA(CA).sub.21 AGAGATTAAAGGCTAAATTC (SEQ. ID. NO: 99) TTCGTAGTTGGTTAAAAT TG (SEQ. ID. NO: 100) (SEQ. ID. NO: 98) Mfd17 (AC).sub.23 (SEQ. ID. NO: 101) TTTCCACTGGGGAACATGGT (SEQ. ID. NO: 102) ACTCTTTGTTGAATTCCCAT (SEQ. ID. NO: 103) Mfd18 (AC).sub.18 (SEQ. ID. NO: 104) AGCTATCATCACCCTATAAAAT (SEQ. ID. NO: 105) AGTTTAACCATGTCTCTCCCG (SEQ. ID. No: 106) Mfd19 (AC).sub.8 AG(AC).sub.3 AG(AC).sub.24 TCAC(TC).sub.6 T (SEQ. ID. NO: 107) TCTAACCCTTTGGCCATTTG (SEQ. ID. NO: 108) GCTTGTTACATTGTTGCTTC (SEQ. ID. NO: 109) Mfd20 (AC).sub.17 (SEQ. ID. NO: 110) TTTGAGTAGGTGGCATCTCA (SEQ. ID. NO: 111) TTAAAATGTTGAAGGCATCTTC (SEQ. ID. NO: 112) Mfd21 (TA).sub.6 TT(TA).sub.2 TC(TA).sub.5 TT(TA).sub.3 CA(TA).sub.7 (CA).sub.8 TACATG(TA).sub.3 GCTCAGGAGTTCGAGATCA (SEQ. ID. NO: 114) CACCACACCCGACATTTTA (SEQ. ID. NO: 115) (SEQ. ID. NO: 113) Mfd22 (AC).sub.20 AG(AGAC).sub.5 AGA (SEQ. ID. NO: 116) TGGGTAAAGAGTGAGGCTG (SEQ. ID. NO: 117) GGTCCAGTAA GAGGACAGT (SEQ. ID. NO: 118) Mfd23 (AC).sub.20 (SEQ. ID. NO: 119) AGTCCTCTGTGCACTTTGT (SEQ. ID. NO: 120) CCAGACATGGCAGTCTCTA (SEQ. ID. NO: 121) Mfd24 (AC).sub.7 AGAG(AC).sub.14 A (SEQ. ID. NO: 122) AAGCTTGTATCTTTCTCAGG (SEQ. ID. NO: 123) ATCTACCTTGG CTGTCATTG (SEQ. ID. NO: 124) Mfd25 (AC).sub.11 (SEQ. ID. NO: 125) TTTATGCGAGCGTATGGATA (SEQ. ID. NO: 126) CACCACCATTGATCTGGAAG (SEQ. ID. NO: 127) Mfd26 (AC).sub.28 A (SEQ. ID. NO: 128) CAGAAAATTCTCTCTGGCTA (SEQ. ID. NO: 129) CTCATGTTCCTGGCAAGAAT (SEQ. ID. NO: 130) Mfd27 (CA).sub.9 AA(CA).sub.19 (GA).sub.7 (SEQ. ID. NO: 131) GATCCACTTTAACCCAAATAC (SEQ. ID. NO: 132) GGCATCAACTTGAACAGCAT (SEQ. ID. NO: 133) Mfd28 (AC).sub.10 AG(AC).sub.21 A (SEQ. ID. NO: 134) AACACTAGTGACATTATTTTCA (SEQ. ID. NO: 135) AGCTAGGCC TGAAGGCTTCT (SEQ. ID. NO: 136) Mfd29 (AC).sub.19.5 (SEQ. ID. NO: 137) AGCTCCCTCGAGATGCACT (SEQ. ID. NO: 138) TTCTTTGCTTTACATGTGGC (SEQ. ID. NO: 139) Mfd30 (AC).sub.18.5 (SEQ. ID. NO: 140) CCATGTCCCATATCTCTACA (SEQ. ID. NO: 141) TGAAATCACTGATGACAATG (SEQ. ID. NO: 142) Mfd31 (AC).sub.13 A (SEQ. ID. NO: 143) TAATAAAGGAGCCAGCTATG (SEQ. ID. NO: 144) ACATCTGATGTAAATGCAAGT (SEQ. ID. NO: 145) Mfd32 (AC).sub.12 A (SEQ. ID. NO: 146) AGCTAGATTTTTACTTCTCTG (SEQ. ID. NO: 147) CTGGTTGTACATGCCTGAC (SEQ. ID. NO: 148) Mfd33 (AC).sub.14 AT(AC).sub.13 (SEQ. ID. NO: 149) AGCCTGGGAGTCAGAGTGA (SEQ. ID. NO: 150) AGCTCCAAATCCAAAGACGT (SEQ. ID. NO: 151) Mfd34 (AC).sub.4 AT(AC).sub.15 (SEQ. ID. NO: 152) GGTTTCTTTTTTCTAGTTCTTC (SEQ. ID. NO: 153) TCATATAGCCTTTTGTTTGCA (SEQ. ID. NO: 154) Mfd35 In preparation GTGGAGAGTAAGACTCTGTC (SEQ. ID. NO: 155) TGATGCAACAC AGGAGACCT (SEQ. ID. NO: 156) Mfd36 (AC).sub.15 AT(AC).sub.6 A (SEQ. ID. NO: 157) AGCTATAATTGCATCATTGCA (SEQ. ID. NO: 158) TGGTCTATAA CTGGTCTATG (SEQ. ID. NO: 159) Mfd37 (AC).sub.10 A (SEQ. ID. NO: 160) AAAAGTGTGTTACTTTCAGAAC (SEQ. ID. NO: 161) ACAAGGTGACAAGGTGCCTA (SEQ. ID. NO: 162) Mfd38 (AT).sub.14 (AC).sub.14.5 (SEQ. ID. NO: 163) ATCTCTGTTCCCTCCCTGTT (SEQ. ID. NO: 164) CTTATTGGCCTTGAAGGTAG (SEQ. ID. NO: 165) Mfd39 (TC).sub.12.5 GTT(TC).sub.11.5 (CA).sub.14 A(CA).sub.5.5 (SEQ. ID. NO: 166) GGGTTGGTTGTAAATTAAAAC (SEQ. ID. NO: 167) TGTCAAATACTTAAGCACA G (SEQ. ID. NO: 168) Mfd40 (CA).sub.13 C(CA).sub.6 T(AC).sub.5 (SEQ. ID. NO: 169) GGCATCATTTTAGAAGGAAAT (SEQ. ID. NO: 170) ACATTTGTTCAGGACCAAAG (SEQ. ID. NO: 171) Mfd41 (AC).sub.17 (SEQ. ID. NO: 172) CAGGTTCTGTCATAGGACTA (SEQ. ID. NO: 173) TTCTGGAAACCTACTCCTGA (SEQ. ID. NO: 174) Mfd42 (CA).sub.16 T(AC).sub.3.5 (SEQ. ID. NO: 175) GGCCTCAAAGAATCCTACAG (SEQ. ID. NO: 176) GACACGTAGTTGCTTATTAC (SEQ. ID. NO: 177) Mfd43 In preparation TTGGAAGCCTTAGGAAGTGC (SEQ. ID. NO: 178) AAGAATTCTAG TTTCAATACCG (SEQ. ID. NO: 179 Mfd44 (CA).sub.17 (SEQ. ID. NO: 180) GTATTTTTGGTATGCTTGTGC (SEQ. ID. NO: 181) CTATTTTGGAATATATGTGCCT (SEQ. ID. NO: 182) Mfd45 (CA).sub.20.5 (SEQ. ID. NO: 183) TCCAGCAGAGAAAGGGTTAT (SEQ. ID. NO: 184) GGCAAAGAGAACTCATCAGA (SEQ. ID. NO: 185) Mfd46 (AC).sub.25 (SEQ. ID. NO: 186) AAAAGGAAGAATCAAATAGAC (SEQ. ID. NO: 187) ATATATTTAAACCATTTGAAAG (SEQ. ID. NO: 188) Mfd47 (AC).sub.17.5 (SEQ. ID. NO: 189) ACAGAGTGAGACCGTGTAAC (SEQ, ID. NO: 190) AGAGAAGCATCTCACTTAGT (SEQ. ID. NO: 191) Mfd48 (AC).sub.17 (SEQ. ID. NO: 192) TGTCTCCTGCTGAGAATAG (SEQ. ID. NO: 193) TAATATCCAAACCACAAAGGT (SEQ. ID. NO: 194) Mfd49 (CA).sub.22 (SEQ. ID. NO: 195) GATAAATGCCAAACATGTTGT (SEQ. ID. NO: 196) TGCTCTCAGGATTTCCTCCA (SEQ. ID. NO: 197) Mfd50 (CA).sub.19 (SEQ. ID. NO: 198) ACATTCTAAGACTTTCCCAAT (SEQ. ID. NO: 199) AGAGCATGCACCCTGAATTG (SEQ. ID. NP: 200) Mfd51 In preparation AGCTGATACACCACTTCTGA (SEQ. ID. NO: 201) GACAGAAATAT CCTTCCCAT (SEQ. ID. NO: 202) Mfd52 (AC).sub.18 TTG(CA).sub.3 (SEQ. ID. NO: 203) AAATCAGACAAGTACAGGTG (SEQ. ID. NO: 204) ATGAACTTGTTCTGGGAGGA (SEQ. ID. NO: 205) Mfd53 In preparation TGCCCCTGCACTCTAGCCT (SEQ. ID. NO: 206) GCTATCAACAAG CTTTAGGT (SEQ. ID. NO: 207) Mfd54 In preparation CTGACAGGTTGAGGCTGCA (SEQ. ID. NO: 208) CAGTTTGTATGT ATGTTTGGA (SEQ. ID. NO: 209) Mfd55 (AC).sub.16 (SEQ. ID. NO: 210) GTCAACATAGTGAGACCCCA (SEQ. ID. NO: 211) ATCCAGCCTGTAACACATTC (SEQ. ID. NO: 212) Mfd56 In preparation CTGGTGAATTCAAACAACCT (SEQ. ID. NO: 213) TTTTCTCTGAC ACCTCAACT (SEQ. ID. NO: 214) Mfd57 (CA).sub.15.5 (SEQ. ID. NO: 215) GATCTATCCCCTCACTTACG (SEQ. ID. NO: 216) TATGAACAGAACAGTGGAGC (SEQ. ID. NO: 217) Mfd58 (CA).sub.16.5 (SEQ. ID. NO: 218) CTCATTTGAAGACTGCAGCA (SEQ. ID. NO: 219) AGGGCTTCCTGTCCATCTA (SEQ. ID. NO: 220) Mfd59 (AC).sub.23.5 (SEQ. ID. NO: 221) AAGAACCATGCGATACGACT (SEQ. ID. NO: 222) CATTCCTAGATGGGTAAAGC (SEQ. ID. NO: 223) Mfd60 In preparation GTCCCTAGCCTCCCAGCAT (SEQ. ID. NO: 224) CAAGAGCGAAAG TCCGTCTC (SEQ. ID. NO: 225) Mfd61 (CA).sub.23 (SEQ. ID. NO: 226) GCCCTATAAAATCCTAATTAAC (SEQ. ID. NO: 227) GAAGGAGAATTGTAATTCCG (SEQ. ID. NO: 228) Mfd62 (AC).sub.21 (SEQ. ID. NO: 229) AGCTTTACAGATGAGACCAG (SEQ. ID. NO: 230) CAGCCAATTTCTTGAGTCCG (SEQ. ID. NO: 231) Mfd63 (CA).sub.20.5 (SEQ. ID. NO: 232) CAAAACCAAAAAACCAAAGGC (SEQ. ID. NO: 233) CAATCTGTGACAGTTTCTCA (SEQ. ID. NO: 234) Mfd64 (AC).sub.15.5 (SEQ. ID. NO: 235) ACGAACATTCTACAAGTTAC (SEQ. ID. NO: 236) TTTCAGAGAAACTGACCTGT (SEQ. ID. NO: 237) Mfd65 (CA).sub.14.5 (SEQ. ID. NO: 238) GCAAACCACAATGGAATGCA (SEQ. ID. NO: 239) CTTTACTTCCTTTGCCTCAG (SEQ. ID. NO: 240) Mfd66 (AC).sub.22 (SEQ. ID. NO: 241) GCCCCTACCTTGGCTAGTTA (SEQ. ID. NO: 242) AACCTCAGCTTATACCCAAG (SEQ. ID. NO: 243) Mfd67 (TC).sub.12 (AC).sub.18 (SEQ. ID. NO: 244) ATCCTGCCCTTATGGAGTGC (SEQ. ID. NO: 245) CCCACTCCTCTGTCATTGTA (SEQ. ID. NO: 246) Mfd68 In preparation ATGTATAGAATTCCATTCCTG (SEQ. ID. NO: 247) TAAAATCAAG TGTTGATGTAG (SEQ. ID. NO: 248) Mfd69 (AC).sub.18.5 A(AC).sub.3 (SEQ. ID. NO: 249) TAGCTGGTGCATAAGCTCAC (SEQ. ID. NO: 250) GTTAGTGGAAGAGCAGAGC (SEQ. ID. NO: 251) Mfd70 In preparation AACATAGTGAAACCCCATCT (SEQ. ID. NO: 252) GTGCCACTACA TGCAGCTA (SEQ. ID. NO: 253) Mfd71 In preparation CCAAACTACAATACCAGCTA (SEQ. ID. NO: 254) CTTGATTTGAG TATAACCAATA (SEQ. ID. NO: 255) Mfd72 (AC).sub.17 G(GA).sub.8 (SEQ. ID. NO: 256) AGAAGACATAAGGATACTGC (SEQ. ID. NO: 257) GATCCCAACTATTTCTTTCT (SEQ. ID. NO: 258) Mfd73 In preparation CCTGGAAAAATGGCTCACC (SEQ. ID. NO: 259) GGAAAATCAGTC TCTAGTTG (SEQ. ID. NO: 260) Mfd74 In preparation TTTCACCTCCTTGGCTTTGT (SEQ. ID. NO: 261) ATCCCTTTTAC AACAACTGC (SEQ. ID. NO: 262) Mfd75 In preparation CTCACTCATGCTTGTTTTGA (SEQ. ID. NO: 263) GATCACGTCAG ACTGGGCT (SEQ. ID. NO: 264) Mfd76 In preparation CCTGTGAGACAAAGCAAGAC (SEQ. ID. NO: 265) GACATTAGGCA CAGGGCTAA (SEQ. ID. NO: 266) Mfd77 In preparation ATAGACTTCCAGACAGATAG (SEQ. ID. NO: 267) CCTCTCTCATT CCTGGTACT (SEQ. ID. NO: 268) Mfd78 In preparation GAATCCATAGCTGTACTCCA (SEQ. ID. NO: 269) AATTGTCTATG GTCCCAGCA (SEQ. ID. NO: 270) Mfd79 (AC).sub.15.5 (SEQ. ID. NO: 271) GATAAAACTGCATAGAAATGCG (SEQ. ID. NO: 272) CAACTGGGATATTGACATTG (SEQ. ID. NO: 273) Mfd80 In preparation TTGAGGCTGCAGTGAGCTAT (SEQ. ID. NO: 274) ATGTTGTGTTT TCACAGCAG (SEQ. ID. NO: 275) Mfd81 In preparation GCACTCATGTCACCAATTCT (SEQ. ID. NO: 276) ATAGTCAATGG TTAATGCTC (SEQ. ID. NO: 277) Mfd82 In preparation AGCTTGGGTGCAAGAAGAG (SEQ. ID. NO: 278) GATCCCATTATT TAAAAGTGTA (SEQ. ID. NO: 279) Mfd83 In preparation GATCTCATGTGCTCAGTTTA (SEQ. ID. NO: 280) CCAAAAAAGTG CAAATTTAGAGT (SEQ. ID. NO: 281) Mfd84 (AC).sub.5 AT(AC).sub.2 AGATT(AC).sub.2 GG(AC).sub.17 (SEQ. ID. NO: 282) AATGTCCTTGTACTTAGGAT (SEQ. ID. NO: 283) CACTTAATATCTCAATGTATAC (SEQ. ID. NO: 284) Mfd85 In preparation GATCCTTTTCATCTTCTGAC (SEQ. ID. NO: 285) GAGGGACGGAG CAACTGAT (SEQ. ID. NO: 286) Mfd86 In preparation CAACATAGCAAGACCCTGTC (SEQ. ID. NO: 287) GCACATGCCAC CAAGACAAG (SEQ. ID. NO: 288) Mfd87 In preparation TCAAAAGCTTGTAATTGGAG (SEQ. ID. NO: 289) TGCAATCTGTA AGCATTCCT (SEQ. ID. NO: 290) Mfd88 In preparation ACCTGAGTGTTCATCAATAC (SEQ. ID. NO: 291) TCCAGAATCAT CCATGTTGT (SEQ. ID. NO: 292) Mfd89 In preparation GTCTTGTTTGCTGGCTCCA (SEQ. ID. NO: 293) AGCTATGAAGTG GGAGTTCA (SEQ. ID. NO: 294) Mfd90 In preparation ATTTTGGATGAGCCAAGCCT (SEQ. ID. NO: 295) ATCTGTATATA TGTGTACCTG (SEQ. ID. NO: 296) Mfd91 In preparation CTACATATTTCTAAATACATGC (SEQ. ID. NO: 297) ACTTAGTAG TTTTAAGCAGGA (SEQ. ID. NO: 298) Mfd92 In preparation ATTTCCACCCACTTCTGGT (SEQ. ID. NO: 299) GATGGTGTTGAG AATTAGGC (SEQ. ID. NO: 300) Mfd93 (AAAT).sub.6 TT(TA).sub.7 (CA).sub.13 (SEQ. ID. NO: 301) GACAGAGTGAGACTCCATCT (SEQ. ID. NO: 302) CTTCCCATTTTCAATCCCTAG (SEQ. ID. NO: 303) Mfd94 AAACGCACAG(AC).sub.21.5 (SEQ. ID. NO: 304) AGTCTTTCTCCTGTTGTGCT (SEQ. ID. NO: 305) CCCTAAGGACAGAACAAGTG (SEQ. ID. NO: 306) Mfd95 GGAGATTTGG(AC).sub.23 (SEQ. ID. NO: 307) TAGGCCCTACTGCAATAATG (SEQ. ID. NO: 308) CTTTATCTTCACACAGCTTC (SEQ. ID. NO: 309) Mfd96 In preparation TCAACAATGGCCGAGGTTA (SEQ. ID. NO: 310) AACCTGACACCA TGCTCCT (SEQ. ID. NO: 311) Mfd97 CTCTCTCTCT(CA).sub.11.5 (SEQ. IS. NO: 312) TTCTATTTCTGAAGGTGAACTA (SEQ. ID. NO: 313) ATAGTTACCATCAGTCACTG (SEQ. ID. NO: 314) Mfd98 In preparation TCTGGAGACCACTAACTGTA (SEQ. ID. NO: 315) ACTCTCCATGA GTCCTGATG (SEQ. ID. NO: 316) Mfd99 TCTCTATCTT(CA).sub.20.5 (SEQ. ID. NO: 317) ATGAGCTAATTCTCTATCTTC (SEQ. ID. NO: 318) TAGCCTACATAAAGGAGGGT (SEQ. ID. NO: 319) Mfd100 In preparation GGAGCCAAATACTAAATTCT (SEQ. ID. NO: 320) TTAGGCACTT TAATCAGGCT (SEQ. ID. NO: 321) Mfd101 (AC).sub.17 (SEQ. ID. NO: 322) CATAAAAGGCTTATTGGTTTG (SEQ. ID. NO: 323) CAAAACAGAGAACAGAGTAG (SEQ. ID. NO: 324) Mfd102 (AC).sub.19 AA(AC).sub.5 A (SEQ. ID. NO: 325) AGGAGAGCTAGAGCTTCTAT (SEQ. ID. NO: 326) GTTTCAACATG AGTTTCAGA (SEQ. ID. NO: 327) Mfd103 GCAGTAAAAG(CA).sub.20 (SEQ. ID. NO: 328) CAGATAAACTAATACAAGCAG (SEQ. ID. NO: 329) CTCTGCCTCCCAAAGTGCT (SEQ. ID. NO: 330) Mfd104 TTATATATAT(AC).sub.14.5 (SEQ. ID. NO: 331) GATCATGTGAGTTAATACTTAA T (SEQ. ID. NO: 332) TCAGCTGCCTGTATTACTCA (SEQ. ID. NO: 333) Mfd105 TCAAACACAA(AC).sub.16 (SEQ. ID. NO: 334) GATCCTGTCTCAAACACAAAC (SEQ. ID. NO: 335) AAGTCTTCAGCTTTATCAAC (SEQ. ID. NO: 336) Mfd106 TCTTCCCCCA(AC).sub.20.5 (SEQ. ID. NO: 337) GATCTGTCTTCCCCCAAC (SEQ. ID. NO: 338) TTTCATGTTGCAGTCAGAGC (SEQ. ID. NO: 339) Mfd107 TGCCCGGCCT(AC).sub.16 (SEQ. ID. NO: 340) CCCAAAGTACTGGGATTACA (SEQ. ID. NO: 341) TTCAAGTGTTACTGTACTGC (SEQ. ID. NO: 342) Mfd108 GTGGCTAAAT(AC).sub.16 (SEQ. ID. NO: 343) GCCTCTGAAGTGGCTAAATA (SEQ. ID. NO: 344) CCCCTCACCACATCACTTG (SEQ. ID. NO: 345) Mfd109 GGGAAATAGG(CA).sub.18 (SEQ. ID. NO: 346) GACACAGAGAAGGGAAATAG (SEQ. ID. NO: 347) TCCCATATCCTATGTAGAAG (SEQ. ID. NO: 348) Mfd110 (ATTT).sub.11 AT (SEQ. ID. NO: 349) GCTAGAGGGAGGTTTAATTG (SEQ. ID. NO: 350) AATTAGCCAGGTGTTGTGGT (SEQ. ID. NO: 351) Mfd111 TGAGACCCTG(AC).sub.15.5 (SEQ. ID. NO: 352) AACCAAGATTGTGCCACTG (SEQ. ID. NO: 353) GATCATGACTCTTTTGTG (SEQ. ID. NO: 354) Mfd112 CCACCCCCAG(CA).sub.24.5 (SEQ. ID. NO: 355) GATCCATGCCCACCCCCA (SEQ. ID. NO: 356) CCTCTCAGACTCATCCCAC (SEQ. ID. NO: 357) Mfd113 (AC).sub.18 (SEQ. ID. NO: 358) CTGCTGACTTTGACTCAGTA (SEQ. ID. NO: 359) GGTCCTGAGCAGGTCTCTTC (SEQ. ID. NO: 360) Mfd114 GTTATCCATT(AC).sub.19.5 (SEQ. ID. NO: 361) TGGCATCTCTAATCATACTG (SEQ. ID. NO: 362) GACTAAAACATTGCAGAATAC (SEQ. ID. NO: 363) Mfd115 ATAGAGAAGG(AC).sub.17.5 (SEQ. ID. NO: 364) AGAAAATAAGAATAGAGAAGG (SEQ. ID. NO: 365) CAAGAACTATGTTATTGGGA (SEQ. ID. NO: 366) Mfd116 CCCCCACCCA(AC).sub.20 (SEQ. ID. NO: 367) CTGCACTAGAAAGGCAGAGT (SEQ. ID. NO: 368) TGCAGCACCAAACACCAAGT (SEQ. ID. NO: 369) Mfd117 GCAGCAACAT(AC).sub.16.5 (SEQ. ID. NO: 370) ACAAGAGCACATTTAGTCAG (SEQ. ID. NO: 371) AGCTTCATTTTTCCCTCTAG (SEQ. ID. NO: 372) Mfd118 (AC).sub.15 (SEQ. ID. NO: 373) CTTTCTTATAGTTAAGGTTAGC (SEQ. ID. NO: 374) TAGCATCAGAAGACCTGGC (SEQ. ID. NO: 375) Mfd119 (AC).sub.16 (SEQ. ID. NO: 376) GAATCTTAAGTAGTTATCCCTC (SEQ. ID. NO: 377) CTACAAAAAGTCAGATACCT (SEQ. ID. NO: 378) Mfd120 (TC).sub.5 (AC).sub.20 (SEQ. ID. NO: 379) CAATGACTTCAAGCACTAAG (SEQ. ID. NO: 380) TCAGAGGTTGAGGCTGAAG (SEQ. ID. NO: 381) Mfd12l (TC).sub.5 (CA).sub.23.5 (SEQ. ID. NO: 382) GATCTGGGTATGTCTTTCTG (SEQ. ID. NO: 383) ACTGGGACTCTAACTAATGT (SEQ. ID. NO: 384) Mfd122 TTTACAGTAG(CA).sub.17 (SEQ. ID. NO: 385) ATGCAGAATCTACAAGGACC (SEQ. ID. NO: 386) CTTTAACATCCTTTAACAGC (SEQ. ID. NO: 387) Mfd123 (AC).sub.21.5 (SEQ. ID. NO: 388) AGCAGCTATTATGGAATTGC (SEQ. ID. NO: 389) CAACATATGCAAGGTGCCTA (SEQ. ID. NO: 390) Mfd124 ATTTAGCATA(AC).sub.20.5 (SEQ. ID. NO: 391) TGCTTAAACAGAAAAGTAGC (SEQ. ID. NO: 392) TAAAACAGTACCCAGTACCT (SEQ. ID. NO: 393) Mfd125 CTGAAACAAA(AC).sub.23 (POST ALU) (SEQ. ID. NO: 394) TGAGACCCTGTCTCTGAAAC (SEQ. ID. NO: 395) TGTATGGGCTCTTGAAATTG (SEQ. ID. NO: 396)
The following examples are intended to illustrate, but not limit, the scope of the invention.
EXAMPLES
Example I
This example describes the method used to identify and isolate specific (dC-dA).sub.n.(dG-dT).sub.n fragments.
General Procedure
Total human genomic DNA or total DNA from a chromosome 10-specific large insert page library (LL19NL01) was digested to completion with Sau3A I, Alu I, Taq I, or a combination of Sau3A I and Taq I. DNA fragments ranging in size from about 150 to 400 base pairs were purified by preparative agarose gel electrophoresis (Weber at al. (1988), J. Biol. Chem. 3:11321-11425), and ligated into mp18 or mp19 m13 vectors. Nitrocellulose plaque lifts (Benton and Davis (1977), Science 196:180-182) prepared from the resulting clones were screened by hybridization to synthetic poly(dC-dA).poly(dG-dT) which had been nick-translated using both .varies..sup.32 P-dATP and .varies..sup.32 P-dTTP to a specific activity of about 5.times.10.sup.7 cpm/ug. Hybridizations were carried out in 6.times.SSC, pH 7.0, 2.5 mM EDTA, 5.0% (v/v) O'Darby Irish Cream Liqueur (Elbrecht, A., March 1987, B. M. Biochemica, 12-13, at 60.degree. C. After hybridization, filters were washed in 2.times.SSC, 25 mM NaPO.sub.4, 0.10% SDS, 5.0 mM EDTA, 1.5 mM Na.sub.4 P.sub.2 O.sub.7, pH 7.0, and then in 1SSC, 0.10% SDS, 5.0 mM EDTA, pH 7.0. Phage from the first screen were usually diluted and then screened a second time to insure plaque purity, Single stranded DNA was isolated from the positive clones and sequenced as described (Biggin et al. (1983), Proc. Natl. Acad, Sci. USA 80:3063-3965).
GenBank DNA databases were screened for the presence of sequences with (dC-dA).sub.6 or (dG-dT).sub.6 (see, for instance, nucleotides 1-12 of SEQ. ID. NO.:59 using the QUEST program made available by Intelligenetics Inc. through the national BIONET computing network. Since the sequences of only one of the two strands of each DNA fragment are compiled in GenBank, separate screens for both CA and GT repeats were necessary.
Results.
The hybridization procedure was used to isolate and sequence over one hundred (dC-dA).sub.n.(dG-dT).sub.n blocks and the DNA immediately flanking the repeats. Examples are listed in Table 1 (Mfd6-125). Numbers of (dC-dA).sub.n.(dG-dT).sub.n dinucleotide repeats within the blocks ranged from 10 to over 30. Many of the blocks had imperfect repeats or were adjacent to tandem repeats with different sequences.
(dC-dA).sub.n.(dG-dT).sub.n sequences obtained from the GenBank screens (Mfd1-5) were similar to those obtained through the hybridization procedure, except that sequences containing as few as six repeats could be selected.
Example II
In this example a subset of the sequences isolated and identified as in Example I were amplified and labeled using the polymerase chain reaction, and were then resolved on polyacrylamide gels to demonstrate length polymorphisms in these sequences.
General Procedures.
Oligodeoxynucleotide primers were synthesized on a Cyclone DNA synthesizer (Biosearch, Inc., San Rafael, Calif.). Primers were 19-22 total bases in length, and contained 7-11 G+C bases. Self-complementary regions in the primers were avoided.
Genomic DNA was isolated from nucleated blood cells as described (Aidridge et al., 1984, Am. J. Hum. Genet. 36:546-564). Standard polymerase chain reactions (Saiki et al., 1985, Science 230:1350-1354; Mullis and Faloona, 1987, Method Enzymol. 155:335-350; Saiki et al., 1988, Science 239:487-491) were carried out in a 25 ul volume containing 10-20 ng of genomic DNA template, 100 ng each oligodeoxynucleotide primer, 200 uM each dGTP, dCTP and dTTP, 2.5 uM dATP, 1-2 uCi of .varies..sup.32 P-dATP at 800 CI/mmole or .varies..sup.35 S-dATP at 500 Ci/mmole, 50 mM KCl, 10 mM Tris, pH 8.3, 1.5 mM MgCl.sub.2, 0.01% gelatin and about 0.75 unit of Taq polymerase (Perkin Elmer Cetus, Norwalk, Conn.). Samples were overlaid with mineral oil and processed through 25 temperature cycles consisting of 1 min at 94.degree. C. (denaturation), 2 min at 55.degree. C. (annealing), and 2.5 min at 72.degree. C. (elongation). The last elongation step was lengthened to 10 min.
Results shown in FIG. 1 were obtained using conditions slightly different that the standard conditions. Templates were 100-200 ng of genomic DNA, annealing steps were 2.5 min at 37.degree. C., elongation steps were 3.5 min at 72.degree. C., and .varies..sup.35 S-dATP was added after the 18th cycle rather than at the beginning of the reactions. The plasmid DNA sample was amplified starting with 50 pg of total plasmid DNA as template.
Primers for the Mfd15 marker shown in Figure 2 are listed in Table 1. Primers for the Mfd26 marker also shown in FIG. 2 are CAGAAAATTCTCTCTGGCTA (SEQ ID NO:129) and CTCATGTTCCTGGCAAGAAT (SEQ ID NO:130), and primers for the Mfd31 marker are TAATAAAGGAGCCAGCTATG (SEQ ID N0:144) and ACATCTGATGTAAATGCAAGT (SEQ ID NO:145).
Aliquots of the amplified DNA were mixed with two volumes of formamide sample buffer and electrophoresed on standard denaturing polyacrylamide DNA sequencing gels. Exposure times were about 2 days. Gel size standards were dideoxy sequencing ladders produced using m13, mp10 or mp8 DNA as template.
Results
FIG. 1 shows the amplified DNA fragments for the IGF1 (dC-dA).sub.n.(dG-dT).sub.n marker in seven unrelated individuals (1-7). Z represents the most frequent allele; Z-2, the allele that is two bp larger than the most frequent; Z-2, the allele that is 2 bp small, etc. K indicates Kpn I digestion of amplified samples 1 and 7. Kpn digestions reduce the number of bands to half the original number because the CA strand, which normally migrates with an apparent size of about four bases less than the GT strand, is after Kpn I digestion, four bases longer than the GT strand resulting in co-migration of the two strands. P refers to DNA amplified from a plasmid DNA sample containing the IGF1 (dC-dA).sub.n.(dG-dT).sub.n block. Sizes of the DNA fragments in bases are indicated on the left. At the top of the figure are shown the sequence of amplified DNA along with the primer sequences and the site of Kpn I cleavage.
Because the CA and GT strands of the amplified DNA fragments migrate with different mobilities under the denaturing electrophoresis conditions (see Example III below), homozygotes yield two bands and heterozygotes four bands. The band corresponding to the faster moving CA strand is more intense on the autoradiographs than the band for the slower GT strand because the adenine content of the CA strand is higher and labeling is with .varies..sup.35 S-dATP. Two of the seven individuals shown in FIG. 1 (1 and 3) were homozygous for the predominant allele (Z) of the IGF1 (dC-dA).sub.n.(dGdT).sub.n block; the remainder were heterozygotes of various types.
Proof that the amplified DNA was really from the IGF1 gene and not from some other portion of the genome includes; that the amplified DNA was of the general expected size range for the primers used, that the amplified DNA hybridized to nick-translated poly(dC-dA).poly(dG-dT) (not shown), that this DNA was cleaved by a restriction enzyme, Kpn I, at the expected position (FIG. 1, lanes 1K and 7K), and that plasmid DNA containing the IGF1 sequence could be used as a polymerase chain reaction template to yield DNA of the same size as was amplified from the genomic DNA templates (FIG. 1, lane P).
FIG. 2 shows additional examples of polymorphic amplified DNA fragments containing (dC-dA).sub.n.(dG-dT).sub.n sequences. In this case three different markers fragments, Mfd15, Mfd26 and Mfd31 were amplified simultaneously from genomic DNA templates from several different individuals.
Example III
Comparison of different labeling approaches.
General Procedures.
The ApoAII (Mfd3) CA or GT strand oligodeoxynucleotide primers were end-labeled for 1 h at 37.degree. C. in a 50 ul reaction containing 90 pmoles (600 ng) of primers, 33 pmoles of y.sup.32 P-ATP at 3000 Ci/mmole, 10 mM MgCl.sub.2, 5 mM DTT 50 mM Tris, pH 7 6, and 50 unites of T4 polynucleotide kinase. Polymerase chain reactions were carried out in 25 ul volumes with 50 ng of end-labeling primer and 86 ng of each unlabeled primer. Interior labeling was performed as in Example II.
Results
Rather than labeling the amplified DNA throughout the interiors of both strands, one or both of the polymerase chain reaction primers can be end-labeled using polynucleotide kinase. FIG. 3 shows the results of such an experiment using as template, DNA from three different individuals (1-3) and labeling throughout the interiors of both strands versus labeling of the GT strand primer only versus labeling of the CA strand primer only. Individual 1 is a homozygote and individuals 2 and 3 are heterozygotes. Because of strand separation during the denaturing gel electrophoresis, labeling of both strands produces two major bands per allele on the autoradiograph, whereas labeling of the GT strand primer gives predominantly only the upper band for each allele and labeling the CA strand primer gives predominantly only the lower band for each allele. Additional fainter bands on the autoradiograph are artifacts of the polymerase chain reaction and will be discussed in Example IV.
Example IV
Estimates of informativeness and allele frequencies for the (dC-dA).sub.n.(dG-dT).sub.n markers.
General Procedure
Estimates of PIC (polymorphism information content) (Bostein et al., 1980, Amer. J. Hum. Genet. 32:314-331 and heterozygosity were obtained by typing DNA from 41-45 unrelated Caucasians for markers Mfd1-Mfd4, or by typing DNA from 75-78 parents of the 40 CEPH (Centre d'Etude du Polymorphisme Humain, Paris, France) reference families for markers Mfd514 Mfd10. The CEPH families are from the U.S.A., France and Venezuela. Estimates of allele frequencies were calculated from the same data.
Results.
PIC and heterozygosity values for the first 10 (dC-dA).sub.n.(dG-dT).sub.n markers are shown in Table 41. Values ranged from 0.31 to 0.80 with an average PIC of 0.54.+-.0.14 and average heterozygosity of 0.56.+-.0.15. The informativeness of the (dC-dA).sub.n.(dG-dT).sub.n block markers is generally superior to standard unique sequence probe polymorphisms (Gilliam et al., 1987, Nucleic Acids Res. 15:4617-4627; Schumm et al., 1988, Am. J. Hum. Genet. 42:143-159) and as good as many minisatellite polymorphisms (Nakamura et al., 1987, Science 235:1612-1622). Considering the vast number of (dC-dA).sub.n.(dG-dT).sub.n blocks in the human genome, it is likely that a subset of up to several thousand can be identified with average heterozygosities of 70% or better.
The number of different alleles detected for the first ten markers (Table 41) ranged from 4 to 11. Alleles always differed in size by multiples of two bases (from CA strand to CA strand bands), consistent with the concept that the number of tandem dinucleotide repeats is the variable factor. Allele frequencies for the first ten markers are shown in Table 42. For most of the test markers, major alleles were clustered in size within about 6 bp on either side of the predominant allele. Amplified fragments must be small enough so that alleles differing in size by as little as two bases can easily be resolved on the polyacrylamide gels. The size differences between the largest and smallest alleles were .ltoreq.20 bp for most of the markers, and therefore several markers can be analyzed simultaneously on the same gel lane (see FIGS. 2 and 4).
Example V
Demonstration of Mendelian codominant Inheritance of (dC-dA).sub.n.(dG-dT).sub.n Markers.
General Procedure
DNA from individuals of the CEPH families and from other three generation families was used as template for the amplification of various (dC-dA).sub.n.(dG-dT).sub.n markers using the procedure described in Example II.
Results
FIG. 4 shows the amplified DNA from markers Mfd1, Mfd3 and Mfd4 from CEPH family 1423. DNA fragment sized in bases are marked to the left of the gel. individual genotypes are listed below the gel. All three markers showed Mendelian behavior for this family.
A total of approximately 500 family/marker combinations have been tested to date. Mendelian codominant inheritance has been observed in all cases; no new mutations have been found. Therefore new mutations are unlikely to be a general problem with the (dC-dA).sub.n.(dG-dT).sub.n markers.
Example VI
Artifacts of the amplification reactions.
General Procedures
Aliquots of two different (dC-dA).sub.n.(dG-dT).sub.n amplified fragments (two different markers from two individuals), untreated from the polymerase chain reaction (C), were brought up to uM dATP and then incubated at 37.degree. for 30 min with 6 units of Klenow enzyme (K), 1 unit of T4 DNA polymerase (P), or with no additional enzyme (T). Samples were then mixed with formamide sample buffer and loaded on polyacrylamide gels.
For the results shown in FIG. 6, DNA from an individual with the Mfd3 genotype Z+6, Z-6 was amplified with modified Mfd3 primers (AGGCTGCAGGATTCACTGCTGTGGACCCA (SEQ. ID. NO:459) and GTCGGTACCGGTCTGGAAGTACTGAGAAA) (SEQ. ID. No:460) so that sites for the restriction enzymes Pst I and Kpn I were located at opposite ends of the amplified DNA. An aliquot of the amplified DNA from a 27 cycle reaction was diluted 60,000 fold with 0.2.times. TE, and 10 .mu.l of the dilution (approximately 10.sup.5 molecules) were amplified with the same primers for another 27 cycles. An aliquot of the second amplification reaction was diluted as above and subjected to a third 27 cycle reaction. Amplified DNA samples were treated with T4 DNA polymerase prior to electrophoresis.
Amplified DNA from the first reaction described in the above paragraph was digested with Kpn I and Pst I, extracted with phenol, and simultaneously concentrated and dialyzed into 0.2.times. TE using a Centricon 30 cartridge. This DNA was then ligated into mp19 and transformed in E. coli. Clear plaques on X-gal/IPTG plates from the transformed cells were picked, amplified and used to isolate single-stranded DNA. DNA from 102 such clones was sequenced. The distribution of the numbers of dinucleotide repeats in these clones is shown in Table 43.
Results.
Additional bands, less intense than the major pair of bands for each allele and smaller in size than the major bands, were usually seen for the amplified DNA fragments. These bands are particularly apparent in FIGS. 2-4. The additional bands were present when cloned DNA versus genomic DNA was used as template (FIG. 1, lane P), and even when such small amounts of heterozygote genomic DNA were used as template that only one of the two alleles was amplified. Also, DNA amplified from 63 lymphocyte clones (gift of J. Nicklas) (Nicklas et al (1987), Mutagenesis 2:341-347) produced from two individuals showed no variation in genotype for all clones from a single donor. These results indicate that the additional bands are generated as artifacts during the amplification reactions, and are not reflections of somatic mosaicism.
Griffin et al. (1988), (Am. J. Hum. Genet. 43(Suppl.):A185) demonstrated that DNA fragments amplified by the PCR could not be efficiently ligated to blunt ended vectors without first repairing the ends of the fragments with T4 DNA polymerase. To test whether "ragged" ends were responsible for the extra bands associated with (dC-dA).sub.n.(dG-dT).sub.n amplified fragments, amplified DNA (C) was treated with the Klenow fragment of DNA polymerase I (K), with T4 DNA polymerase (P) or with no enzyme (T) as shown in FIG. 5. Both Klenow enzyme and T4 DNA polymerase simplified the banding pattern somewhat by eliminating extra bands which differed in size from the most intense bands by 1 base. These enzymes also reduced the size of the most intense bands by 1 base. The most likely explanation for these results is the Taq polymerase produces a mixture of fragments during the PCR with different types of ends. The most intense bands in the untreated samples are likely derived from double stranded molecules with single base noncomplementary 3' overhangs. The fainter bands which are 1 base smaller than the major bands are likely to be blunt ended. The 3'-5' exonuclease activity of Klenow or T4 DNA polymerase converts the molecules with overhangs into blunt ended molecules. Clark (1988), (Nucleic Acids Res. 16:9677) showed that a variety of DNA polymerases including Taq polymerase can add a noncomplementary extra base to the 3' ends of blunt ended molecules.
Remaining after Klenow or T4 DNA polymerase treatment are extra bands which differ in size from the major bands by multiples of two bases. The data in Table 4 and FIG. 6 strongly indicate that these particular extra bands are the result of the skipping of repeats by the Taq polymerase during the amplification cycles. Sequencing of individual clones of DNA amplified for 27 cycles (first lane, FIG. 6) verifies that repeats have been deleted in the amplified DNA (Table 4). The largest of the predominant-sized fragments in the amplified DNA (with the exception of the one 20-mer) are 19 and 13 repeats in length, consistent with the Z+6, Z-6 genotype of the donor of the template DNA. Substantial numbers of clones containing fewer repeats are also seen, and this distribution matches the pattern of fragments shown in the first lane of FIG. 6. If the Taq polymerase is skipping repeats during the amplification cycles, as seems likely from the sequencing data, then further cycles of amplification in addition to the first 27 should reduce the intensities of the bands corresponding to the original DNA in the template and increase the intensities of the bands corresponding to fragments with skipped repeats. This is exactly what is observed in FIG. 6 for the amplified DNA from the reactions with 54 and 81 total amplification cycles.
Example VII
Use of (dC-dA).sub.n.(dG-dT).sub.n polymorphisms to identify individuals.
General Procedure
Genomic DNA from a collection of 18 unrelated individuals and from an unknown individual was isolated from blood and amplified with various (dC-dA).sub.n.(dG-dT).sub.n markers as described under Example II. Genotype frequencies were calculated from the allele frequencies shown in Table 43 assuming Hardy-Weinberg equilibria.
Results
One individual out of a group of 18 unrelated volunteers was selected so that the identity of this individual was unknown. The unknown DNA sample and the control samples were then typed for the Mfd3 and Mfd4 markers. As shown in Table 44, only three of the 18 controls had Mfd3 and Mfd4 genotypes consistent with the unknown sample, namely, individuals 22, 35 and 42. Further typing of these three samples and the unknown with four additional markers as shown in Table 44, conclusively demonstrated that the unknown DNA sample came from individual 22.
Table 45 shows the expected genotype frequencies for the six typed markers for individual 22 using the allele frequency data from Table 46. Single genotypes range in frequency from 0.04 to 0.49. The frequency for the entire collection of six markers is the product of the individual probabilities or 1.5.times.10.sup.-5 or about 1 in 65,000 people. Choosing a better, more informative collection of (dC-dA).sub.n.(dG-dT).sub.n markers would result in considerably greater discriminant ion.
EXAMPLE VIII
Analyzing the dependence of (dC-dA).sub.n.(dG-dT).sub.n marker informativeness of repeat sequence length and sequence type.
Repeat Sequences
(dC-dA).sub.n.(dG-dT).sub.n sequences were obtained either by computer search of GenBank, Version 54, for all sequences with six or more consecutive CA or GT repeats, or by sequencing m13 clones selected by stringent hybridization to poly (dC-dA).poly(dG-dT) (Weber and May, 1989, supra.). m13 libraries were produced by ligating size-selected human genomic Sau3A I, Alu I, or Sau3A I/Taq I fragments in the range of 150-400 bp into mp10, mp18 or mp19 vectors. Size-selected fragments were obtained by electrophoresis on preparative agarose gels and elution using GENECLEAN (Bio 101 Inc., LaJolla, Calif.).
Rules for categorizing (dC-dA).sub.n.(dG-dT).sub.n sequences, i.e. perfect sequence, imperfect sequence, compound repeat sequence; are described elsewhere in this application. Examples are listed in Table 46 as follows:
A total of 55 human (dC-dA).sub.n.(dG-dT).sub.n sequences were obtained by searching GenBank Version 54 for all sequences with six or more CA or GT repeats. 57 (dC-dA).sub.n.(dG-dT).sub.n sequences were determined in the lab after selecting random human genomic DNA clones by hybridization to poly(dC-dA).poly(dG-dT). These repeat sequences were classified as either perfect (no interruptions in the run of dinucleotide repeats), imperfect (one or more interruptions in the run of repeats), or compound (a run of perfect or imperfect CA or GT repeats adjacent to a run of another simple sequence repeat). Examples of repeat sequences in each of the categories are shown in Table 46 above; numbers of sequences in each category are listed in Table 47 as follows:
Perfect repeat sequences predominated and compound repeat sequences were relatively infrequent. The proportions of sequences of the various types were similar for the GenBank and clone sequences.
Length distributions of the GenBank and clone (dC-dA).sub.n.(dG-dT).sub.n sequences are shown in FIG. 7. The two distributions were similar except that a much greater proportion of the GenBank (dC-dA).sub.n.(dG-dT).sub.n sequences contained 6-9 repeats. Few sequences were found with more than 24 repeats.
Example IX
PIC Determination.
Conditions for the amplification reactions were essentially as described in Weber and May, 1989, supra. except that a recombinant Taq polymerase (AmpliTaq, Perkin Elmer Cetus, Norwalk, Conn.) was used and samples were processed through 27 temperature cycles consisting of 1 min at 94.degree., 2 min at 55.degree. and 2 min at 72.degree.. Denaturing polyacrylamide DNA sequencing gels contained 6.5% acrylamide and were 0.4 mm thick. Size standards were mp8 dideoxynucleotide sequencing ladders obtained using a -20 primer (Catalog #1211, New England Biolabs, Beverly, Mass.).
Allele sizes were measured as the midpoint between the CA-strand and GT-strand bands (Weber and May 1989). The error in this procedure is estimated at 1-2 bases. Relative sizes of the different alleles for a single marker, however, were determined without appreciable error by loading a common set of amplified DNA fragments on all gels used to determine PIC values for that marker.
For all polymorphisms, the size of one of the allelic fragments either matched exactly or differed by only one base from the size determined by knowledge of the original genomic sequences used to synthesize the PCR primers. In cases where the allele sizes differed from the expected fragment size by a single base, the most frequent of the two alleles nearest in size to the expected allele was taken to represent the sequenced allele. The size-matched allele was used to set the number of repeats for each of the other alleles simply by assuming that every 2 bp difference in allele size corresponded to a difference of one repeat unit. For amplified fragments which did not show any size polymorphism, the identity of the fragment of the expected size was verified by digesting the amplified DNA with a restriction enzyme which was predicted to cut within the interior of the fragment.
Allele sizes from 82-146 chromosomes were measured for estimations of allele frequencies and PIC values. For four of the markers, DNA samples were obtained from unrelated Caucasian volunteers from the Marshfield, Wis. area; for another six markers, DNA samples were from CEPH family parents (Caucasians from France, Venezuela, and the US); and for the remaining markers, DNA samples were from selected CEPH family grandparents.
Statistical Analysis
The weighted average number of repeats for a marker was defined as ##EQU1## where p is the total number of alleles for that marker, f.sub.i was the estimate of allele frequency for the ith allele and n.sub.i was the number of repeats for the ith allele. The relationship between PIC and the weighted average number of repeats (for the perfect repeat sequence markers) was approximated by applying SPSS/PC+ software (SPSS Inc., Chicago, Ill.) to fit a logistic type curve to the data using an iterative least squares method. Other types of curves were not judged to fit the data as well as the logistic curve. The equation of the least squares curve shown in FIGS. 2 and 3 is ##EQU2## where w is the weighted average number of repeats, a=4.388, b=0.297 and c=0.85. To compare the fit of the logistic curve to the two sets of data shown in FIG. 2, residuals were calculated for each data set, and differences between residuals for the two sets were analyzed using a paired t-test.
A plot of PIC values versus numbers of repeats for markers with perfect repeat sequences is shown in the FIG. 8A. Information describing the least informative markers, including PCR primer sequences, is listed in Table 48 as follows:
The weighted average number of repeats for each marker was chosen as the most appropriate measure of repeat length. Only sequences with 12 or fewer repeats were found to be nonpolymorphic (PIC=0). Informativeness of the markers generally increased as the average number of repeats increased, especially in the range of about 11-17 repeats. Maximum informativeness of the markers tested was 0.81 PIC.
Usually, at the time of PCR primer synthesis, the repeat sequence of only one allele is available. To determine if the numbers of repeats in the original sequences were good predictors of marker informativeness, these repeat lengths were plotted versus PIC values as shown in FIG. 8B. The fit of this data to the curve (the identical curve was shown in both plots) was not significantly different from the weighted average points shown at the top (p=0.57).
Informativeness for markers with imperfect repeat sequences was plotted versus repeat length in FIG. 9. The open squares were obtained by counting all the repeats, including imperfect repeats, within the repeat sequence. These points were all below the least squares curve (from FIGS. 8A and 8B) with especially disappointing PIC values for some sequences with the greatest numbers of repeats. The probability that all nine data points lie below the curve on the basis of chance alone is very small (0.002). The filled squares were obtained using the longest run of perfect repeats as a measure of repeat length. When plotted in this fashion, the data for the imperfect repeat sequence markers exhibited a much better fit to the least squares curve.
Various modes of carrying out the invention are contemplated as being within the scope of the following claims particularly pointing out and distinctly claiming the subject matter which is regarded as the invention.
Claims
I claim:
1. An isolated nucleic acid molecule consisting of a sequence selected from the group consisting of SEQ. ID. NOS: 1 to 26, 28 to 53, 56, 59, 62, 65, 68, 71, 74, 77, 80, 98, 113, 137, 143, 163, 186, 215, 249, 301, 304, 307, 312, 317, 322, 325, 328, 331, 334, 337, 340, 343, 346, 352, 355, 358, 361, 364, 370, 373, 376, 379, 382, 388, 391, and 394.
2. A method for detecting a polymorphic genetic marker of the form (dC dA).sub.n.(dG dT).sub.n, wherein n.gtoreq.6, comprising: a. isolating genomic DNA from a subject; b. analyzing the isolated genomic DNA for the presence of said polymorphic genetic marker using a nucleic acid molecule consisting of a sequence selected from the group consisting of SEQ. ID. NOS: 1 to 26, 28 to 53, 56, 59, 62, 65, 68, 71, 74, 77, 80, 98, 113, 137, 143, 163, 186, 215, 249, 301, 304, 307, 312, 317, 322, 325, 328, 331, 334, 337, 340, 343, 346, 352, 355, 358, 361, 364, 370, 373, 376, 379, 382, 388, 391, and 394 to detect said polymorphic genetic marker.
3. The method of claim 2 wherein step (b) further comprises: a. amplifying DNA molecules from the genomic DNA by the polymerase chain reaction; b. resolving the amplified DNA molecules by electrophoresis; c. detecting the amplified DNA molecules; and d. determining the allele of the polymorphic marker present in the genomic DNA.
4. The method of claim 2 wherein the sequence of the nucleic acid molecule used in step (b) is a perfect repeat sequence.
5. The method of claim 2 wherein the sequence of the nucleic acid molecule used in step (b) is an imperfect repeat sequence.
6. The method of claim 2 wherein the sequence of the nucleic acid molecule used in step (b) is a compound repeat sequence.
7. A kit for performing the method of claim 2 comprising: a. oligonucleotide primer pairs consisting of sequence pairs selected from the group consisting of: SEQ. ID. NOS: 54 and 55, SEQ. ID. NOS: 57 and 58, SEQ. ID. NOS: 60 and 61, SEQ. ID. NOS: 63 and 64, SEQ. ID. NOS: 66 and 67, SEQ. ID. NOS: 69 and 70, SEQ. ID. NOS: 72 and 73, SEQ. ID. NOS: 75 and 76, SEQ. ID. NOS: 78 and 79, SEQ. ID. NOS: 81 and 82, SEQ. ID. NOS: 99 and 100, SEQ. ID. NOS: 114 and 115, SEQ. ID. NOS: 138 and 139, SEQ. ID. NOS: 144 and 145, SEQ. ID. NOS: 164 and 165, SEQ. ID. NOS: 187 and 188, SEQ. ID. NOS: 211 and 212, SEQ. ID. NOS: 216 and 217, SEQ. ID. NOS: 250 and 251, SEQ. ID. NOS: 302 and 303, SEQ. ID. NOS: 305 and 306, SEQ. ID. NOS: 308 and 309, SEQ. ID. NOS: 313 and 314, SEQ. ID. NOS: 318 and 319, SEQ. ID. NOS: 323 and 324, SEQ. ID. NOS: 326 and 327, SEQ. ID. NOS: 329 and 330, SEQ. ID. NOS: 332 and 333, SEQ. ID. NOS: 335 and 336, SEQ. ID. NOS: 338 and 339, SEQ. ID. NOS: 341 and 342, SEQ. ID. NOS: 344 and 345, SEQ. ID. NOS: 347 and 348, SEQ. ID. NOS: 350 and 351, SEQ. ID. NOS: 353 and 354, SEQ. ID. NOS: 356 and 357, SEQ. ID. NOS: 359 and 360, SEQ. ID. NOS: 362 and 363, SEQ. ID. NOS: 365 and 366, SEQ. ID. NOS: 371 and 372, SEQ. ID. NOS: 374 and 375, SEQ. ID. NOS: 377 and 378, SEQ. ID. NOS: 380 and 381, SEQ. ID. NOS: 383 and 384, SEQ. ID. NOS: 389 and 390, SEQ. ID. NOS: 392 and 393, and SEQ. ID. NOS: 395 and 396; and b. reagents necessary for the amplification of DNA sequences using the polymerase chain reaction.
8. The kit of claim 7 wherein the sequence of the amplified DNA is a perfect repeat sequence.
9. The kit of claim 7 wherein the sequence of the amplified DNA is an imperfect repeat sequence.
10. The kit of claim 7 wherein the sequence of the amplified DNA is a compound repeat sequence.
Patent Citations (3)
Non-Patent Literature (43)
- Nathans et al "Isolation and Nucleotide Sequence of the Gene Encoding . . . , " PNAS 81: 4851-5, Aug. 1984.
- Litt and Luty "A Hypervariable Microsatellite Revealed by in vitro Amplification . . . , " Mar. 1989 Am. J. Hum. Genet. 44:397-401.
- Litt et al "A Highly Polymorphic (TG)n Microsatellite at the D11S35 Locus" Mar. 29, 1989 J Cell Biochem Supp. 13D, p. 32 abstract #K219.
- Hamada et al "Characterization of Genomic Polyd2T-dG).multidot.Poly (dC-dA) Sequences . . . ," 1984. Mol. Cell. Biol. 4:2610-21.
- Aldridge et al., 1984, "A Strategy to Reveal High-Frequency RFLPs Along the Human X Chromosome." Am. J. Hum. Genet. 36:546-564.
- Benton and Davis, 1977, "Screening .lambda.gt Recombinant Clones by Hybridization to Single Plaques In Situ." Science 196:180-182.
- Biggin et al., 1983, "Buffer Gradient Gels and .sup.35 S Label as an Aid to Rapid DNA Sequence Determination." Proc. Natl. Acad. Sci. USA 80:3963-3965.
- Bostein et al., 1980, "Construction of a Genetic Linkage Map in Man Using Restriction Fragment Length Polymorphisms." Am. J. Hum. Genet. 32:314-331.
- Clark, 1988, "Novel Non-Templated Nucleotide Addition Reactions Catalyzed by Procaryotic and Eucaryotic DNA Polymerases." Nucleic Acids Res. 16:9677-9686.
- Das et al., 1987, "The Human Apolipoprotein C-II Gene Sequence Contains a Novel Chromosome 19-Specific Minisatellite in Its Third Intron." J. Biol. Chem. 262:4787-4793.
- Donis-Keller et al., 1987, "A Genetic Linkage Map of the Human Genome." Cell 51:319-337.
- Elbrecht, A., 1987, "Lab Hints: Irish Cream Liqueur as a Blocking Agent for DNA Dot Blots." B.M. Biochemica, 4:12-13.
- Floyd-Smith, G., Whitehead, A. S., Colten, H. R., and Francke, U., 1986, "The Human C-Reactive Protein Gene (CRP) and Serum Amyloid P Component Gene (APCS) are Located on the Proximal Long Arm of Chromosome I." Immunogenetics 24:171-176.
- Griffin et al., 1988, "Synthesis of Hexokinase 1 (HK1) cDNA Probes by Mixed Oligonucleotide Primed Amplification of cDNA (MOPAC) Using Primer Mixtures of High Complexity". Am. J. Hum. Genet. 43 (Suppl.):A185.
- Gross and Garrard, 1986, "The Ubiquitous Potential Z-Forming Sequence of Eucaryotes, (dT-dG).sub.n .cndot.(dC-dA).sub.n, is not Detectable in the Genomes of Eubacteria, Archaebacteria, or Mitochondria." Mol. Cell. Biol. 6:3010-3013.
- Gusella, J. F., 1986, "DNA Polymorphism and Human Disease," Ann. Rev. Biochem. 55:831-854.
- Hamada, H., Petrino, M. G. and Kakunaga, T., 1982, "A Novel Repeated Element with Z-DNA-Forming Potential is Widely Found in Evolutionarily Diverse Eukaryotic Genomes," Proc. Nat. Acad. Sci. U.S.A. 79:6465-6469.
- Hamada and Kakunaga, 1982, "Potential Z-DNA Forming Sequences are Highly Dispersed in the Human Genome." Nature 298:396-398.
- Jeffreys et al., 1985, "Hypervariable `Minisatellite` Regions in Human DNA." Nature 314:67-73.
- Jeffreys et al., 1988, "Spontaneous Mutation Rates to New Length Alleles at Tandem-Repetitive Hypervariable Loci in Human DNA." Nature 332:278-281.
- Ledbetter, D. H., Rich, D. C., O'Connell, P., Leppert, M., and Carey, J. C., 1989, "Precise Localization of NFI to 17q11.2 by Balanced Translocation." Am. J. Hum. Genet. 44:20-24.
- Ledbetter, S. A., Schwartz, C. E., Davies, K. E. and Ledbetter D. H., 1991, "New Somatic Cell Hybrids for Physical Mapping in Distal Xq and the Fragine X Region." Am. J. Med. Genet. 38:418-420.
- Litt, M., Buroker, N. E., Kondoleon, S., Douglass, J., Liston, D., Sheehy, R., and Magenis, R. E., 1988, "Chromosomal Localization of the Human Proenkephalin and Prodynorphin Genes." Am. J. Hum. Genet. 42:327-334.
- Luty, J. A., Guo, Z., Willard, H. F., Ledbetter, D. H., Ledbetter, S. and Litt, M., 1990, "Five Polymorphic Microsatellite VNTRs on the Human X Chromosome." Am. J. Hum. Genet. 46:776-783.
- Matthews and Kricka, 1988, "Analytical Strategies for the Use of DNA Probes." Anal. Biochem. 169:1-25.
- Miesfeld et al., 1981, "A Member of a New Repeated Sequence Family Which is Conserved Throughout Eucaryotic Evolution is Found Between the Human .delta. and .beta. Globin Genes." Nucleic Acids Res. 9:5931-5947.
- Mullis and Faloona, 1987, "Specific Synthesis of DNA in Vitro via a Polymerase-Catalyzed Chain Reaction." Method Enzymol. 155:335-350.
- Nakamura et al., 1987, "Variable Number of Tandem Repeat VNTR Markers for Human Gene Mapping." Science 235:1616-1622.
- Nicklas et al., 1987, "Molecular Analyses of in Vivo hprt Mutations in Human T-Lymphocytes: I. Studies of How Frequency `Spontaneous` Mutants by Southern Blots." Mutagenesis 2:341-347.
- Overhauser et al., 1987, "Identification of 28 DNA Fragments that Detect RFLP's in 13 Distinct Physical Regions of the Short Arm of Chromosome 5." Nucleic Acids Res. 15:4617-4627.
- Saiki et al., 1985, "Enzymatic Amplification of .beta.-Globin Genomic Sequences and Restriction Site Analysis for Diagnosis of Sickle Cell Anemia." Science 230:1350-1354.
- Saiki et al., 1988, "Primer-Directed Enzymatic Amplification of DNA with a Thermostable DNA Polymerase." Science 239:487-491.
- Schumm et al., 1988, "Identification of More than 500 RFLPs by Screening Random Genomic Clones." Am. J. Hum. Genet. 42:143-159.
- Shen and Rutter, 1984, "Sequence of the Human Somatostatin I Gene," Science 224:168-171.
- Slightom et al., 1980, "Human Fetal .sup.G .gamma. and .sup.A .gamma. Globin Genes: Complete Nucleotide Sequences Suggest that DNA can be Exchanged Between These Duplicated Genes." Cell 21:627-638.
- Tautz and Renz, 1984, "Simple Sequences are Ubiquitous Repetitive Components of Eukaryotic Genomes." Nucleic Acids Res. 12:4127-4138.
- vanTuinen, P., Rich, D. C. Summers, K. M., and Ledbetter, D. H., 1987, "Regional Mapping Panel for Human Chromosome 17: Application to Neurofibromatosis Type 1." Genomics 1:374-381.
- vanTuinen, P., Dobyns, W. B., Rich, D. C., Summers, K. M., Robinson, T. J., Nakamura, Y., and Ledbetter, D. H., 1988, "Molecular Detection of Microscopic and Submicroscopic Deletions Associated with Miller-Dieker Syndrome." Am. J. Hum. Genet. 43:587-596.
- Weber and May, 1988, "An Abundant New Class of Human DNA Polymorphisms." Am. J. Hum. Genet. 43:A161 Abstract.
- Weber et al., 1988, "Primary Structure of a Plasmodium falciparum Malaria Antigen Located at the Merozoite Surface and Within the Parasitophorous Vacuole." J. Biol. Chem. 263:11421-11425.
- Weber, J. L. and May, P. E., 1989, "Abundant Class of Human DNA Polymorphisms Which Can Be Typed Using the Polymerase Chain Reaction." Am. J. Hum. Genet. 44:388-396.
- White and Lalouel, 1988, "Sets of Linked Genetic Markers for Human Chromosomes." Annu. Rev. Genet., 22:259-279.
- Zoghbi, H. Y., McCall, A. E., LeBorgne-Demarquoy, F., 1990, "The Use of Radiation Hybrids and Interspersed Repetitive Sequence (IRS) PCR for the Rapid Isolation of 23 DNA Fragments Near the Spinocerebellar Ataxia (SCA1) Locus." Am. J. Hum. Genet. 47:A206.