US 6,048,689 AGrant
Method for Identifying Variations in Polynucleotide Sequences
Issue Date:2000-04-11
•25 Claims
•10 Drawing Sheets
Abstract
A step-wise integrated process for identifying sequence variations in polynucleotide sequences is disclosed. The identification process is composed of three stages, including allele specific hybridization assays of known sequence variations (Stage I), sequence variation locating assays (Stage II), and direct sequencing (Stage III). The methods can be used for efficient and accurate detection of mutations in any test gene sample.
Metadata
Assignee
- Gene Logic, Inc.
Inventors
- Patricia D. Murphy
- Marga B. White
Application Information
Application Number:US 8254877
Filing Date:1997-03-28
Priority Date:1997-03-28
Art Unit:163
Classifications
IPC:
C12Q 168
Field of Search:
4355369356;91.224.33;24.377;78
Patent Drawings (10 sheets)
Description
1. Introduction
The present invention relates to methods of detecting and identifying sequence variations in polynucleotide sequences. More specifically, this invention relates to a screening process whereby the presence of sequence variations is detected in sequential steps. The methods are applicable to detecting mutations in any isolated gene. The invention is described in detail, for the purpose of illustration and not by way of limitation, for detecting mutations in the human BRCA1 gene.
2. Background of the Invention
An increasing number of genes which play a role in many different diseases are being identified. Detection of mutations in such genes is instrumental in determining susceptibility to or diagnosing these diseases. Some diseases, such as sickle cell disease, are monomorphic, i.e., the disease is generally caused by a single mutation present in the population. In such cases where one or only a few known mutations are responsible for the disease, methods for detecting the mutations are targeted to the site within the gene at which they are known to occur.
In many other cases, however, individuals affected by a given disease display extensive allelic heterogeneity. For example, more than 125 mutations in the human BRCA1 gene have been reported (Breast Cancer Information Core world wide web site at http://www.nchgr.nih.gov/dir/lab.sub.-- transfer/bic, which became publicly available on Nov. 1, 1995; Friend, S. et al., 1995, Nature Genetics 11: 238). Mutations in the BRCA1 gene are thought to account for roughly 45% of inherited breast cancer and 80-90% of families with increased risk of early onset breast and ovarian cancer (Easton, 1993, et al., American Journal of Human Genetics 52: 678-701).
Other examples of genes for which the population displays extensive allelic heterogeneity and which have been implicated in disease include CFTR (cystic fibrosis), dystrophin (Duchenne muscular dystrophy, and Becker muscular dystrophy), and p53 (Li-Fraumeni syndrome).
Breast cancer is also an example of a disease in which, in addition to allelic heterogeneity, there is genetic heterogeneity. In addition to BRCA1, the BRCA2 and BRCA3 genes have been linked to breast cancer. Similarly, the NFI and NFII genes are involved in neurofibromatosis (types I and II, respectively).
Accuracy in detection of mutations is extremely important, particularly in clinical settings. Direct end-to-end sequencing of the gene potentially provides the most accurate results, given that an accurate reference sequence is available. However, sequencing can also be a cumbersome technique. Detection of one of many known or unknown mutations is further complicated when the gene is large and/or has a complex structure. The human BRCA1 gene, for example, is approximately 100,000 base pairs long and contains 24 exons (Weber, B., Science and Medicine, Scientific American January-February 1996, 12-21). Furthermore, in order to be practical and available to the general population, detection methods must be efficient enough to accommodate a large number of different samples.
A number of techniques that are more rapid but less comprehensive than direct sequencing have been developed for detecting nucleotide sequence variations. Many of these techniques are based on detecting differences, between normal and mutant nucleotide sequences, in hybridization (e.g., allele specific hybridization), secondary structure (e.g., single strand conformation polymorphism analysis, heteroduplex analysis), melting (constant denaturing gel electrophoresis, denaturing gradient gel electrophoresis), and susceptibility to cleavage (either chemical or restriction enzyme cleavage). Other techniques, such as the protein truncation test, detect changes on the protein level. For a summary of such techniques, see Marajver & Petty, 1996, Clinics in Lab. Med. 16: 139-167, especially Table 5 at p. 152.
Efforts have independently focused primarily on increasing the rapidity of processing sequencing analyses or increasing the comprehensiveness of the hybridization-based techniques. There remains a need, however, for a systematic method of detecting mutations in individual gene samples that is both accurate enough to provide a reliable diagnosis to an individual patient and efficient enough to be practical for application to the general population.
3. Summary of the Invention
It is an object of the invention to provide a step-wise integrated process for efficient and accurate detection of variations in polynucleotide sequences. For the purpose of convenience, the steps are described herein according to the following categories of analysis: Stage I, Stage II, and Stage III. Stage I involves allele specific hybridization assays for detection of known mutations in the test gene sample. Stage II analysis involves sequence variation locating assays, with subsequent targeted confirmatory sequencing of detected sequence variations. Stage III analysis involves direct sequencing analysis of any regions of the polynucleotide not sequenced in Stage II.
Stages I, II, and III are components of an overall sequence variation detection process. In one embodiment of the invention, Stage I analysis is followed by Stage II analysis, then by Stage III analysis. In another embodiment of the invention, Stage I analysis is followed by Stage II analysis without proceeding on to Stage III. In yet another embodiment of the invention, Stage I analysis is followed by Stage III analysis. In still another embodiment of the invention, the process is initiated with Stage II analysis, followed by Stage III analysis.
4. Brief Description of the Figures
FIGS. 1A-1J. Nucleotide sequence and encoded amino acid sequence of the BRCA1.sup.(omi1) gene.
5. Detailed Description of the Invention
Described below is a step-wise integrated process for efficient and accurate detection of variations in polynucleotide sequences. For the purpose of convenience, the steps are described herein according to the following categories of analysis: Stage I, Stage II, and Stage III.
The present invention can be used to identify sequence variations in any polynucleotide sequence. For example, the present invention can be used to detect mutations in any gene that plays a role in disease. Examples of genes which can be analyzed for mutations and other sequence variations in accordance with the invention include, but are not limited to, BRCA1, BRCA2, BRCA3, cystic fibrosis transmembrane regulator (CFTR), dystrophin, neurofibromatosis I (NFI), neurofibromatosis II (NFII), and p53. Reference sequences for known genes can be found, for example, by searching the world wide web site of GenBank at http://www.ncbi.nlm.nih.gov (for example, BRCA1 at accession #Y08757 (or FIG. 1, herein); BRCA2 at accession #U43746; dystrophin at locus HUMDYS; CFTR at accession #M28668; NF1 at accession #M89914, NF2 at accession #L27131; and p53 at accession #X54156).
The following definitions are provided for the purpose of understanding this invention.
"Coding sequence" or "DNA coding sequence" refers to those portions of a gene which, taken together, code for a peptide (protein), or which nucleic acid itself has function.
"BRCA1.sup.(omi) " refers collectively to the "BRCA1.sup.(omi1) ", "BRCA1.sup.(omi2) " and "BRCA1.sup.(omi3) " coding sequences.
"BRCA1.sup.(omi1) " refers to the most frequently occurring coding sequence for the BRCA1 gene. This coding sequence was found by end-to-end sequencing of BRCA1 alleles from individuals randomly drawn from a Caucasian population found to have no family history of breast or ovarian cancer. The sequenced gene was found not to contain any mutations. BRCA1.sup.(omi1) was determined to be a consensus sequence by calculating the frequency with which the coding sequence occurred among the sample alleles sequenced. BRCA1.sup.(omi1) is shown in FIG. 1.
"BRCA1.sup.(omi2) " and "BRCA1.sup.(omi3) " refer to two additional, less frequently occurring coding sequences for the BRCA1 gene which were also isolated from individuals randomly drawn from a Caucasian population found to have no family history of breast or ovarian cancer. The differences among BRCA1.sup.(omi1), BRCA1.sup.(omi2) and BRCA1.sup.(omi3) are summarized in Table 1, below.
"Primer" as used herein refers to a sequence comprising about 20 or more nucleotides of a gene.
A "target polynucleotide" refers to the nucleic acid sequence of interest e.g., the BRCA1 encoding polynucleotide.
"Consensus" means the most commonly occurring in the population.
"Substantially complementary to" refers to a probe or primer sequences which hybridize to the sequences provided under stringent conditions and/or sequences having sufficient homology with test polynucleotide sequences, such that the allele specific oligonucleotide probe or primers hybridize to the test polynucleotide sequences to which they are complimentary.
"Isolated" as used herein refers to substantially free of other nucleic acids, proteins, lipids, carbohydrates or other materials with which they may be associated. Such association is typically either in cellular material or in a synthesis medium.
"Sequence variation" as used herein refers to any difference in nucleotide sequence between two different oligonucleotide or polynucleotide sequences.
"Polymorphism" as used herein refers to a sequence variation in a gene which is not associated with known pathology.
"Mutation" as used herein refers to a sequence variation in a gene which results in the gene coding for a non-functioning protein or a protein with substantially reduced or altered function.
"Pre-determined sequence variation" as used herein refers to a nucleotide sequence that is designed to be different than the corresponding sequence in a reference nucleotide sequence. A pre-determined sequence variation can be a known mutation in a gene.
"Allele specific hybridization assay" as used herein refers to an assay to detect the presence or absence of a pre-determined sequence variation in a test polynucleotide or oligonucleotide by hybridizing the test polynucleotide or oligonucleotide with a polynucleotide or oligonucleotide of pre-determined sequence such that differential hybridization between matching nucleotide sequences versus nucleotide sequences containing a mismatch is detected.
"Sequence variation locating assay" as used herein refers to an assay that detects a sequence variation in a test polynucleotide or oligonucleotide and localizes the position of the sequence variation to a sub-region of the test polynucleotide, without necessarily determining the precise base change or position of the sequence variation.
"Targeted confirmatory sequencing" as used herein refers to sequencing a polynucleotide in the region wherein a sequence variation has been located by a sequence variation locating assay in order to determine the precise base change and/or position of the sequence variation.
The aforementioned three categories of analysis, Stage I, Stage II, and Stage III, are components of an overall sequence variation detection process. Stage I involves allele specific hybridization assays for detection of known mutations in the test gene sample. Stage II analysis involves sequence variation locating assays, with subsequent targeted confirmatory sequencing of detected sequence variations. Stage III analysis involves direct sequencing analysis of any regions of the polynucleotide not sequenced in Stage II. Stage I, II and III may be combined differently to identify sequence variations in any polynucleotide sequence. Stage I analysis may be followed by Stage II analysis, then by Stage III analysis. Stage I analysis may be followed by Stage II analysis without proceeding on to Stage III. Stage I analysis may be followed by Stage III analysis. The analysis may also be initiated with Stage II analysis, followed by Stage III analysis.
One embodiment of the present invention is a method for determining the presence or absence of a sequence variation in a gene sample, comprising:
(a) performing an allele specific hybridization assay for the presence or absence of one or more pre-determined sequence variations;
(b) if no pre-determined sequence variation is found in step (a), then performing a sequence variation locating assay;
(c) if no sequence variation is found in step (b), then sequencing the gene sample; and
(d) determining the presence or absence of a sequence variation by analyzing the sequences obtained in step (c) against a reference sequence.
The method may further comprise repeating the allele specific hybridization until a predetermined number of known sequence variations have been tested for. The allele specific hybridization may also comprise testing for a predetermined number of sequence variations in a single step not requiring repetition. The predetermined sequence variation in step (a) may be a known mutation. The sequence variation may also be a known mutation. The allele specific hybridization assay may be performed using a dot blot format, a multiplex format, a reverse dot blot format, a MASDA format, or a chip array format. The sequence variation locating assay may be performed using a protein truncation assay, a chemical cleavage assay, a heteroduplex analysis, a single strand conformation polymorphism assay, a constant denaturing gel, electrophoresis assay, or a denaturing gradient gel electrophoresis assay. The sequencing of a gene sample may be performed in only the forward or reverse direction, or in both the forward and reverse directions. Both exons and introns of the gene or parts thereof may be sequenced. All exons and introns may also be sequenced from end to end. Alternatively, only exons or only intronic sequences may be sequenced. The reference sequence may be a coding sequence, a genomic sequence, or one or more exons of the gene. The present method may be used to identify sequence variations in any polynucleotide sequence. For example, the gene of interest may be a human BRCA1 gene. Accordingly, the reference sequence may be a BRCA1 coding sequence or a BRCA1 genomic sequence.
A further embodiment of the invention is a method for determining the presence or absence of a sequence variation in a gene sample, comprising:
(a) performing an allele specific hybridization assay for the presence of one or more pre-determined sequence variations;
(b) if no pre-determined sequence variation is found in step (a), then performing a sequence variation locating assay;
(c) if a sequence variation is detected in step (b), then performing targeted confirmatory sequencing; and
(d) determining the presence or absence of a sequence variation by analyzing the sequences obtained in step (c) against a reference sequence.
The method may further comprise repeating the allele specific hybridization until a predetermined number of known sequence variations have been tested for. The allele specific hybridization may also comprise testing for a predetermined number of sequence variations in a single step not requiring repetition. The predetermined sequence variation in step (a) may be a known mutation. The sequence variation may also be a known mutation. The allele specific hybridization assay may be performed using a dot blot format, a multiplex format, a reverse dot blot format, a MASDA format, or a chip array format. The sequence variation locating assay may be performed using a protein truncation assay, a chemical cleavage assay, a heteroduplex analysis, a single strand conformation polymorphism assay, a constant denaturing gel, electrophoresis assay, or a denaturing gradient gel electrophoresis assay. The sequencing of a gene sample may be performed in only the forward or reverse direction, or in both the forward and reverse directions. Both exons and introns of the gene or parts thereof may be sequenced. All exons and introns may also be sequenced from end to end. Alternatively, only exons or only intronic sequences may be sequenced. The reference sequence may be a coding sequence, a genomic sequence, or one or more exons of the gene. The present method may be used to identify sequence variations in any polynucleotide sequence. For example, the gene of interest may be a human BRCA1 gene. Accordingly, the reference sequence may be a BRCA1 coding sequence or a BRCA1 genomic sequence.
A further embodiment of the invention is a method for determining the presence or absence of a sequence variation in a gene sample, comprising:
(a) performing an allele specific hybridization assay for the presence or absence of one or more pre-determined sequence variations; and
(b) if no pre-determined sequence variation is found in step (a), then sequencing the gene sample; and
(c) determining the presence or absence of a sequence variation by analyzing the sequences obtained in step (b) against a reference sequence.
The method may further comprise repeating the allele specific hybridization until a predetermined number of known sequence variations have been tested for. The allele specific hybridization may also comprise testing for a predetermined number of sequence variations in a single step not requiring repetition. The predetermined sequence variation in step (a) may be a known mutation. The sequence variation may also be a known mutation. The allele specific hybridization assay may be performed using a dot blot format, a multiplex format, a reverse dot blot format, a MASDA format, or a chip array format. The sequencing of a gene sample may be performed in only the forward or reverse direction, or in both the forward and reverse directions. Both exons and introns of the gene or parts thereof may be sequenced. All exons and introns may also be sequenced from end to end. Alternatively, only exons or only intronic sequences may be sequenced. The reference sequence may be a coding sequence, a genomic sequence, or one or more exons of the gene. The present method may be used to identify sequence variations in any polynucleotide sequence. For example, the gene of interest may be a human BRCA1 gene. Accordingly, the reference sequence may be a BRCA1 coding sequence or a BRCA1 genomic sequence.
A further embodiment of the invention is a method for determining the presence or absence of a sequence variation in a gene sample, comprising:
(a) performing a sequence variation locating assay;
(b) if no sequence variation is found in step (a), then sequencing the gene sample; and
(c) determining the presence or absence of a sequence variation by analyzing the sequences obtained in step (b) against a reference sequence.
The sequence variation may be a known mutation. The sequence variation locating assay may be performed using a protein truncation assay, a chemical cleavage assay, a heteroduplex analysis, a single strand conformation polymorphism assay, a constant denaturing gel, electrophoresis assay, or a denaturing gradient gel clectrophoresis assay. The sequencing of a gene sample may be performed in only the forward or reverse direction, or in both the forward and reverse directions. Both exons and introns of the gene or parts thereof may be sequenced. All exons and all introns may also be sequenced from end to end. Alternatively, only exons or only intronic sequences may be sequenced. The reference sequence may be a coding sequence, a genomic sequence, or one or more exons of the gene. The present method may be used to identify sequence variations in any polynucleotide sequence. For example, the gene of interest may be a human BRCA1 gene. Accordingly, the reference sequence may be a BRCA1 coding sequence or a BRCA1 genomic sequence.
5.1. STAGE I: DETECTION OF PRE-DETERMINED SEQUENCE VARIATIONS
Stage I analysis is used to determine the presence or absence of a pre-determined nucleotide sequence variation; preferably a known mutation or set of known mutations in the test gene. In accordance with the invention, such pre-determined sequence variations are detected by allele specific hybridization. An allele specific hybridization assay detects the differential ability of mismatched nucleotide sequences (e.g., normal:mutant) to hybridize with each other, as compared with matching (e.g., normal:normal or mutant:mutant) sequences.
5.1.1. DETECTION OF PRE-DETERMINED SEQUENCE VARIATIONS USING ALLELE SPECIFIC HYBRIDIZATION
A variety of methods well-known in the art can be used for detection of pre-determined sequence variations by allele specific hybridization. Preferably, the test gene is probed with allele specific oligonucleotides (ASOs); and each ASO contains the sequence of a known mutation. ASO analysis detects specific sequence variations in a target polynucleotide fragment by testing the ability of a specific oligonucleotide probe to hybridize to the target polynucleotide fragment. Preferably, the oligonucleotide contains the mutant sequence (or its complement). The presence of a sequence variation in the target sequence is indicated by hybridization between the oligonucleotide probe and the target fragment under conditions in which an oligonucleotide probe containing a normal sequence does not hybridize to the target fragment. A lack of hybridization between the sequence variant (e.g., mutant) oligonucleotide probe and the target polynucleotide fragment indicates the absence of the specific sequence variation (e.g., mutation) in the target fragment. In a preferred embodiment, the test samples are probed in a standard dot blot format. Each region within the test gene that contains the sequence corresponding to the ASO is individually applied to a solid surface, for example, as an individual dot on a membrane. Each individual region can be produced, for example, as a separate PCR amplification product using methods well-known in the art (see, for example, the experimental embodiment set forth in Mullis, K. B., 1987, U.S. Pat. No. 4,683,202). The use of such a dot blot format is described in detail in the example in section 6.1, below, detailing the Stage I analysis of the human BRCA1 gene to detect the presence or absence of eight different known mutations using eight corresponding ASOs.
Membrane-based formats that can be used as alternatives to the dot blot format for performing ASO analysis include, but are not limited to, reverse dot blot, multiplex format, and multiplex allele-specific diagnostic assay (MASDA).
In a reverse dot blot format, oligonucleotide or polynucleotide probes having known sequence are immobilized on the solid surface, and are subsequently hybridized with the labeled test polynucleotide sample.
In a multiplex format, individual samples contain multiple target sequences within the test gene, instead of just a single target sequence. For example, multiple PCR products each containing at least one of the ASO target sequences are applied within the same sample dot. Multiple PCR products can be produced simultaneously in a single amplification reaction using the methods of Caskey et al., U.S. Pat. No. 5,582,989. The same blot, therefore, can be probed by each ASO whose corresponding sequence is represented in the sample dots.
A MASDA format expands the level of complexity of the multiplex format by using multiple ASOs to probe each blot (containing dots with multiple target sequences). This procedure is described in detail in U.S. Pat. No. 5,589,330 by A. P. Shuber, and in Michalowsky et al., October 1996, American Journal of Human Genetics 59(4): A272, poster 1573, each of which is incorporated herein by reference in its entirety. First, hybridization between the multiple ASO probe and immobilized sample is detected. This method relies on the prediction that the presence of a mutation among the multiple target sequences in a given dot is sufficiently rare that any positive hybridization signal results from a single ASO within the probe mixture hybridizing with the corresponding mutant target. The hybridizing ASO is then identified by isolating it from the site of hybridization and determining its nucleotide sequence.
Suitable materials that can be used in the dot blot, reverse dot blot, multiplex, and MASDA formats are well-known in the art and include, but are not limited to nylon and nitrocellulose membranes.
When the target sequences are produced by PCR amplification, the starting material can be chromosomal DNA in which case the DNA is directly amplified. Alternatively, the starting material can be mRNA, in which case the mRNA is first reversed transcribed into cDNA and then amplified according to the well known technique of RT-PCR (see, for example, U.S. Pat. No. 5,561,058 by Gelfand et al.).
Alternative methods to the membrane-based methods described above include, but are not limited to, chip array hybridization. This method is described in detail in Hacia et al., 1996, Nature Genetics 14: 441-447, which is hereby incorporated by reference in its entirety. As described in Hacia et al., high density arrays of over 96,600 oligonucleotides, each 20 nucleotides in length, can be applied to a single glass or silicon chip. Each oligonucleotide can be designed to contain a pre-determined sequence that varies from the corresponding test gene sequence such that, collectively, the applied oligonucleotides contain small variations in the test gene sequence potentially spanning the entire test gene. The oligonucleotides applied to the chip, therefore, can contain pre-determined sequence variations that are not yet known to occur in the population, or they can be limited to mutations that are known to occur in the population. The immobilized oligonucleotides are then hybridized with a test polynucleotide sample. The hybridization pattern of the test polynucleotide sample is compared with that of the reference polynucleotide. Differences in hybridization can be localized to specific positions on the chip array containing individual oligonucleotides with pre-determined sequence variations.
In accordance with the invention, when no sequence variations are detected in a given test gene using the Stage I analysis described above, the presence of alternative sequence variations is subsequently tested by analyzing the test gene according Stage II, or Stage III, or Stage II followed by Stage III, as described in Sections 5.2 and 5.3, below.
5.2. STAGE II: SEQUENCE VARIATION LOCATING AND TARGETED CONFIRMATORY SEQUENCING
Test genes, such as those determined to contain none of the specific mutations tested in Stage I, can be further investigated for the presence of a variation in the nucleotide sequence of the test gene. According to Stage II analysis, a sequence variation locating assay is used to localize a sequence variation to a sub-region within the test polynucleotide. Such localization allows for subsequent targeted confirmatory sequencing to identify the specific sequence variation. In accordance with the invention, sequence variations can be located by initial detection at either the nucleic acid level (nucleotide sequence) or the protein level (amino acid sequence).
Stage II analysis can be used to identify sequence variations that represent mutations in the test gene. When no sequence variation is detected by Stage II analysis, or when the sequence variation is determined not to represent a mutation of the test gene, the test gene is subsequently subjected to Stage III analysis, as described in Section 5.3, below.
In a preferred embodiment of Stage II analysis, sequence variations that result in mutant polypeptides that are shorter than the corresponding normal polypeptide are initially detected at the protein level using a protein truncation assay.
5.2.1. PROTEIN TRUNCATION ASSAY
Mutations that cause a truncated protein product are detected by producing a protein product encoded by the test gene and analyzing its apparent molecular weight. A protein truncation assay provides the advantage of directly detecting nucleotide sequence variations that affect the encoded polypeptide, often in a dramatic fashion. Thus, in addition to localizing the site of the sequence variation for targeted nucleic acid sequencing, a protein truncation assay can provide immediate information of the effect of the sequence variation on the encoded protein.
Mutant truncated proteins can result from nonsense substitution mutations, frameshift mutations, in-frame deletions, and splice site mutations.
A nonsense substitution mutation occurs when a nucleotide substitution causes a codon that normally encodes an amino acid to code for one of the three stop signals (TGA, TTA, TAG). For such mutations, the protein truncation point occurs at the corresponding position in the gene at which the mutation occurs.
Frameshift mutations result from the addition or deletion of any number of bases that is not a multiple of three (e.g., one or two base insertion or deletion). For such frameshift mutations, the reading frame is altered from the point of mutation downstream. A stop codon, and resulting truncation of the corresponding encoded protein product, can occur at any point from the position of the mutation downstream.
In-frame deletions result from the deletion of one or more codons from the coding sequence. The resulting protein product lacks only those amino acids that were encoded by the deleted codons.
Splice site mutations result in an improper excision and/or joining of exons. These mutations can result in inclusion of some or all of an intron in the mRNA, or deletion of some or all of an exon from the mRNA. In some instances, these insertions or deletions result in stop codon being encountered prematurely, as typically occurs with frameshift mutations. In other instances, one or more specific exons is deleted from the mature mRNA in such a manner that the proper reading frame is maintained for the remaining exons, i.e., non-contiguous exons are fused in frame with each other. For such splice mutations, the encoded protein may terminate at the appropriate stop codon, but is shortened by the absence of the unspliced internal exon.
In a preferred embodiment, the length of the protein encoded by the test gene is analyzed by isolating all or a portion of the test gene for use as a template in an in vitro transcription/translation reaction. The template for transcription can be a cDNA copy of all or a portion of the gene. Using a cDNA template allows for translation initiation at the endogenous translation start site.
Alternatively, individual exons isolated from genomic DNA are tested separately. Such exons can be amplified using PCR from genomic DNA template. Since internal exons generally require a translation initiation signal (i.e., an ATG start codon), amplification primers are carefully selected such that the amplification product contains translation start signal in the proper reading frame through incorporation of the primer sequence into the amplification product.
In addition, particularly long cDNAs or individual exons that encode very large protein products can be divided into separate fragments. Each fragment can then serve individually as a separate template for transcription. Such fragments can be designed and generated by amplification using carefully chosen primers. In amplifying any fragment that does not contain a translation initiation signal (e.g., ATG) at its 5'-end, primers that incorporate such a signal at the 5'-end in frame with the coding sequence are used.
The translated protein products are then analyzed by well-known techniques for determining the apparent molecular weight of the polypeptide chains, such as SDS polyacrylamide gel electrophoresis (SDS/PAGE).
Some types of truncated protein products may be difficult to detect by such SDS/PAGE analysis. Extremely short truncated protein products may be too small to be detected on the gel. In addition, protein products that are only slightly truncated, i.e., by one or a few amino acids, may not be resolvable from the normal length protein product also present on the gel (derived from the normal allele amplification product template present in the in vitro transcription/translation reaction). Therefore, if no truncated protein products are detected, the terminal regions of the amplification products must be sequenced to ensure that extremely short and nearly full-length protein products are not encoded by the test polynucleotide sample. Preferably, even if a truncation product is detected, the terminal regions are sequenced to rule out the presence of additional upstream mutations.
In addition, if a slightly truncated protein results from a small internal deletion (e.g., a small in-frame deletion), not only will such a protein be difficult to resolve on the gel, the mutation will not be detected by sequencing the termini of the polynucleotide. Such mutations can be detected, however, in Stage III analysis, as described in Section 5.3, below.
The use of such a protein truncation assay is described in detail in the example in Section 6.2, below, for the detection of truncated human BRCA1 proteins resulting from mutations in exon 11 of the human BRCA1 gene. Furthermore, the detection of the C3508G mutation in the human BRCA1 gene using the protein truncation assay described in Section 6.2 is described in the example in Section 7, below.
5.2.2. ASSAYS FOR NUCLEOTIDE SEQUENCE VARIATIONS
A variety of methods, in addition to protein truncation assay, can be used to detect and locate a sequence variation within a region of a test gene, including but not limited to single strand conformation polymorphism (SSCP) analysis, heteroduplex analysis (HA), denaturing gradient gel electrophoresis (DGGE), constant denaturant gel electrophoresis (CDGE), and chemical cleavage analysis. Each of these techniques can be used alone or in conjunction with each other. In each case, once the nucleotide sequence variation is localized, the implicated region is sequenced to identify the exact variation in a process referred to herein as targeted confirmatory sequencing. The altered nucleotide sequence is then analyzed to determine its predicted effect on the encoded protein, e.g., whether the change results in a change in the amino acid sequence.
In single strand conformation polymorphism (SSCP) analysis, a double stranded PCR product is heated to melt the duplex and then run single stranded on a gel. For a general discussion of SSCP analysis, see Weber, 1996, supra. The single stranded DNA folds on itself in a characteristic fashion. Sequence variations between two amplified allelic sequences alter this conformation and change the migration on the gel, either more quickly or more slowly. When a shift with respect to the migration rate of the normal allele is detected on the gel, the "shifted" fragment is sequenced to identify the specific mutation.
Heteroduplex analysis (HA) can also be used to detect and localize sequence variations to a region within the test gene. HA is described in detail in Gayther, et al., 1996, Am. J. Hum. Genet. 58: 451-456. PCR products from individuals who are heterozygous for a mutation, when heated and then allowed to reanneal, form four types of products: two homoduplexes (one normal:normal, one mutant:mutant) and two heteroduplexes (normal:mutant of each strand). The homoduplexes and heteroduplexes will migrate at different rates in a non-denaturing gel. The fragment that is abnormal is sequenced to identify the specific mutation.
In denaturing gradient gel electrophoresis (DGGE), DNA duplexes start to melt at one end in a gel of an increasing denaturant. Mutations in the DNA alter the way it melts as it moves through the gel. Sequencing is performed on the fragment that is abnormal to detect the specific mutation. For more detail, see, for example, Fodde & Losekoot, 1994, Hum. Mutat. 3: 83-94; and Sheffield et al., 1989, Proc. Natl. Acad. Sci. USA 86: 232-236.
Constant denaturant gel electrophoresis (CDGE) is described in detail in Smith-Sorenson et al., 1993, Human Mutation 2: 274-285 (see also, Anderson & Borreson, 1995, Diagnostic Molecular Pathology 4: 203-211). A given DNA duplex melts in a predetermined, characteristic fashion in a gel of a constant denaturant. Mutations alter this movement. An abnormally migrating fragment is isolated and sequenced to determine the specific mutation.
Chemical cleavage analysis is described in U.S. Pat. No. 5,217,863, by R. G. H. Cotton. Like heteroduplex analysis, chemical cleavage detects different properties that result when mismatched allelic sequences hybridize with each other. Instead of detecting this difference as an altered migration rate on a gel, the difference is detected in altered susceptibility of the hybrid to chemical cleavage using, for example, hydroxylamine, or osmium tetroxide, followed by piperidine.
5.3. STAGE III: NUCLEOTIDE SEQUENCING ANALYSIS
In accordance with the invention, when no mutations are detected after analysis by Stage I, or Stage II, or Stage I followed by Stage II, the test gene is subjected to Stage III sequencing analysis. Stage III sequencing differs from the confirmatory sequencing of Stage II by not being targeted to a particular region of the gene proposed or known to contain a sequence variation. Thus, Stage III sequencing is used to detect sequence variations in the test polynucleotide (e.g., part or all of a gene) by directly sequencing the test polynucleotide and comparing its nucleotide sequence to a corresponding reference sequence (e.g., a normal sequence of the gene).
In one embodiment of the invention, Stage III sequencing analysis is performed by sequencing the entire gene including all introns and exons from end-to-end. In another embodiment, only exons (one or more) are sequenced. In yet another embodiment, only introns (one or more) are sequenced. In additional embodiments only portions of one or more introns, or portions of one or more exons, or combinations of partial or whole introns and partial or whole exons are sequenced.
Sequencing can be carried out only so far as a mutation is detected. Alternatively, sequencing can be continued to any point beyond the detection of a mutation up to and including the sequencing of every nucleotide in the entire gene, including each exon, intron, regulatory region, and other non-transcribed regions constituting a part of the functional gene.
The nucleotide sequence, in accordance with the invention, can be determined for just one of the DNA strands of the gene (e.g., "forward direction" sequencing, or sequencing the "sense" strand), or, preferably, by sequencing both strands (i.e., both forward/sense strands and reverse/antisense strands).
In a preferred embodiment for the human BRCA1 gene, each of the 22 coding exons is sequenced (exons 2, 3, 5, 5-24).
A number of methods well-known in the art can be used to carry out the sequencing reactions. Preferably, enzymatic sequencing based on the Sanger dideoxy method is used.
The sequencing reactions can be analyzed using methods well-known in the art, such as polyacrylamide gel electrophoresis. In a preferred embodiment for efficiently processing multiple samples, the sequencing reactions are carried out and analyzed using a fluorescent automated sequencing system such as the Applied Biosystems, Inc. ("ABI", Foster City, Calif.) system. For example, PCR products serving as templates are fluorescently labeled using the Taq Dye Terminator.RTM. Kit (Perkin-Elmer cat# 401628). Dideoxy DNA sequencing is performed in both forward and reverse directions on an ABI automated Model 377.RTM. sequencer. The resulting data can be analyzed using "Sequence Navigator.RTM." software available through ABI.
Alternatively, large numbers of samples can be prepared for and analyzed by capillary electrophoresis, as described, for example, in Yeung et al., U.S. Pat. No. 5,498,324.
6. Example 1
Detection of Sequence Variations in Polynucleotides
The methods of the invention, which can be used to detect sequence variations in any polynucleotide sample, are demonstrated in the example set forth in this section, for the purpose of illustration, for one gene in particular, namely, the human BRCA1 gene. The BRCA1 gene is approximately 100,000 base pairs of genomic DNA encoding the 1836 amino acid BRCA1 protein. The sequence is divided into 24 separate exons. Exons 1 and 4 are noncoding, in that they are not part of the final functional BRCA1 protein product. Each exon consists of 200-400 bp, except for exon 11 which contains about 3600 bp (Weber, B., Science & Medicine (1996).
The consensus sequence for the coding region of the human BRCA1 gene, referred to herein as BRCA1.sup.(omi1) (SEQ ID NO:1, herein), was first disclosed in co-pending application Ser. No. 08/598,591 (as SEQ ID NO:1, therein), which is hereby incorporated by reference in its entirety. The BRCA1.sup.(omi1) is more highly represented in the population than the previously disclosed BRCA1 coding sequence available as GenBank Accession Number U14680. In particular, it has a different nucleotide at each of seven polymorphic positions in the gene. These polymorphisms are summarized in Table 1, below. In addition to the BRCA1.sup.(omi1) sequence, co-pending application Ser. No. 08/798,691, filed Feb. 12, 1997, describes the previously unknown BRCA1.sup.(omi2) SEQ ID NO.3 and BRCA1.sup.(omi3) alleles SEQ ID NO.3 of the BRCA1 gene which also differ from the U14680 sequence: BRCA1.sup.(omi2) SEQ ID NO.3 differs from the U14680 sequence at six of the seven polymorphic sites and differs from BRCA1.sup.(omi1) at only one of the seven polymorphic sites; BRCA1.sup.(omi3) differs from the U14680 sequence at one of the seven polymorphic sites (see Table 1, below). The BRCA1.sup.(omi2) SEQ ID NO.3 and BRCA1.sup.(omi3) SEQ ID NO.5 sequences are described as SEQ ID NOS: 3 and 5, respectively, in co-pending application Ser. No. 08/798,691.
The seven polymorphisms most commonly occurring in the population are summarized in Table 1, below.
6.1 ALLELE-SPECIFIC OLIGONUCLEOTIDE (ASO) ANALYSIS OF MUTATIONS IN THE BRCA1 GENE
6.1.1 Isolation of Genomic DNA
White blood cells were collected from the patient and genomic DNA was extracted from the white blood cells according to well-known methods (Sambrook, et al., Molecular Cloning, A Laboratory Manual, 2nd Ed., 1989, Cold Spring Harbor Laboratory Press, at 9.16-9.19).
6.1.2. PCR Amplification
The genomic DNA was used as a template to amplify a separate DNA fragment encompassing the site of each of the eight mutations to be tested. Each 50 .mu.l PCR reaction contained the following components: 1 .mu.l template (100 ng/.mu.l) DNA, 5.0 .mu.l 10.times.PCR Buffer (Perkin-Elmer), 5.0 .mu.l dNTP (2 mM each dATP, dCTP, dGTP, dTTP), 5.0 .mu.l Forward Primer (10 .mu.mM), 5.0 .mu.l Reverse Primer (10 mM), 0.5 .mu.l Taq DNA Polymerase (Perkin-Elmer). 25 mM MgCl.sub.2 was added to each reaction according to Table 2, below, and H.sub.2 O was added to 50 .mu.l. All reagents for each exon except the genomic DNA can be combined in a master mix and aliquoted into the reaction tubes as a pooled mixture.
For each exon analyzed for this mutation, the following control PCRs were set up:
(1) "Negative" DNA control (100 ng placental DNA (Oncor, Inc., Gaithersburg, Md.)
(2) Three "no template" controls
PCR for all exons was performed using the following thermocycling conditions:
Quality control agarose gel of PCR amplification:
The quality of the PCR products were examined prior to further analysis by electrophoresing an aliquot of each PCR reaction sample on an agarose gel. 5 .mu.l of each PCR reaction was run on an agarose gel along side a DNA mass ladder (Gibco BRL Low DNA Mass Ladder, cat# 10068-013). The electrophoresed PCR products were analyzed according to the following criteria:
Each patient sample must show a single band of the corresponding size indicated in Table 2. If a patient sample demonstrates smearing or multiple bands, the PCR reaction must be repeated until a clean, single band is detected. The only exceptions to this are for the mutations 1294del40 and T>Gins59. If a patient sample from one of these two exons (11B and 6, respectively) demonstrates two bands instead of one, it may indicate the presence of the mutation. The 1294del40 mutation shortens the size of the corresponding PCR product by 40 bp and the T>Gins59 mutation lengthens the size of the corresponding PCR product by 59 bp. Patients heterozygous for these mutations would have a normal sized PCR product from the normal allele, and an altered sized PCR product from the mutant allele.
If no PCR product is visible or if only a weak band is visible, but the control reactions with placental DNA template produced a clear band, the patient sample should be re-amplified with 2.times. as much template DNA.
All three "no template" reactions must show no amplification products. Any PCR product present in these reactions is the result of contamination. If any one of the "no template" reactions shows contamination, all PCR products should be discarded and the entire PCR set of reactions should be repeated after the appropriate PCR decontamination procedures have been taken.
The optimum amount of PCR product on the gel should be 50-100 ng, which can be determined by comparing the intensity of the patient sample PCR products with that of the DNA mass ladder. If the patient sample PCR products contain less than 50-100 ng, the PCR reaction should be repeated until sufficient quantity is obtained.
6.1.3. Binding PCR Products to Nylon Membrane
A. The PCR products were denatured no more than 30 minutes prior to binding the PCR products to the nylon membrane. To denature the PCR products, the remaining PCR reaction (45 .mu.l) and the appropriate positive control mutant gene amplification product were diluted to 200 ul final volume with PCR Diluent Solution (500 mM NaOH, 2.0 M NaCl, 25 mM EDTA) and mixed thoroughly. The mixture was heated to 95.degree. C. for 5 minutes, and immediately placed on ice and held on ice until loaded onto dot blotter, as described below.
The PCR products were bound to 9 cm by 12 cm nylon Zeta Probe Blotting Membrane (Bio-Rad, Hercules, Calif., catalog number 162-0153) using a Bio-Rad dot blotter apparatus. Forceps and gloves were used at all times throughout the ASO analysis to manipulate the membrane, with care taken never to touch the surface of the membrane with bare hands.
3MM filter paper [Whatman.RTM., Clifton, N.J.] and nylon membrane were pre-wet in 10.times.SSC from 20.times.SSC buffer stock (Appligene, France, catalog #130431), making sure the nylon membrane was wet thoroughly. The vacuum apparatus was rinsed thoroughly with dH.sub.2 O prior to applying the membrane. 100 .mu.l of each denatured PCR product sample was added to the wells of the blotting apparatus. Each row of the blotting apparatus contained a set of reactions for a single exon to be tested, including a placental DNA (negative) control, a positive control, and three no template DNA controls. Information on the positive controls is listed in Table 3, below. Plasmid clones and PCR products were obtained from individual patients, and are interchangeable with synthetic oligonucleotides containing the respective mutant sequence.
When all liquid was suctioned through, the nylon membrane was placed on a piece of dry 3MM filter paper to prepare for fixing the DNA to the membrane.
The nylon filter was placed DNA side up on a piece of 3MM filter paper saturated with denaturing solution (1.5M NaCl, 0.5 M NaOH) for 5 minutes. The membrane was transferred to a piece of 3MM filter paper saturated with neutralizing solution (1M Tris-HCl, pH 8, 1.5 M NaCl) for 5 minutes. The neutralized membrane was then transferred to a dry 3MM filter DNA side up, and exposed to ultraviolet light (Stralinker, Stratagene, La Jolla, Calif.) for exactly 45 seconds. Shorter exposure time can result in the DNA not binding to the filter, and longer exposure time can result in degradation of the DNA. This UV crosslinking should be performed within 30 min. of the denaturation/neutralization steps. The nylon membrane was then cut into strips such that each strip contained a single row of blots of one set of reactions for a single exon.
6.1.4. Hybridizing Labeled oligonucleotides to the Nylon Membrane
6.1.4.1. Prehybridization
The eight strips were prehybridized using the Hybaid.RTM. (Savant Instruments, Inc., Holbrook, N.Y.) hybridization apparatus and bottles. The Hybaid.RTM. bottles, each containing approximately 2ml of 2.times.SSC, were preheated to 52.degree. C. in the Hybaid.RTM. oven. For each nylon strip, a single piece of nylon mesh cut slightly larger than the nylon membrane strip (approximately 1".times.5") was pre-wet with 2.times.SSC. Each single nylon membrane was removed from the prehybridization solution and placed on top of the nylon mesh. The membrane/mesh "sandwich" was then transferred onto a piece of parafilm. The membrane/mesh sandwich was rolled lengthwise and place into the appropriate Hybaid.RTM. bottle, such that the rotary action of the Hybaid.RTM. apparatus caused the membrane to unroll. The bottle was capped and gently rolled to cause the membrane/mesh to unroll and to evenly distribute the 2.times.SSC, making sure that no air bubbles formed between the membrane and mesh or between the mesh and the side of the bottle. The 2.times.SSC was discarded and replaced with 5 mls TMAC Hybridization Solution, which contained 3 M TMAC (Sigma T-3411), 100 mM Na.sub.3 PO.sub.4 (pH6.8), 1 mM EDTA, 5.times.Denhardt's (1% Ficoll, 1% polyvinylpyrrolidone, 1% BSA (fraction V)), 0.6% SDS, and 100 ug/ml Herring Sperm DNA. The filter strips were prehybridized at 52.degree. C. with medium rotation (approx. 8.5 setting on the Hybaid.RTM. speed control) for at least one hour. Prehybridization can also be performed overnight.
6.1.4.2. Labeling Oligonucleotides
The DNA sequences of the oligonucleotide probes used to detect the eight BRCA1 mutations were as follows (for each mutation, a mutant and a normal oligonucleotide must be labeled):
Each labeling reaction contained 2 .mu.l 5.times.Kinase buffer (or 1 .mu.l of 10.times.kinase buffer), 5 .mu.l gamma-ATP-.sup.32 P (not more than one week old), 1 .mu.l T4 polynucleotide kinase, 3 .mu.l oligonucleotide (20 .mu.M stock), sterile H.sub.2 O to 10 .mu.l final volume if necessary. The reactions were incubated at 37.degree. C. for 30 minutes, then at 65.degree. C. for 10 minutes to heat inactivate the kinase. The kinase reaction was diluted with an equal volume of sterile dH.sub.2 O (10 .mu.l). The labeled oligonucleotides were stored at -20.degree. C. until ready for use, but not for more than 21 days.
The oligos were purified on STE Micro Select-D, G-25 spin columns (catalog no. 5303-356769), according to the manufacturer's instructions. The 20 .mu.l eluate was diluted into 80 .mu.l dH.sub.2 O. The amount of radioactivity in the oligo sample was determined by measuring the radioactive counts per minute (cpm) in Probe Count. The total radioactivity must be at least 2 million cpm. For any samples containing less than 2 million total, the labeling reaction was repeated.
6.1.4.3. Hybridization with Mutant Oligonucleotides
2-5 million counts of each labeled mutant oligo probe was diluted into 5 mls of TMAC hybridization solution. 40 .mu.l of 20 .mu.M stock of unlabeled normal oligo was added to each tube containing probe and TMAC as an additional blocking reagent. Each probe mix was preheated to 52.degree. C. in the hybridization oven. The prehybridization solution was removed from each bottle and replaced with the probe mix, making sure that no air bubbles formed as with the prehybridization. The filters were hybridized for 1 hour at 52.degree. C. with moderate rotation rotation (approx. 8.5-setting on the Hybaid.RTM. speed control). Following hybridization, the probe mix was decanted into a storage tube and stored at 20.degree. C. Each filter was rinsed by adding approximately 20 mls of 2.times.SSC+0.1% SDS at room temperature and rolling the capped bottle gently for approximately 30 seconds and pouring off the rinse. At this point, all filters, except those for mutants 1249del40 and 876delAG can be pooled and washed together in the same container. Each filter, including those for mutants 1249del40 and 876delAG, were washed with 2.times.SSC+0.1% SDS at room temperature for 20 minutes, with shaking. The filters for mutants 1249del40 and 876delAG should not be removed from the hybridization bottle. After the initial room temperature wash another 20 mls of wash buffer was added to the bottle. The 8763delAG filter was washed at 52.degree. C. for 10 minutes. The 1294del40 filter was washed at 52.degree. C. for 20 minutes.
The eight separate strips were placed on 2 pieces of pre-cut Whatman paper and covered with plastic wrap. The efficiency of the hybridization was tested using a survey meter. The strips were autoradiographed at -80 using Kodak Biomax MS film in a cassette containing a biomax intensifying screen.
6.1.4.3. Control Hybridization with Normal Oligonucleotides
The purpose of this step is to ensure that the PCR products were all transferred efficiently to the nylon membrane. In other words, this step ensures that any samples that do not yield a hybridization signal with the mutant oligonucleotides are truly negative for the mutation, rather than false negatives resulting from poor transfer of the PCR products to the membrane.
Following hybridization with the mutant oligonucleotides, as described in Section 6.1.4.2, above, each nylon membrane was washed in 2.times.SSC, 0.1% SDS for 20 minutes at 65.degree. C. to strip off the mutant oligos. All eight nylon strips were prehybridized together in 35 mls of TMAC hybridization solution for at least 1 hour at 52.degree. C. in a shaking water bath. 2-5 million counts of each of the normal labeled oligo probes plus 40 .mu.l of 20 .mu.M stock of unlabeled normal oligo were added directly to the container containing the nylon membranes and the prehybridization solution. The filters and probes were hybridized at 52.degree. C. with shaking for at least 1 hour. Hybridization can be performed overnight. The hybridization solution was poured off; the nylon membrane was rinsed in 2.times.SSC, 0.1% SDS for 1 minute with gentle swirling by hand. The rinse was poured off and the membrane was washed in 2.times.SSC, 0.1% SDS at room temperature for 20 minutes with shaking.
The nylon membranes were removed and allowed to air dry for a few minutes, taking care not to let the nylon membranes dry completely. The nylon membranes were wrapped in one layer of plastic wrap and place on autoradiography film, as described in Section 6.1.2.4, above, except that exposure was for at least 1 hour.
For each sample, adequate transfer to the membrane is indicated by a strong autoradiographic hybridization signal. For each sample, an absence or weak signal when hybridized with its normal oligonucleotide, indicates an unsuccessful transfer of PCR product did not transfer successfully, and it is a false negative. The ASO analysis experiment must be repeated for any sample that did not successfully transfer to the nylon membrane.
6.1.5. Interpreting Results
After hybridizing with mutant oligos, the results for each exon are interpreted as follows:
After hybridization with normal oligos, interpret the results as follows:
6.2. PROTEIN TRUNCATION ASSAY
6.2.1. Amplification of Exon 11 from Genomic DNA
Exon 11 of the BRCA1 gene is approximately 3.4 kb long, and thus contains approximately 61% of the entire BRCA1 coding region. Three overlapping segments encompassing all of exon 11 were amplified directly from genomic DNA, isolated as described in Section 6.1.1, above. Fragment 1 encodes a 70.5 kD protein fragment, fragment 2 encodes a 74.8 kD protein fragment, and fragment 3 encodes a 41.5 kD protein fragment. The primers used for amplification are listed in Table 5 below. The ATG start codon that provided the translation initiation signal for each fragment is underlined.
6.2.2. In Vitro Transcription/Translation
The PCR amplification products were translated using the Tnt Coupled Reticulocyte Lysate System from Promega (cat #L4610). .sup.35 S-methionine from Amersham International (cat #SJ1015) was used to label the translation products. This particular "translation grade" methionine contains a stabilizer that retards the degradation of the .sup.35 S, and was not used more than 4 weeks after arrival. RNase inhibitor (Boehringer Mannheim--cat #799017) was added to the reactions.
Each 25 .mu.l sample contained 12.5 .mu.l rabbit reticulocyte extract, 1.0 .mu.l TnT reaction buffer, 0.5 .mu.l TnT T7 RNA polymerase, 0.5 .mu.l 1 mM amino acid mixture minus methionine, 0.5 .mu.l RNase inhibitor, 2.0 .mu.l S-35 Methionine (1 mCi/100 .mu.l), 500-750 ng DNA template, and deionized H.sub.2 O to 25 .mu.l.
The reaction was incubated at 30.degree. C. for 1 hour and 30 minutes, making sure that the temperature did not exceed 30.degree. C.
6.2.3. Gel Analysis
The patient samples are compared to the corresponding fragments from a normal gene by SDS PAGE analysis. Truncated proteins are possible at any point in the sequence. Areas in the sequence where a mutation can occur that are difficult to detect because of small molecular weight or a stop signal occurring at the end of the sequence are accounted for by two methods. First, 288 base pairs of the 5' end of exon 11 fragment 1 are sequenced. This assures detectability of proteins 10 kD or less by means of sequencing and detectablity of proteins greater than 10 kD by means of protein truncation. A stop occurring at the end of exon 11 fragment 3 is identified by sequencing the final 257 base pairs of exon 11. Second, the areas between fragments one and two and two and three overlap each other by at least 27 kD.
The truncated protein should be of equal intensity on the gel as the normal fragment. When a potential truncated band is identified, the protein must be sized using the S-35 radiolabelled marker. For example, if a band appears on the gel from fragment 3 (in addition to the normal 43 kD band), then the sample is positive for a protein truncation. In order to identify the exact mutation causing the truncation, the size of the truncated protein is estimated from the size ladder on the gel autoradiograph. The position of the stop signal does not always indicate the position where the mutation occurred. Using the conversion factor, 270 bp=10 kD, convert the molecular weight into base pairs. Refer to the breakdown sequence of exon 11 to locate the region needed to be sequenced for the mutation. Once this calculation is performed, the particular region of exon 11 in which the mutation is predicted to lie is sequenced according to the standard sequencing procedure (see Section 6.3 below).
6.3. NUCLEOTIDE SEQUENCING ANALYSIS
6.3.1. PCR Amplification
Genomic DNA (100 nanograms) extracted from white blood cells of the patient, as described in Section 6.1.1, above. Each sample was amplified in a final volume of 25 microliters containing 1 microliter (100 nanograms) genomic DNA, 2.5 microliters 10.times.PCR buffer (100 mM Tris, pH 8.3, 500 mM KCl, 1.2 mM MgCl.sub.2), 2.5 microliters 10.times.dNTP mix (2 mM each nucleotide), 2.5 microliters forward primer, 2.5 microliters reverse primer, and 1 microliter Taq polymerase (5 units), and 13 microliters of water.
The primers in Table 6, below were used to carry out amplification of the various sections of the BRCA1 gene samples. The primers were synthesized on an DNA/RNA Model 394.RTM. Synthesizer.
Thirty-five cycles were performed, each consisting of denaturing (95.degree. C.; 30 seconds), annealing (55.degree. C.; 1 minute), and extension (72.degree. C.; 90 seconds), except during the first cycle in which the denaturing time was increased to 5 minutes, and during the last cycle in which the extension time was increased to 5 minutes.
PCR products were purified using Qia-quick.RTM. PCR purification kits (Qiagen cat# 28104; Chatsworth, Calif.). Yield and purity of the PCR product determined spectrophotometrically at OD.sub.260 on a Beckman DU 650 spectrophotometer.
6.3.2. DNA Sequence Analysis
Fluorescent dye was attached to PCR products for automated sequencing using the Taq Dye Terminator.RTM. Kit (Perkin-Elmer cat# 401628). Dideoxy DNA sequencing was performed in both forward and reverse directions on an Applied Biosystems, Inc. (ABI) Foster City, Calif., automated Model 377.RTM. sequencer. The software used for analysis of the resulting data was "Sequence Navigator.RTM. software" obtained through ABI.
7. Example 2
Identification of the C3508g Mutation in the Brca1 Gene
A test BRCA1 gene sample was analyzed for mutations first using ASO analysis, as described in Section 6.1, above. The test gene was determined to be negative for all eight of the ASO mutations. Exon 11 of the test gene was then analyzed according to the protein truncation assay as described in Section 6.2, above.
Protein products of fragments 1 and 2 of the test sample were normal length of 70 kD and 74 kD, respectively. Fragment 3 of the test sample yielded a 15 kD protein in addition to the normal length 43 kD protein. The occurrence of a nonsense mutation at the point of truncation, which was about one third of the way or about 400 bp (270 bp/10 kD.times.15 kD=approximately 400 bp) into fragment 3, was tested first, prior to testing for a frameshift mutation at some point upstream of the truncation site. DNA sequence analysis of the coding region around 400 bp from the 5'-end of DNA fragment 3 confirmed the latter possibility that the mutation was a nonsense mutation, designated C3508G, resulting in a premature terminator codon (TCA to TGA) at position 1130 within exon 11 (see FIG. 1).
8. Example 3
Identification of the 5053delg Mutation in the Brca1 Gene
A test BRCA1 gene sample was tested for mutations first using ASO analysis, as described in Section 6.1, above. The test gene was determined to be negative for all eight of the ASO mutations. Exon 11 of the test gene was then analyzed according to the protein truncation assay as described in Section 6.2, above. All three protein fragments were the normal size indicating that exon 11 of the test gene did not contain a protein truncating mutation.
The DNA sequence of the test gene was then analyzed according to the sequencing analysis of Section 6.3, above. Comparison of the patient sequence with the BRCA1.sup.(omi) sequence revealed that the patient guanine (G) residue present at nucleotide position 5053 of the BRCA1 gene was deleted in the patient's BRCA1 gene. This deletion causes a frame shift in the coding region of exon 16. The frame shift results in an altered translation of the BRCA1 gene after nucleotide position 5053 until nucleotide positions 5089-5091, where a stop codon (TGA) occurs in the altered reading frame (see FIG. 1).
The present invention is not to be limited in scope by the specific embodiments described herein, which are intended as single illustrations of individual aspects of the invention, and functionally equivalent methods and components are within the scope of the invention. Indeed, various modifications of the invention, in addition to those shown and described herein will become apparent to those skilled in the art from the foregoing description and accompanying drawings. Such modifications are intended to fall within the scope of the appended claims.
Claims
What is claimed is:
1. A method for determining the presence or absence of a sequence variation in a gene sample, comprising the sequential steps of: (a) performing an allele specific hybridization assay for the presence or absence of one or more pre-determined sequence variations; (b) if no pre-determined sequence variation is found in step (a), then performing a sequence variation locating assay selected from the group consisting of protein truncation assay, single strand conformation polymorphism analysis, heteroduplex analysis, denaturing gradient gel electrophoresis, constant denaturant gel electrophoresis, and chemical cleavage analysis; (c) if no sequence variation is found in step (b), then sequencing the gene sample; and (d) determining the presence or absence of a sequence variation by analyzing the sequence(s) obtained in step (c) against a reference sequence.
2. A Method for determining the presence or absence of a sequence variation in a gene sample, comprising the sequential steps of: (a) performing an allele specific hybridization assay for the presence of one or more pre-determined sequence variations; (b) if no pre-determined sequence variation is found in step (a), then performing a sequence variation locating assay selected from the group consisting of protein truncation assay, single strand conformation polymorphism analysis, heteroduplex analysis, denaturing gradient gel electrophoresis, constant denaturant gel electrophoresis, and chemical cleavage analysis; (c) if a sequence variation is detected in step (b), then performing targeted confirmatory sequencing; and (d) determining the presence or absence of a sequence variation by analyzing the sequence(s) obtained in step (c) against a reference sequence.
3. A method for determining the presence or absence of a sequence variation in a gene sample, comprising the sequential steps of: (a) performing an allele specific hybridization assay for the presence or absence of one or more pre-determined sequence variations; and (b) if no pre-determined sequence variation is found in step (a), then sequencing the gene sample; and (c) determining the presence or absence of a sequence variation by analyzing the sequence(s) obtained in step (b) against a reference sequence.
4. A method for determining the presence or absence of a sequence variation in a gene sample, comprising the sequential steps of: (a) performing a sequence variation locating assay selected from the group consisting of protein truncation assay, single strand conformation polymorphism analysis, heteroduplex analysis, denaturing gradient gel electrophoresis, constant denaturant gel electrophoresis, and chemical cleavage analysis; (b) if no sequence variation is found in step (a), then sequencing the gene sample; and (c) determining the presence or absence of a sequence variation by analyzing the sequence(s) obtained in step (b) against a reference sequence.
5. The method of claim 1, 2, or 3 further comprising repeating the allele specific hybridization until a predetermined number of known sequence variations have been tested for.
6. The method of claim 5 wherein the allele specific hybridization assay is performed using a dot blot format.
7. The method of claim 5 wherein the allele specific hybridization assay is performed using a multiplex format.
8. The method of claim 1, 2, or 3 wherein the allele specific hybridization comprises testing for a predetermined number of sequence variations in a single step not requiring repetition.
9. The method of claim 8 wherein the allele specific hybridization assay is performed using a reverse dot blot format, a MASDA format, or a chip array format.
10. The method of claim 1, 2, or 4 wherein the sequence variation locating assay is performed using a protein truncation assay.
11. The method of claim 1, 2, or 4 wherein the sequence variation locating assay is performed using a chemical cleavage assay, a heteroduplex analysis, a single strand conformation, polymorphism assay, a constant denaturing gel electrophoresis assay, or a denaturing gradient gel electrophoresis assay.
12. The method of claim 1, 2, 3, or 4 wherein sequencing is performed in only the forward or reverse direction.
13. The method of claim 1, 2, 3, or 4 wherein sequencing is performed in both the forward and reverse directions.
14. The method of claim 1, 2, 3, or 4 wherein sequencing comprises sequencing both exons and introns of the gene or parts thereof.
15. The method of claim 14 wherein all exons and all introns are sequenced from end to end.
16. The method of claim 1, 2, 3, or 4 wherein sequencing comprises sequencing only exons.
17. The method of claim 1, 2, 3, or 4 wherein sequencing comprises sequencing only intronic sequences.
18. The method of claim 1, 2, 3, or 4 wherein the gene sample is a human BRCA1 gene.
19. The method of claim 1, 2, 3, or 4 wherein the reference sequence is a coding sequence.
20. The method of claim 19 wherein the reference sequence is a BRCA1 coding sequence.
21. The method of claim 1, 2, 3, or 4 wherein the reference sequence is a genomic sequence.
22. The method of claim 21 wherein the reference sequence is a BRCA1 genomic sequence.
23. The method of claim 1, 2, 3, or 4 wherein the reference sequence is one or more exons of a gene of interest.
24. The method of claim 1, 2, or 3 wherein the predetermined sequence variation in step (a) is a known mutation.
25. The method of claim 1, 2, 3, or 4, wherein the sequence variation is a known mutation.
Patent Citations (10)
| Patent | Date | Inventor | Cited By |
|---|---|---|---|
| US4683202 | 1987-07-01 | Mullis | |
| US5217863 | 1993-06-01 | Cotton et al. | |
| US5498324 | 1996-03-01 | Yeung et al. | |
| US5545527 | 1996-08-01 | Stevens et al. | |
| US5561058 | 1996-10-01 | Gelfand et al. | |
| US5582989 | 1996-12-01 | Caskey et al. | |
| US5589330 | 1996-12-01 | Shuber | |
| US5654155 | 1997-08-01 | Murphy et al. | |
| US5710001 | 1998-01-01 | Skolnick et al. | |
| US5750400 | 1998-05-01 | Murphy et al. |
Non-Patent Literature (14)
- Struewing et al. Am. J. Hum. Genet. 57:1-5, Jul. 1995.
- Hacia et al. Nature genetics 14:441-447, Dec. 1996.
- Dowton et al. Clinical Chemistry 41:785-794, May 1995.
- Sutcharitchan et al. Curr. Op. Hematol. 3:131-138, Mar. 1996.
- Andersen, T. and Borresen, A., "Alterations of the TP53 Gene as a Potential Prognostic Marker in Breast Carcinomas," Diagnostic Molecular Pathology 4(3):203-211 (1995).
- Easton et al., "Genetic Linkage Analysis in Familial Breast and Ovarian Cancer: Results from 214 Families," American Journal of Human Genetics 52:678-7091 (1993).
- Friend et al., "Breast Cancer Information on the web," Nature genetics 11:238-239 (1995).
- Gayther et al., "Rapid detection of Regionally Clustered Germ-Line BRCA 1 Mutations by Multiplex Heteroduplex analysis," Am. J. Hum. Genet. 58:451-456 (1996).
- Hacia et al., "Detection of Heterozygous Mutations in BRCA 1 Using High Density Oligonucleotide Arrays and Two-Colour Fluorescence Analysis," Nature Genetics 14(4):441-447 (1996).
- Merajver, S. and Petty, E., "Risk assessment and presymptomatic molecular diagnosis in hereditary breast cancer," Clinics in Lab. Med. 16(1):139-167 (1996).
- Michalowsky et al., "Combinatorial Probes for Identifications of > 100 Known Mutations in Hundreds of patients Samples Simultaneousl Using MASDA (Multiplex Allele-Specific Diagnostic Assay)," American Journal of Human Genetics 59(4):A272, poster 1573 (1996).
- Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Ed., Cold Spring Harbor Laboratory Press (1989).
- Sheffield et al., "Attachment of a 40-base-pair G+C-rich sequence (GC-clamp) to genomic DNA fragments by the polymerase chain reaction results in improved detection of Single-base changes," Proc. Natl. Acad. Sci. USA 86:232-236 (1989).
- Weber, B., "Genetic Testing for Breast Cancer," Scientific American SCIENCE & MEDICINE Jan.-Feb.:12-21 (1996).