TAIR Glossary
This glossary defines terms as they are used in TAIR's database and website.
Many of the definitions used for TAIR terms were taken from dictionaries, text books, websites and other glossaries, including the ones listed in the following section.
Controlled Vocabularies
Gene Ontology Consortium. Enter a search term and click on Amigo to view, terms, defintions and annotated gene products.
Plant Ontology Consortium. Enter a search term and click on the ontology browser to view, terms, defintions and annotated gene products.
MESH Browser. Medical Subject Headings from the National Library of Medicine.
Swiss-Prot Knowledge Base:Definitions for keywords used at Swiss Prot.
Plants
Angiosperm Phylogeny website glossary:Glossary of botanical terms.
Taxonomy
NCBI Taxonomy: Search and browse NCBI's database of organisms.
Genetics/Molecular Biology
Dictionary of Cell and Molecular Biology:For unregistered users, access is limited to one day per every ninety days.
NHGRI Genome Glossary:Talking glossary from the National Human Genome Research Institute.
activation tag A DNA construct containing multiple enhancers of gene expression. When inserted into the genome the enhancers activate expression of adjacent genes, essentially creating a dominant mutation.
AFLP Amplified Fragment Length Polymorphism (AFLP) methods are used to detect multiple sequence variations in a single sample. The technology utilizes DNA fragments that are ligated to complementary adaptor oligonucleotide and subsequent rounds of PCR amplification using primers complementary to the adaptor sequences. The multiple rounds of amplification are used to reduce the complexity of the PCR product population so that the amplified fragments can be easily resolved by gel electrophoresis.
amino acid sequence Amino acid sequences for proteins from the start of translation to the terminator. Unless specifically noted, the sequences contain all amino acids present before any post translational modification occurs (e.g. cleavage of signal peptides).
anatomy A term describing a part of an organism and the components of those parts.
Annotation Annotations include explanatory notes or comments made regarding data in the database that are stored and displayed as comments on the detail pages. Comments may be contributed by curators from TAIR or other databases and members of the research community. Functional annotations generally take the form of associations of data such as genes, to keywords such as GO terms that describe the function, subcellular localization and processes involving the gene product. Structural annotations (e.g. transcription start sites, exons, introns, protein domains) are displayed on the SeqViewer as well as the detail pages for genes, loci and proteins.
Annotation Unit An annotation unit is a sequenced region used for annotating the genome.They correspond to clones used to sequence the Arabidopsis genome. However, annotation unit sequences frequently differ from the clone sequences in GenBank. This is because sequences may be trimmed or added to clones, to facilitate annotation of genes in the regions of overlap.
antimorphic A type of mutation in which the mutated gene product has an altered function that acts antagonistically to the wild type allele. Antimorphic mutants are usually dominant or semi-dominant.
antisense A construct in which a gene is fused to a promoter in 3' to 5' orientation. The resulting transcript is an 'antisense' transcript. Used in plants to suppress the expression of the endogenous gene.
Arabidopsis paralogs List of Arabidopsis genes that reside in the same gene family as the query Arabidopsis gene.
AraCyc Aracyc is a tool for querying and visualizing Arabidopsis biochemical pathways.Pathways can be browsed or searched by a variety of parameter such as reactions, compunds, genes and enzymes. Each pathway can be viewed from a gross perspective or zoomed in to display detailed sections. Reactions, compounds, genes and proteins are hyperlinked to detailed information about the selected objects.The expression viewer can be used to display pathways overlaid with gene expression data and supports visualization of time course experimental data.Pathways were initially derived computationally and then manually curated using available information from the literature.
attenuator An attenuator is the terminator sequence at which attenuation occurs. Attenuation describes the regulation of transcription that is involved in controlling the expression of bacterial operons.
BAC Bacterial artificial chromosome. Used for cloning DNA fragments between 100-200 kb in length. Based on a naturally occurring F-factor plasmid from Esherichia.coli.
backcross A seed generative method in which the progeny of a cross is mated back to one of the parental types.
BAC_end A type of clone end containing DNA fragments obtained from the end regions of a bacterial artificial chromosome (BAC) clone. Also refers to the BAC end sequences are used to order contigs and assemble scaffolds.
BiBAC Binary bacterial artificial chromosome. BiBAC vectors are used for cloning large DNA fragments (typically 100-200kb) in E.coli. BiBACs can be directly used as a plant transformation vector.
binding site A DNA sequence to which a protein ,such as a transcription factor, binds.
Biological Replicate True for replicate hybridizations that use samples obtained from different RNA extracts from different biological material
bulk protein download Used to obtain tables and lists of proteins and associated data such as subcellular localization, molecular weight, pI and domain information. Search all Arabidopsis proteins or search a subset of proteins by uploading a list of locus identifiers.
Bulk Sequence Download TAIR's Bulk Sequence Download tool can be used to obtain a defined set of nucleotide or amino acid sequences. For example it can be used to retrieve 1000 base pairs of upstream sequences from a set of co-regulated genes which can then be used to identify potential shared regulatory motifs. Sequence files can be obtained in one of two formats. You can choose to obtain FASTA formatted sequence files which are suitable for uploading into a variety of analysis programs such as CLUSTAL W, BLAST or a motif finding program. You can also obtain a tab delimited text file which can be opened in a standard spreadsheet program such as Microsoft Excel.
calculated molecular weight Predicted molecular weight (in daltons) of a protein.The predictions were calculated using the Bioperl function get_mol_wt (found in the SeqStats object of Bioperl 0.7).
calculated PI Predicted isoelectric point calculated using the iep program in the EMBOSS package. EMBOSS version: 2.0.1.
CAPS Cleaved Amplified Polymorphisms (CAPS) markers are a type of PCR based marker. In this method, genomic DNA is amplified by PCR using specific sets of primers complementary to a given region. The PCR product is then digested with a restriction enzyme and the products separated on an agarose gel. The polymorphisms detected by this method typically show an absence or presence of a restriction site leading to a difference in the size of DNA fragments.
cds The cds is the translated part of the gene from start codon to stop codon and excluding introns.
chemical Terms describing substances obtained by a chemical process or used for producing a chemical effect.
chromosome A linkage structure consisting of a specific linear sequence of genetic information. In TAIR there are 5 nuclear chromosomes (1-5) , the chloroplast chromosome (C) and the mitochondrial chromosome (M).
Chromosome Map Tool This tool can be used to display a customized map of all 5 Arabidopsis chromosomes decorated with a defined set of genes. The maps can be displayed at a variety of zoom levels or used to generate 'publication ready' maps for a gene family or any set of genes of interest.
classical mapping line Strains that are primarily used for classical genetic mapping. Individual plants are scored for visible traits and map locations are based upon recombination frequencies between visible alleles.
Clone A DNA fragment has been inserted into a vector molecule, such as a plasmid or a phage chromosome, and then replicated to form many identical copies.
clone seq Clones that are mapped onto the genome sequence include annotation units (e.g. BACS used to assemble the genome sequence), full length cDNAs and expressed sequence tags. The clone sequence is matched to the genome sequence according to a set of parameters defined for each clone type.
clone set Refers to stocks that are available as sets of individual clones such as sets of EST clones.
CloneEnd A term used at TAIR that includes sequences from the ends of clones as well as clones where the inserted DNA fragment comes from the END of a larger clone. Examples of subclones are BAC ends in which the end fragment of the BAC was inserted into another vector and sequenced or used for mapping. Clone ends that are sequences of clones are predominantly EST and plasmid ends.
coding region The DNA sequence of a gene that is translated into a protein.
Control replicate Indicates the replicate that is the baseline or reference of the corresponding replicate on the Replicate (id:name) column. This only applies to one-channel platforms.
cosmid A plasmid into which phage lambda cos sites have been inserted. As a result,the plasmid DNA can be packaged in vitro into the phage coat.
cre-lox recombination DNA constructs that contain either the Cre recombinase gene or the lox target sites for the Cre protein. The lox sites are used to direct the site of recombination when the Cre protein is present. When introduced into the plant genome, they can be used to generate deletions and other types of chromosomal rearrangements. Cre-lox systems have also been generated as tools for making chimeric sectored plants for analyzing the effect of loss of gene function in an otherwise wild type background.
Detection Affymetrix Detection algorithm uses probe pair intensities to generate a Detection p-value and assign a Present, Marginal, or Absent call. Each probe pair in a probe set is considered as having a potential vote in determining whether the measured transcript is detected (Present) or not detected (Absent). The p-value associated with this test reflects the confidence of the Detection call.
developmental stage A term describing a stage in the life cycle of an organism or a point in the ontogeny of an organ.
di-epoxybutane DNA cross linking agent that induces mutations.
directed acyclic graph A way of organizing objects according to their relationships to one another. The relationship between objects is directed; parent objects can have children. In contrast to simple hierarchies, children are allowed to have more than one parent, however they are not allowed to be their own parent (hence they are acyclic). The representation of controlled vocabularies in TAIR follows the form of directed acyclic graphs where each child (node) can have one or more parent (node).
dominant An allele that produces the same character whether present in the homozygous or heterozygous state.
dsRNA silencing A method used to induce post-transcriptional silencing of a target gene with the intention of generating a 'knock out' mutant phenotype. The RNAi (RNA interference) construct , which produces a dsRNA, is introduced into the plant genome via Agrobacterium mediated transformation. Subsequent cleavage of the dsRNA produces small interfering RNA's (siRNA) leading to degredation of the target gene transcript.
ecotype A genetic variety of a single species adapted for local ecological conditions.
enhancer trap A type of DNA construct used to identify regulatory elements controlling spatial or temporal patterns of gene expression. The construct includes a reporter gene fused to a minimal or weak promoter. When the construct is introduced into the genome of the organism , depending upon its position, the reporter gene may come under the control of nearby regulatory elements.
environmental Terms describing the environmental conditions an organism is subjected to.
ethyl-nitrosourea A nitrosoamine used as a chemical mutagen. Causes oxidative deamination of adenine and cytosine. Deamination of cytosine gives a C-G to T-A transition.
ethylmethane_sulfonate A chemical frequently used as a mutagen that typically induces base substitutions (such as transversions and transitions).
evidence code A controlled vocabulary of codes used to identify the type of evidence supporting an annotation. The codes are intended to provide a quick way to ascertain the strength of the evidence. For example, the code IEA is used for annotations that are made based purely by computational methods whereas IDA means that the annotation is supported by a direct assay (experimental data). The codes can be used to rank the quality of annotations.
evidence description A controlled vocabulary used to elaborate the types of evidence used to support annotations. Generally, the descriptions indicate the method (experimental, computational,or other) used to generate the supporting data.
exon A segment of an interrupted gene that is represented in the mature RNA product.
Experiment A set of microarray hybridizations that together comprise a single experiment designed to test a given hypothesis.
fast neutrons Mutagen that typically produces small deletions. Cloning of mutated genes is facilitated by PCR and subtractive hybridization methods to detect deletions.
filter A membrane containing DNA bound to the surface. Filters in TAIR are generally from BAC, YAC, or plasmid colony libraries.
gain of function A type of mutation in which the altered gene product possesses a new function or a new pattern of gene expression. Gain of function mutants are usually dominant or semidominant.
gamma rays Used to induce mutations.Induceds many types of mutations such as small deletions, large deletions, chromosome inversions and rearrangements.
Gene Class Symbol A symbolic name used as a core name for one or more genes. Symbolic names are generally derived based upon some functional characteristics of a class of genes or similar mutant phenotypes. An example of a functional gene symbol class is CYP71A which encompasses a set of homologous genes which may have CYtochromeP450 activity (hence CYP). An example of a mutant gene class symbol is CER (eceriferum) which includes a number of altered surface wax mutants.
Gene Hunter Gene hunter performs a meta-search for a gene by querying a number of different databases. Databases searched include TAIR, TIGR, PubMed, GenBank, Protein Information Resource, Swiss-PROT, and the Arabidopsis Genome Resource. Enter a gene name and choose to search one or all of these databases.
Gene Model In TAIR each gene entry corresponds to a single gene model. A gene model is defined as any description of a gene product from a variety of sources including computational prediction, mRNA sequencing, or genetic characterization.
gene trap Gene traps, also known as promoter traps or exon traps, are type of DNA construct containing a reporter gene downstream of a splice acceptor site. When the construct is inserted into the intron of a transcribed gene a transcriptional fusion with the reporter gene is generated. The reporter gene should then have the same temporal and spatial pattern of expression as the gene it is inserted into.
GeneModelType TAIR's classification of gene model types.
generation The number of generations represented by a germplasm. For example, the first generation of a cross between two parental lines would be the F1 (first filial generation), the second , F2 and so on.
generative method Methods used to create or propagate the line or pool.
GeneticMarker In TAIR, genetic markers are any biological object that is used to distinguish between two or more polymorphic states. Markers are distinct from polymorphisms in that they are associated with a specific method of detection of the polymorphism.
genic The parts of the genome that contain genes. The current definition of a genic region in TAIR does not include upstream of any 5' UTR's or downstream of any 3' UTRs. The presence of UTR sequences is determined by comparison to full length cDNA sequences and is not inferred.
genomic_DNA Genomic DNA includes all nuclear, mitochondrial and plastid DNA. Genomic DNA stocks from the ABRC are derived from an individual line.
genomic_DNA_pooled Stocks of genomic DNA isolated from pools of mutagenized plants. The pooled genomic DNA is usually screened using PCR to identify what pools contain an insertion/mutation in gene of interest.
GenPept An amino acid sequence database of translated proteins from GenBank.
GEO accession The unique identifier of the hybridization in NCBI
Germplasm In TAIR, germplasms correspond to individual strains having unique genotypes, pooled strains and sets of pools and strains.
GO Biological Process One of the three aspects covered by the Gene Ontology Consortium.A biological process is accomplished via one or more ordered assemblies of molecular functions. Usually there is some temporal aspect to it, although a process event may be essentially instantaneous. It often involves transformation, in the sense that something goes into a process and something different comes out of it. Examples of broad biological process terms are "cell growth and maintenance," or "signal transduction." Examples of more specific terms are "pyrimidine metabolism" or "cAMP biosynthesis. It is not equivalent to a pathway.
GO Cellular Component One of the three aspects covered by the Gene Ontology Consortium. A component of a cell with the proviso that the component is part of some larger object, which may be an anatomical structure, e.g. rough endoplasmic reticulum or nucleus, or a gene product group, e.g. ribosome, proteasome or a heterodimeric protein.
GO Molecular Function One of the three aspects covered by the Gene Ontology Consortium. A capability that a physical gene product (or gene product group) carries as a potential. It describes only what it can do without specifying where or when this usage actually occurs.
GO Slim A version of the Gene Ontology consisting of 10-20 terms from each of the three aspects of the GO ontologies (function, process and component). The terms represent broad categories that encompass all of their 'children' terms. For example, the GO slim process term "response to abiotic and biotic stimulus" encompasses all of the more specific processes such as response to abiotic stress and response to abiotic stress and each of their children.
haplo-insufficient A description applied to a gene that produces a mutant phenotype when present in a diploid individual heterozygous for an amorphic allele.
hyb_based Markers that are not RFLPs, where the primary method of detecting the polymorphism is by hybridization. This includes SNP polymorphisms detected by hybridization of genomic DNA to a microchip. Also included are markers (such as the MIT set) in which genomic DNA amplified using PCR and then probed with a radiolabelled allele- specific probe.
hypermorphic A type of mutation in which the altered gene product possesses an increased level of activity, or in which the wild-type gene product is expressed at a increased level.
hypomorphic A type of mutation in which the altered gene product possesses an increased level of activity, or in which the wild-type gene product is expressed at a increased level.
HZE High energy, radiocative nucleii used to induce mutations. The radiocative particles disrupt the chemical bonds of DNA and cause a variety of mutations.
HZE-C High energy, radiocative carbon nucleii used to induce mutations. The radiocative particles disrupt the chemical bonds of DNA and cause a variety of mutations.
HZE-Kr High energy, radiocative krypton nucleii used to induce mutations. The radiocative particles disrupt the chemical bonds of DNA and cause a variety of mutations.
HZE-Ne High energy, radiocative neon nucleii used to induce mutations. The radiocative particles disrupt the chemical bonds of DNA and cause a variety of mutations.
HZE-U High energy, radiocative uranium nucleii used to induce mutations. The radiocative particles disrupt the chemical bonds of DNA and cause a variety of mutations.
incompletely dominant An allele combination that produces a phenotype in the heterozygous state that is distinct from the dominant homozygote and the recessive homozygote phenotypes. Also known as semi-dominant.
INDEL Insertion/deletions (INDELS) are sequence variants where the Columbia (reference) polymorphism is an insertion relative to one ecotype and a deletion relative to another, different ecotype.
individual line A line descended from a single lineage. Strains may be propagated either through single seed descent or bulking the seeds from sibling plants.
individual pool A collection of germplasm from multiple individuals which have been combined together as a single unit.
intergenic The portions of a genome that are not considered to lie within a defined gene. Intergenic regions may overlap with genic regions on the complement strand. Intergenic region include sequences upstream and downstream of experimentally determined 5' and 3' UTRs. If no UTR sequences are determined, upstream and downstream intergenic sequences are derived from the position of the most 5' and 3' base pairs of the gene model.
interspecific cross A generative method in which the pollen donor and egg donor come from different species.
intron A segment of DNA in a gene that is transcribed but removed from the transcript by splicing together exon sequences from either side of the intron.
inverted repeat Identical adjacent nucleotide sequences in which one sequence is inverted with respect to the other.
ionizing radiation
back to topback to topback to top
lambda Any type of cloning vector (cosmids) containing cos sites necessary for propagation in bacteriophage lambda. Cosmid clones can be used to clone DNA fragments of up to 15 kb.
library A collection of clones in any type of vector.
Locus In TAIR a locus is a mapped element that corresponds to a transcribed region in Arabidopsis genome, or a genetic locus that segregates as a single genetic locus or quantitative trait. Loci are mapped based upon sequence match, or by recombination frequencies. A locus can have one or more associated gene models such as alternatively spliced variants.
loss of function A type of mutation in which the altered gene product lacks the function of the wild-type gene. Also called amorphic or null mutation.
Map Element A map element is any biological object that can be positioned on a map.
MapViewer The MapViewer graphically displays genetic, sequence and physical maps of each Arabidopsis chromosome. Each map can be searched, browsed and zoomed in to display more detailed views. Users can define which maps they want to display-for example show only the genetic, recombinant inbred line map and AGI sequence map for chromosome 1. The MapViewer was designed to aid forward and reverse genetic approaches to analyze gene functions. Maps can be aligned using common markers to facilitate identification of candidate genes, or to find candiate mutant phenotypes for a gene of interest.
match segment Regions of the genome upon which a segment of similar sequence is mapped. For example sequences flanking polymorphisms such as T-DNA insertion flanks or sequences at the 5' and 3' flanks of a substitution are all sequence matched segments.
method Terms describing procedures or techniques used in scientific research.
molecular mapping line Strains that are typically used for molecular mapping are generally recombinant inbred lines. RI lines are derived from a cross between parents with polymorphic genotypes. The progeny of the cross are selfed over several generations in so that they are homozygous at all loci, but each RI has a distinct recombinant geneotype. In these lines, genetic distance is calculated based upon the recombination frequencies between molecular markers (e.g. CAPS, SSLP markers).
Motif finder The Motif finder is used to find six base pair stretches (6-mers) that are over-represented within a given set of upstream sequences. It can be used to identify potential cis-regulatory sequences in a set of genes. The program compares the frequency of occurence of the 6-mers in a set of sequences against the frequency in all upstream regions of the genome.
multichannel True when more than one sample is hybridized at the same time onto the same slide, each one labeled differently.
NCBI BLink Protein records in TAIR link to the BLink record at the NCBI database. BLink displays the output of NCBI's precomputed BLASTP results for that protein searched against the non-redundant protein database. BLink results are displayed in a graphical format that includes alignments of the best 20 matches, protein domains, taxonomic data, and links to similar proteins with three dimensional structure information. BLink also allows you to filter the results based upon BLAST score and taxonomic groups.
nitroguanidine
nitrosomethyl biuret
nitrosomethyl urea An alkylating agent that is used as a chemical mutagen. Alkylation of reactive sites on bases with methyl or ethyl groups alters their H bonding and consequently,base pairing. Mutations can be transitions or transversions.
Normalization Description The description of the normalization method used
normalization factor The value used for normalizing all the spots on the slide. For two-channel platforms, this factor is used for balancing the Cy5 and Cy3. For one-channel platforms this value is used to make the Average Intensity of an experimental array the same as that of the Baseline array
open pollinated SF A method of seed generation in which natural self-fertilization is allowed to take place. Both the pollen donor and egg donor are from the plant.
ORF An open reading frame (ORF) corresponds to any nucleotide sequence that can potentially encode a protein. In TAIR ORFs include those that are experimentally verified by comparison to cDNA sequence as well as ORFs predicted by computational methods. Predicted ORFs include those encoding hypothetical proteins (e.g. there is no evidence that a transcript is generated that corresponds to the open reading frame).
origin of replication A sequence of DNA at which replication is initiated.
over expression A DNA construct in which a gene is fused to a promoter conferring a constitutive and or high level of expression.
P1 A plasmid vector derived from Bacteriophage P1 used to clone DNA fragments between 70-95 kb in length.
P1 end A type of clone end containing DNA fragments obtained from the end regions of a phage P1 clone. Also refers to the phage end sequences are used to order contigs and assemble scaffolds.
PatMatch The PatMatch program is used to find short nucleotide (less than 30 base pairs) and amino acid motifs in DNA or amino acid sequences. The program takes as input a simple text string (or regular expression) and finds all instances in the selected target dataset. The query string can allow for degenerate as well as exact matches. For example, it can be used to find all occurances of a given cis regulatory region in upstream sequences, or to find all proteins having a similar domain. The results are hyperlinked to TAIR database detail pages and can be downloaded as a text file.
PCR_based A type of genetic marker that uses PCR as the primary mode of detection. These include SNAP markers, in which allele specific primers are used to amplify the polymorphic region.
Pedigree The lineage or record of ancestors of a germplasm.
PFAM Pfam is a database of multiple alignments of protein domains or conserved protein regions. Pfam domains are generated using multiple sequence alignments to generate a seed alignment. Hidden Markov Models are then used to group related sequences. Pfam families include PfamA which are curated and PfamB which are computationally derived and not inspected by experts.
phenotype Any detectable manifestation of the genotype of an organism.
PIR The Protein Information Resource (PIR) produces the Protein Sequence Database (PSD) of functionally annotated protein sequences, which grew out of the Atlas of Protein Sequence and Structure (1965-1978) edited by Margaret Dayhoff and has been incorporated into an integrated knowledge base system of value-added databases and analytical tools.
Plant homologs Data displayed in this section are derived from the gene families of PANTHER 17.0 release.
Plant orthologs List of plant orthologs of the query Arabidopsis gene.
Plant Tissue the keywords that define the anatomy or developmental stage of the plant material employed in an experiment
plasmid An autonomously replicating, circular DNA molecule of bacteria that is independent of the chromosomal DNA. Used to clone small DNA fragments (ca.5 kb).
poly A site The polyadenylation site defines the place in the gene where addition of a sequence of polyadenylic acid to the 3' end of an RNA after transcription will occur.
pre trna The primary product of transcription of genes coding for transfer RNAs (t-RNA). The pre-tRNA of prokaryotes and eukaryotes has extra nucleotides at the 5' and 3' extremities and in some eukaryotic pre-tRNAs introns are also present. Maturation of tRNA precursors is a multistep enzymatic process consisting of nucleolytic size reducing reactions and of nucleotide modifications.
PRINTS The PRINTS database is a collection of protein fingerprints; conserved motifs characteristic of a protein family. Usually the motifs do not overlap, but are separated along a sequence, though they may be contiguous in 3D-space. Fingerprints can encode protein folds and functionalities more flexibly and powerfully than can single motifs, full diagnostic potency deriving from the mutual context provided by motif neighbors.
PRODOM Prodom families are built using a recursive PSI-BLAST homology search. Only those ProDom families included in INTERPRO are included in the ProDom assignments for Arabidopsis proteins.
promoter The part of a gene required for its transcriptional regulation. This includes specific regions for initiation of transcription and regions required for temporal, spatial and quantitative regulation.
promoter fusion A DNA construct in which a promoter from one gene is used drive expression of another gene . The promoter and gene can be from different organisms.
promoter reporter In DNA cloning, a construct in which the promoter of one gene is fused to another 'reporter' gene. Examples of reporter genes include, various fluorescent proteins such as green and yellow fluorescent proteins (GFP,YFP); beta glucoronidase (GUS) and others.
promoter trap A DNA construct which contains a reporter gene but lacks a promoter (regulatory sequences). If the reporter gene is inserted into a gene such that a transcriptional fusion is generated,the expression of the reporter is now under the control of the endogenous plant promoter.
PROSITE Prosite includes protein families and domains; biologically significant sites, patterns and profiles that help to reliably identify to which known protein family (if any) a new sequence belongs. Annotation of proteins with Prosite domains can be done automatically by searching with regular expressions for matches to Prosite patterns or by matching a sequence to a profile for finding more distantly related proteins.
protein coding A gene model whose product is a transcript which is subsequently translated into a protein.
protein domain Protein domains are conserved regions of amino acid /structural similarity in protein sequences. Domains generally represent functional units having some form of biological activity. Domains are useful in grouping proteins with little overall sequence similarity. Functions for unknown proteins may be inferred from the presence of conserved domains.
pseudogene A sequence of DNA that is very similar to a normal gene but that has been altered slightly so it is not expressed. Such genes were probably once functional but over time acquired one or more mutations that rendered them incapable of producing a protein product
quantitative trait locus A genetic locus identified through the statistical analysis of complex traits. These traits are typically affected by more than one gene, and also by the environment.
RAPD Randomly Amplified Polymorphic DNA (RAPD) markers. Genomic DNA are amplified by PCR using non-specific primers that are complementary to a number of sites within the genome. The PCR products are separated on an agarose gel to identify co-segregating polymorphism.
Replicate Hybridization The set of hybridizations that are performed with similar samples and arrays, which can be averaged. Replicates can be technical or biological.Technical replicate indicates that the same RNA sample was used in both replicates, either with the same label or with the dyes swapped. Biological replicate means that both hybridizations use independently isolated RNA from identically treated plants.
Replicate Hybridization Set The set of hybridizations that are replicates of each other, which can be averaged.
Replicate Set Signal Percentile Shows how the average signal value of the array element compares to average signal values of other array elements in the same replicate hybridization set. For example, an array element at the 90th percentile has average signal value greater than 90% of array elements in the replicate set. The standard error is the standard deviation of the sampling distribution of the signal percentile for the array element in the replicate hybridization set.
representative gene model In TAIR, this is the reference gene model for the locus. The Sequences and structural features for the locus are derived from the representative gene model.
reverse-dye replicate True if the RNA label was swap in a multichannel hybridization. Cy3 for the control sample and Cy5 for the test sample is considered the standard, otherways is considered a reversed-dye hybridization.
RFLP Restriction Fragment Length Polymorphisms. These markers detect polymorphic sequences using a labelled probe hybridized against digested genomic DNA. Polymorphic sequences which alter the restriction site (adding, removing a site or changing the size of the digest fragment) are separated on an agarose gel, transferred to a membrane and hybridized to a cloned piece of DNA.
ribosomal RNA The transcribed product of ribosomal DNA, also known as rRNA. These rRNA's are part of the ribosome
RNAi RNA interference (RNAi) is a method used to silence the expression of a target gene in order to determine the gene's function by analyzing the mutant phenotype. RNAi constructs are generated using sequences from the target gene are inserted into a vector in opposite orientations resulting in an inverted repeat. When the dsRNA is expressed in the target organism the endogenous gene is silenced.
Scale factor The value to make the average intensity of an array equal to an arbitrarily defined target intensity.
selected bulk A method of seed propigation in which seed from selected individuals in one generation is pooled.
selfing A method of seed generation in which pollination is performed manually and the pollen donor and egg donor are from the same plant.
sequence features Types of gene or genome features that can be mapped on the genome or a gene. These may be structural features of a gene, or characteristic sequence features of a genome (such as repeats or duplicated regions).
sequenced locus Sequenced loci include predicted and experimentally verified transcribed regions.
SeqViewer SeqViewer is a tool for viewing the Arabidopsis genome and its associated annotation from a whole genome view down to the nucleotide level. The SeqViewer displays gene annotations,annotation units (BAC clones used to generate and assemble the genome sequence), transcripts (including ESTs and full length cDNAs),polymorphisms, T-DNA/Transposon insertions, and markers on the Arabidopsis genome sequence. Users can search with up to 250 names or up to four short (< 150 nucleotide) nucleotide sequences, and visualize the locations of one or many search hits on the whole genome, in a closeup view (zoomable from 50 megabases to 10 kilobases) or in a 10 kilobase nucleotide window. Objects displayed in the SeqViewer are clickable and linked to TAIR database where you can find detailed records. Specific objects can be directly located on the SeqViewer from search results and detail pages.
set of lines A set of germplasm from individual plants.
set of pools A collection of pools each of which contains germplasm from multiple individuals.
Signal Quantitative metric calculated for each probe set using Affymetrix software, which represents the relative level of expression of a transcript. Affymetrix Signal calculation algorithm uses the One-Step Tukey?s Biweight Estimate which yields a robust weighted mean that is relatively insensitive to outliers, even when extreme. Each probe pair in a probe set is considered as having a potential vote in determining the Signal value. The vote is defined as an estimate of the real signal due to hybridization of the target. The mismatch intensity is used to estimate stray signal. The real signal is estimated by taking the log of the Perfect Match intensity after subtracting the stray signal estimate. The probe pair vote is weighted more strongly if this probe pair Signal value is closer to the median value for a probe set. Once the weight of each probe pair is determined, the mean of the weighted intensity values for a probe set is identified. This mean value is corrected back to linear scale and is output as Signal.
Signal Percentile Shows how a signal value of the array element compares to signal values of other array elements on the same slide or chip. For example, an array element at the 90th percentile has signal value greater than 90% of array elements on the chip.
simple insert DNA constructs that are primarily designed for the purpose of T-DNA or transposon tagging and contain little additional DNA other than plant selectable markers, T-DNA sequences or transposon sequences.
single genetic locus A genetically defined region corresponding to a single gene.
single seed descent A seed generation method where all of the progeny are descended from a single parent plant. Each subsequent generation is obtained in this manner.
Slides Information about a hybridized slide or chip.
small nucleolar RNA Small nuclear RNAs (snoRNA) that are involved in the processing of pre-ribosomal RNA in the nucleolus. Box C/D containing snoRNAs (U14, U15, U16, U20, U21 and U24-U63) direct site-specific methylation of various ribose moieties. Box H/ACA containing snoRNAs (E2, E3, U19, U23, and U64-U72) direct the conversion of specific uridines to pseudouridine. Site-specific cleavages resulting in the mature ribosomal RNAs are directed by snoRNAs U3, U8, U14, U22 and the snoRNA components of RNase MRP and RNase P.
SMART The SMART database includes more than 500 domain families found in signaling, extracellular and chromatin-associated proteins are detectable. These domains are extensively annotated with respect to phyletic distributions, functional class, tertiary structures and functionally important residues.SMART alignments are optimized manually and following construction of corresponding hidden Markov models (HMMs). Association of proteins to SMART families is done by searching with the protein sequence against the SMART library of HMMs.
sodium ascorbate
species variant In TAIR species variants include natural variants of Arabidopsis thaliana collected from a variety of sources (ecotypes) as well as related Genera and species of Arabidopsis.
spontaneous Mutations that occur in the absence of treatment with a chemical or biological mutagen. Usually refers to mutations in natural populations.
SSLP Simple Sequence Length Polymorphisms (SSLPs) are markers that detect differences in the length of a PCR product. Typically the differences are due to small insertions/deletions such as those caused by differences in the number of simple sequence repeats. Primers complementary to a given genomic region are used to amplify the region from genomic DNA and the resulting PCR products are separated on an agarose gel. Usually these markers are co-dominant, although in some cases, the difference is that one variant is detected and the other does not give a PCR product.
substitution Replacement of one nucleotide sequence by another nucleotide or one amino acid in a protein by another amino acid. Also known as a single nucleotide polymorphism (SNP).
SUPERFAM A designation used by SCOP for proteins that have low sequence identities, but whose structural and functional features suggest that a common evolutionary origin is probable are placed together in superfamilies. SCOP uses structural similarity to group proteins into families, superfamilies and folds.
SwissPROT Swiss-Prot is a curated protein sequence database which strives to provide a high level of annotations (such as the description of the function of a protein, its domains structure, post-translational modifications, variants, etc.), a minimal level of redundancy and high level of integration with other databases.
T-DNA The portion of the Agrobacterium tumafaciens Ti plasmid (including the terminal repeats) that is integrated into the host genome.
T-DNA insertion A physical mutagen that causes mutations by insertion of transfer DNA (with or without additional DNA sequences) into the genome. These tend to be large insertions in the kilobase range and thus are expected to knockout the gene it is inserted in.
TAC Transformation artificial chromosome used as a cloning vector for larger DNA fragments as well as for plant transformation. The TAC vector contains the P1 bacteriophage replicon, which maintains the vector in a single copy, and therefore renders foreign DNA fragments stable, in E. coli cells. The vector also contains the pRiA4 replicon of the Ri plasmid, which ensures a vector copy number of 1 in Agrobacterium tumafaciens.
taxonomy An ordered classification of organisms based upon presumed natural relationships.
Technical Replicate True for replicate hybridizations that use RNA samples obtained from the same or different RNA extracts of the same biological material
terminator A sequence of DNA found at the end of a transcript, that causes RNA polymerase to terminate transcription.
TIGRFAM TIGRFAM protein families are generated from curated multiple sequence alignments and Hidden Markov Models. Curated families are also associated to functional roles to facilitate automated functional identification of proteins by sequence homology.
tissue culture
transition A type of point mutation in which one purine or pyrimidine is replaced by another base of the same type. Examples: A->G and C->T.
transposon insertion Transposable elements (transposons) include a diverse class of DNA sequences that are capable of inserting, excising and relocating into chromosomal or extrachromosomal DNA. Autonomous transposons encode a transposase and are capable of transposing on their own. Transposition of non-autonomous elements requires trans-activation from the autonomous element. Insertion of a transposable element into a gene may create a knock-out, loss of function allele.