Nucleotide Databases¶
Core_nt Database¶
What is the Core_nt Database?¶
The Core_nt database is a refined nucleotide sequence database optimized for speed and search relevance. Core_nt excludes some large eukaryotic chromosome assemblies that can be found in the NCBI Genome resource. It is ideal for characterizing sequences and finding homologs
What is in the Core_nt Database?¶
Core_nt contains sequences from the following categories:
Organisms from across the tree of life, including unknown or unclassified organisms.
Single-species environmental samples including DNA or RNA sequences obtained directly from environmental samples, such as soil, water, or air.
Artificially created sequence of nucleotides that can be designed to replicate natural genetic material.
How is the Core_nt Database Created?¶
The diagram below depicts how the Core_nt database is generated.
Number of sequences¶
Last updated¶
Protein Databases¶
ClusteredNR Database¶
What is the ClusteredNR Database?¶
NCBI ClusteredNR is a protein sequence database where similar sequences from the nr database are grouped into clusters. This creates a significantly smaller database that offers the same function for query identification as the nr database. The benefits of clustering include faster BLAST searches and broader taxonomic coverage.
What is in the ClusteredNR database?¶
It contains sequences for organisms from across the tree of life.
How is the ClusteredNR database created?¶
ClusteredNR is made from the nr protein database using the MMseqs2 computer program. This program groups protein sequences that are very similar to each other (90% identical and 90% length coverage) and provides a representative sequence that is most similar to other sequences in each group. The representative sequences are further processed such that they are replaced by a well-annotated NCBI RefSeq from the same group if available.
The diagram below depicts how the ClusteredNR database is generated.