MeShClust<sup>2</sup>: Application of alignment-free identity scores in clustering long DNA sequences

Bt, James; Hz, Girgis

doi:10.1101/451278

Cited by 7 publications

(3 citation statements)

References 41 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Further, clustering algorithms, e.g. k -means and mean shift ( 35 , 36 ), can utilize Identity in grouping similar sequences. Finally, this methodology was also utilized in Look4TRs ( 34 ) where it was applied to finding repeated motifs in tandem repeats.…”

Section: Discussionmentioning

confidence: 99%

“…This design was inspired by our earlier research. We have successfully implemented adaptive software tools using self-supervised learning algorithms for locating cis-regulatory modules ( 32 ), identifying DNA repeats ( 33 , 34 ), and for clustering DNA sequences ( 35 , 36 ). Multiple software tools we developed earlier utilize general linear models (GLM) ( 34–40 ).…”

Section: Introductionmentioning

confidence: 99%

See 1 more Smart Citation

Identity: rapid alignment-free prediction of sequence alignment identity scores using self-supervised general linear models

Girgis

James

Luczak

2021

NAR Genomics and Bioinformatics

View full text Add to dashboard Cite

Pairwise global alignment is a fundamental step in sequence analysis. Optimal alignment algorithms are quadratic—slow especially on long sequences. In many applications that involve large sequence datasets, all what is needed is calculating the identity scores (percentage of identical nucleotides in an optimal alignment—including gaps—of two sequences); there is no need for visualizing how every two sequences are aligned. For these applications, we propose Identity, which produces global identity scores for a large number of pairs of DNA sequences using alignment-free methods and self-supervised general linear models. For the first time, the new tool can predict pairwise identity scores in linear time and space. On two large-scale sequence databases, Identity provided the best compromise between sensitivity and precision while being faster than BLAST, Mash, MUMmer4 and USEARCH by 2–80 times. Identity was the best performing tool when searching for low-identity matches. While constructing phylogenetic trees from about 6000 transcripts, the tree due to the scores reported by Identity was the closest to the reference tree (in contrast to andi, FSWM and Mash). Identity is capable of producing pairwise identity scores of millions-of-nucleotides-long bacterial genomes; this task cannot be accomplished by any global-alignment-based tool. Availability: https://github.com/BioinformaticsToolsmith/Identity.

show abstract

Section: Discussionmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

Identity: rapid alignment-free prediction of sequence alignment identity scores using self-supervised general linear models

Girgis

James

Luczak

2021

NAR Genomics and Bioinformatics

View full text Add to dashboard Cite

show abstract

“…The adaptation of the mean shift algorithm in MeSh-Clust v2.0 [41] is the same as that of MeShClust v1.0. The difference between the two versions is that the second version's classifier utilized in selecting similar sequences does not utilize alignment algorithms to generate identity scores for training; it uses similar idea to that implemented in Identity.…”

Section: Meshclust V20mentioning

confidence: 99%

MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores

Girgis

2022

Preprint

Self Cite

View full text Add to dashboard Cite

Background: Tools for accurately clustering biological sequences are among the most important tools in computational biology. Two pioneering tools for clustering sequences are CD-HIT and UCLUST, both of which are fast and consume reasonable amounts of memory; however, there is a big room for improvement in terms of cluster quality. Motivated by this opportunity for improving cluster quality, we applied the mean shift algorithm in MeShClust v1.0. The mean shift algorithm is an instance of unsupervised learning. Its strong theoretical foundation guarantees the convergence to the true cluster centers. Our implementation of the mean shift algorithm in MeShClust v1.0 was a step forward; however, it was not the original algorithm. In this work, we make progress toward applying the original algorithm while utilizing alignment-free identity scores in a new tool: MeShClust v3.0. Results: We evaluated CD-HIT, MeShClust v1.0, MeShClust v3.0, and UCLUST on 22 synthetic sets and five real sets. These data sets were designed or selected for testing the tools in terms of scalability and different similarity levels among sequences comprising clusters. On the synthetic data sets, MeShClust v3.0 outperformed the related tools on all sets in terms of cluster quality. On two real data sets obtained from human microbiome and maize transposons, MeShClust v3.0 outperformed the related tools by wide margins, achieving 55%-300% improvement in cluster quality. On another set that includes degenerate viral sequences, MeShClust v3.0 came third. On two bacterial sets, MeShClust v3.0 was the only applicable tool because of the long sequences in these sets. MeShClust v3.0 requires more time and memory than the related tools; almost all personal computers at the time of this writing can accommodate such requirements. MeShClust v3.0 can estimate an important parameter that controls cluster membership with high accuracy. Conclusions: These results demonstrate the high quality of clusters produced by MeShClust v3.0 and its ability to apply the mean shift algorithm to large data sets and long sequences. Because clustering tools are utilized in many studies, providing high-quality clusters will help with deriving accurate biological knowledge.

show abstract

Approximate Hashing for Bioinformatics

Arbitman

Klein

Peterlongo

et al. 2021

Implementation and Application of Automata

View full text Add to dashboard Cite

MeShClust²: Application of alignment-free identity scores in clustering long DNA sequences

Cited by 7 publications

References 41 publications

Identity: rapid alignment-free prediction of sequence alignment identity scores using self-supervised general linear models

Identity: rapid alignment-free prediction of sequence alignment identity scores using self-supervised general linear models

MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores

Approximate Hashing for Bioinformatics

Contact Info

Product

Resources

About

MeShClust2: Application of alignment-free identity scores in clustering long DNA sequences

Cited by 7 publications

References 41 publications

Identity: rapid alignment-free prediction of sequence alignment identity scores using self-supervised general linear models

Identity: rapid alignment-free prediction of sequence alignment identity scores using self-supervised general linear models

MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores

Approximate Hashing for Bioinformatics

Contact Info

Product

Resources

About

MeShClust²: Application of alignment-free identity scores in clustering long DNA sequences