Information retrieval in single cell chromatin analysis using TF-IDF transformation methods

12/10/2022
by   Mehrdad Zandigohar, et al.
0

Single-cell sequencing assay for transposase-accessible chromatin (scATAC-seq) assesses genome-wide chromatin accessibility in thousands of cells to reveal regulatory landscapes in high resolutions. However, the analysis presents challenges due to the high dimensionality and sparsity of the data. Several methods have been developed, including transformation techniques of term-frequency inverse-document frequency (TF-IDF), dimension reduction methods such as singular value decomposition (SVD), factor analysis, and autoencoders. Yet, a comprehensive study on the mentioned methods has not been fully performed. It is not clear what is the best practice when analyzing scATAC-seq data. We compared several scenarios for transformation and dimension reduction as well as the SVD-based feature analysis to investigate potential enhancements in scATAC-seq information retrieval. Additionally, we investigate if autoencoders benefit from the TF-IDF transformation. Our results reveal that the TF-IDF transformation generally leads to improved clustering and biologically relevant feature extraction.

READ FULL TEXT
research
09/21/2019

Application of Fuzzy Clustering for Text Data Dimensionality Reduction

Large textual corpora are often represented by the document-term frequen...
research
03/14/2023

Improving information retrieval through correspondence analysis instead of latent semantic analysis

Both latent semantic analysis (LSA) and correspondence analysis (CA) are...
research
12/16/2017

Taming Wild High Dimensional Text Data with a Fuzzy Lash

The bag of words (BOW) represents a corpus in a matrix whose elements ar...
research
03/19/2015

Reduced Basis Decomposition: a Certified and Fast Lossy Data Compression Algorithm

Dimension reduction is often needed in the area of data mining. The goal...
research
03/18/2021

Unsupervised Doppler Radar-Based Activity Recognition for e-healthcare

Passive radio frequency (RF) sensing and monitoring of human daily activ...
research
10/16/2018

Fast Randomized PCA for Sparse Data

Principal component analysis (PCA) is widely used for dimension reductio...
research
08/18/2020

Clustering and Analysis of Vulnerabilities Present in Different Robot Types

Due to the new advancements in automation using Artificial Intelligence,...

Please sign up or login with your details

Forgot password? Click here to reset