TriSig: Assessing the statistical significance of triclusters

06/01/2023
by   Leonardo Alexandre, et al.
0

Tensor data analysis allows researchers to uncover novel patterns and relationships that cannot be obtained from matrix data alone. The information inferred from the patterns provides valuable insights into disease progression, bioproduction processes, weather fluctuations, and group dynamics. However, spurious and redundant patterns hamper this process. This work aims at proposing a statistical frame to assess the probability of patterns in tensor data to deviate from null expectations, extending well-established principles for assessing the statistical significance of patterns in matrix data. A comprehensive discussion on binomial testing for false positive discoveries is entailed at the light of: variable dependencies, temporal dependencies and misalignments, and p-value corrections under the Benjamini-Hochberg procedure. Results gathered from the application of state-of-the-art triclustering algorithms over distinct real-world case studies in biochemical and biotechnological domains confer validity to the proposed statistical frame while revealing vulnerabilities of some triclustering searches. The proposed assessment can be incorporated into existing triclustering algorithms to mitigate false positive/spurious discoveries and further prune the search space, reducing their computational complexity. Availability: The code is freely available at https://github.com/JupitersMight/TriSig under the MIT license.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/21/2017

Why "Redefining Statistical Significance" Will Not Improve Reproducibility and Could Make the Replication Crisis Worse

A recent proposal to "redefine statistical significance" (Benjamin, et a...
research
10/07/2021

Ranking Warnings of Static Analysis Tools Using Representation Learning

Static analysis tools are frequently used to detect potential vulnerabil...
research
07/25/2019

NoduleNet: Decoupled False Positive Reductionfor Pulmonary Nodule Detection and Segmentation

Pulmonary nodule detection, false positive reduction and segmentation re...
research
09/19/2021

Approximate Conditional Sampling for Pattern Detection in Weighted Networks

Assessing the statistical significance of network patterns is crucial fo...
research
04/09/2018

Cluster Failure Revisited: Impact of First Level Design and Data Quality on Cluster False Positive Rates

Methodological research rarely generates a broad interest, yet our work ...
research
09/03/2019

The Dynamics of Software Composition Analysis

Developers today use significant amounts of open source code, surfacing ...
research
07/06/2021

Furthering a Comprehensive SETI Bibliography

In 2019, Reyes Wright used the NASA Astrophysics Data System (ADS) t...

Please sign up or login with your details

Forgot password? Click here to reset