Scalable privacy-preserving cancer type prediction with homomorphic encryption

04/12/2022
by   Esha Sarkar, et al.
3

Machine Learning (ML) alleviates the challenges of high-dimensional data analysis and improves decision making in critical applications like healthcare. Effective cancer type from high-dimensional genetic mutation data can be useful for cancer diagnosis and treatment, if the distinguishable patterns between cancer types are identified. At the same time, analysis of high-dimensional data is computationally expensive and is often outsourced to cloud services. Privacy concerns in outsourced ML, especially in the field of genetics, motivate the use of encrypted computation, like Homomorphic Encryption (HE). But restrictive overheads of encrypted computation deter its usage. In this work, we explore the challenges of privacy preserving cancer detection using a real-world dataset consisting of more than 2 million genetic information for several cancer types. Since the data is inherently high-dimensional, we explore smaller ML models for cancer prediction to enable fast inference in the privacy preserving domain. We develop a solution for privacy preserving cancer inference which first leverages the domain knowledge on somatic mutations to efficiently encode genetic mutations and then uses statistical tests for feature selection. Our logistic regression model, built using our novel encoding scheme, achieves 0.98 micro-average area under curve with 13 test accuracy than similar studies. We exhaustively test our model's predictive capabilities by analyzing the genes used by the model. Furthermore, we propose a fast matrix multiplication algorithm that can efficiently handle high-dimensional data. Experimental results show that, even with 40,000 features, our proposed matrix multiplication algorithm can speed up concurrent inference of multiple individuals by approximately 10x and inference of a single individual by approximately 550x, in comparison to standard matrix multiplication.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
11/09/2020

Privacy-Preserving XGBoost Inference

Although machine learning (ML) is widely used for predictive tasks, ther...
research
01/29/2022

A Novel Matrix-Encoding Method for Privacy-Preserving Neural Networks (Inference)

In this work, we present , a novel matrix-encoding method that is partic...
research
05/19/2020

Scalable Privacy-Preserving Distributed Learning

In this paper, we address the problem of privacy-preserving distributed ...
research
05/27/2023

Improved Privacy-Preserving PCA Using Space-optimized Homomorphic Matrix Multiplication

Principal Component Analysis (PCA) is a pivotal technique in the fields ...
research
04/07/2017

Privacy-Preserving Visual Learning Using Doubly Permuted Homomorphic Encryption

We propose a privacy-preserving framework for learning visual classifier...
research
05/16/2021

Private Facial Diagnosis as an Edge Service for Parkinson's DBS Treatment Valuation

Facial phenotyping has recently been successfully exploited for medical ...
research
01/23/2020

Data Inference from Encrypted Databases: A Multi-dimensional Order-Preserving Matching Approach

Due to increasing concerns of data privacy, databases are being encrypte...

Please sign up or login with your details

Forgot password? Click here to reset