Semi-Supervised Clustering for the Identification of Different Cancer Types Using the Gene Expression Profiles

View Sample PDF

Author(s): Manuel Martín-Merino (University Pontificia of Salamanca, Spain)
Copyright: 2012
Pages: 17
Source title: Medical Applications of Intelligent Data Analysis: Research Advancements
Source Author(s)/Editor(s): Rafael Magdalena-Benedito (Intelligent Data Analysis Laboratory, University of Valencia, Spain), Emilio Soria-Olivas (Intelligent Data Analysis Laboratory, University of Valencia, Spain), Juan Guerrero Martínez (Intelligent Data Analysis Laboratory, University of Valencia, Spain), Juan Gómez-Sanchis (Intelligent Data Analysis Laboratory, University of Valencia, Spain)and Antonio Jose Serrano-López (Intelligent Data Analysis Laboratory, University of Valencia, Spain)
DOI: 10.4018/978-1-4666-1803-9.ch004

Keywords: Health Information Systems / Medical Information Science Reference / Medicine & Healthcare

Purchase

View Semi-Supervised Clustering for the Identification of Different Cancer Types Using the Gene Expression Profiles on the publisher's website for pricing and purchasing information.

Abstract

DNA Microarrays allow for monitoring the expression level of thousands of genes simultaneously across a collection of related samples. Supervised learning algorithms such as -NN or SVM (Support Vector Machines) have been applied to the classification of cancer samples with encouraging results. However, the classification algorithms are not able to discover new subtypes of diseases considering the gene expression profiles. In this chapter, the author reviews several supervised clustering algorithms suitable to discover new subtypes of cancer. Next, he introduces a semi-supervised clustering algorithm that learns a linear combination of dissimilarities from the a priory knowledge provided by human experts. A priori knowledge is formulated in the form of equivalence constraints. The minimization of the error function is based on a quadratic optimization algorithm. A norm regularizer is included that penalizes the complexity of the family of distances and avoids overfitting. The method proposed has been applied to several benchmark data sets and to human complex cancer problems using the gene expression profiles. The experimental results suggest that considering a linear combination of heterogeneous dissimilarities helps to improve both classification and clustering algorithms based on a single similarity.

The IRMA Community

Research IRM

Semi-Supervised Clustering for the Identification of Different Cancer Types Using the Gene Expression Profiles

Purchase

Abstract

Related Content

IRMA Sponsors