Breast Cancer Clasification Based on Gene Expression Data Using Singular Value Decompotition and Support Vector Machine

Authors

  • M Nadzir Soleh Matematika
  • Respitawulan Study Program of Mathematics, Faculty of Pharmacy and Science
  • Erwin H. Harahap Department of Mathematics , Faculty of Pharmacy and Science

Keywords:

Truncated Singular Value Decomposition, Support Vector Machine, Breast Cancer, Gene Expression

Abstract

Breast cancer gene expression data are typically high-dimensional, with the number of genes far exceeding the number of samples, which increases computational complexity and can reduce classification performance. This study aims to apply Truncated Singular Value Decomposition (TSVD) as a dimension reduction method and to evaluate the effect of rank selection on the performance of Support Vector Machines (SVMs) in classifying six breast cancer subtypes. The study employs a quantitative experimental approach using the GSE45827 dataset from the Curated Microarray Database (CuMiDa), which consists of 151 samples and 54,675 genes. The rank values were determined based on the threshold for retained information, greater than 1% and greater than 0.5%, resulting in rank-18 and rank-59. SVM models with linear and Radial Basis Function (RBF) kernels were evaluated using the F1-score and training time. The results show that the combination of TSVD with rank-18 and a linear kernel delivers the best performance, with an F1-score of 94.43% and a training time of 0.0054 seconds, outperforming the model without dimensionality reduction, which achieved an F1-score of 90.17% and a training time of 5.9684 seconds. These findings demonstrate that TSVD is capable of improving both computational efficiency and SVM classification performance on high-dimensional gene expression data.

References

Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA: A Cancer Journal for Clinicians. 2021;71(3): 209–249. https://doi.org/10.3322/caac.21660.

Khadijah K, Hartati S. Klasifikasi Data Microarray Menggunakan Discrete Wavelet Transform dan Extreme Learning Machine. IJCCS (Indonesian Journal of Computing and Cybernetics Systems). 2015;9(1): 33. https://doi.org/10.22146/ijccs.6638.

Golub TR, Slonim DK, Tamayo P, Huard C, Gaasenbeek M, Mesirov JP, et al. Molecular Classification of Cancer: Class Discovery and Class Prediction by Gene Expression Monitoring. Science. 1999;286(5439): 531–537. https://doi.org/10.1126/science.286.5439.531.

Ultimescu F, Hudita A, Popa DE, Olinca M, Muresean HA, Ceausu M, et al. Impact of Molecular Profiling on Therapy Management in Breast Cancer. Journal of Clinical Medicine. 2024;13(17): 4995. https://doi.org/10.3390/jcm13174995.

Clarke R, Ressom HW, Wang A, Xuan J, Liu MC, Gehan EA, et al. The properties of high-dimensional data spaces: implications for exploring gene and protein expression data. Nature Reviews Cancer. 2008;8(1): 37–49. https://doi.org/10.1038/nrc2294.

Berisha V, Krantsevich C, Hahn PR, Hahn S, Dasarathy G, Turaga P, et al. Digital medicine and the curse of dimensionality. NPJ digital medicine. 2021;4(1): 153.

Peryoga B, Adiwijaya A, Astuti W. Deteksi Kanker Berdasarkan Data Microarray Menggunakan Metode Naïve Bayes dan Hybrid Feature Selection. JURNAL MEDIA INFORMATIKA BUDIDARMA. 2020;4(3): 486. https://doi.org/10.30865/mib.v4i3.2096.

Simon G, Aliferis C. An Appraisal and Operating Characteristics of Major ML Methods Applicable in Healthcare and Health Science. In: Simon GJ, Aliferis C (eds.) Artificial Intelligence and Machine Learning in Health Care and Medical Sciences. Cham: Springer International Publishing; 2024. p. 95–195. https://doi.org/10.1007/978-3-031-39355-6_3.

Jia W, Sun M, Lian J, Hou S. Feature dimensionality reduction: a review. Complex & Intelligent Systems. 2022;8(3): 2663–2693. https://doi.org/10.1007/s40747-021-00637-x.

Anton H, Rorres C. Elementary linear algebra: applications version.. 11th edition. Hoboken, NJ: John Wiley & Sons Inc.; 2014.

Golub GH, Van Loan CF. Matrix computations.. Fourth edition. Johns Hopkins studies in the mathematical sciences. Baltimore: The Johns Hopkins University Press; 2013.

Eckart C, Young G. The approximation of one matrix by another of lower rank. Psychometrika. 1936;1(3): 211–218. https://doi.org/10.1007/BF02288367.

Boser BE, Guyon IM, Vapnik VN. A training algorithm for optimal margin classifiers. In: COLT92: 5th Annual Workshop on Computational Learning Theory. ACM; 1992. p. 144–152. https://doi.org/10.1145/130385.130401.

Cortes C, Vapnik V. Support-vector networks. Machine Learning. 1995;20(3): 273–297. https://doi.org/10.1007/BF00994018.

Ansari ZA, Arif M, Rajaboina NB, Shaikh AA, Singh Y. Integrated Ensemble Strategy for Breast Cancer Detection Using Dimensionality Reduction Technique. ADCAIJ: Advances in Distributed Computing and Artificial Intelligence Journal. 2025;14: e31899. https://doi.org/10.14201/adcaij.31899.

Yaqoob A, Verma NK. Feature Selection in Breast Cancer Gene Expression Data Using KAO and AOA with SVM Classification. Journal of Medical Systems. 2025;49(1): 40. https://doi.org/10.1007/s10916-025-02171-6.

Feltes BC, Chandelier EB, Grisci BI, Dorn M. CuMiDa: An Extensively Curated Microarray Database for Benchmarking and Testing of Machine Learning Approaches in Cancer Research. Journal of Computational Biology. 2019;26(4): 376–386. https://doi.org/10.1089/cmb.2018.0238.

Published

2026-08-04