Comparative Analysis of Spectral Clustering and K-Means for Educational Resource Clustering in West Java, Indonesia

Authors

  • Hana Nurul Wahidah Statistika
  • Suwanda Fakultas FDS, Universitas Islam Bandung

Keywords:

Clustering, Comparative, Education, K-Means, Spectral Clustering

Abstract

Abstract. Clustering is a widely used unsupervised learning technique for identifying groups of objects with similar characteristics. Among various clustering methods, K-Means is widely adopted because of its simplicity and computational efficiency; however, its performance is often limited when data exhibit complex structures. Spectral Clustering addresses this limitation by transforming the original data into a lower-dimensional spectral space derived from a similarity matrix before the clustering process. This study compares the performance of Spectral Clustering and K-Means in clustering educational resources across the 27 regencies and cities of West Java Province, Indonesia. The analysis was conducted using three educational resource indicators, namely the numbers of teachers, students, and schools. After Z-score standardization, both methods were evaluated using the Silhouette Score to determine the optimal number of clusters. The results show that Spectral Clustering consistently outperformed K-Means and achieved the highest Silhouette Score of 0.723 indicating a strong cluster structure with two clusters, indicating a strong cluster structure. The resulting clusters distinguish regions with relatively higher educational resources from those with comparatively lower educational resources, providing meaningful insights into regional disparities. These findings demonstrate that Spectral Clustering provides better clustering performance than K-Means for the educational resource dataset analyzed in this study and may support more equitable educational resource planning and policy development in West Java.

References

J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, Third. Morgan Kaufman, 2012.

B. S. Everitt, S. Landau, M. Leese, and D. Stahl, Cluster Analysis, Fifth. Wiley, 2011.

M. Qori’atunnadyah, “Pengelompokkan Wilayah Berdasarkan Rasio Guru-Murid Pada Jenjang Pendidikan Menggunakan Algoritma K-Means,” Journal of Informatics Development, 2022, doi: https://doi.org/10.30741/jid.v1i1.898.

A. Y. Ng, M. I. Jordan, and Y. Weiss, “On Spectral Clustering: Analysis and an algorithm,” Adv. Neural Inf. Process. Syst., 2002.

S. Trivedi and N. T. Heffernan, “Spectral Clustering in Educational Data Mining,” in 4th International Conference on Educational Data Mining, 2011. Accessed: Jul. 28, 2026. [Online]. Available: https://www.researchgate.net/publication/221570527_Spectral_Clustering_in_Educational_Data_Mining

X. Hong, J. Wang, and G. Qi, “Comparison of spectral clustering, K-clustering and hierarchical clustering on e-nose datasets: Application to the recognition of material freshness, adulteration levels and pretreatment approaches for tomato juices,” Chemometrics and Intelligent Laboratory Systems, vol. 133, pp. 17–24, Apr. 2014, doi: 10.1016/j.chemolab.2014.01.017.

O. Arbelaitz, I. Gurrutxaga, J. Muguerza, J. M. Pérez, and I. Perona, “An extensive comparative study of cluster validity indices,” Pattern Recognit., vol. 46, no. 1, pp. 243–256, Jan. 2013, doi: 10.1016/j.patcog.2012.07.021.

A. Khusaeri, “Regional Priority Analysis for Equal Distribution of Educational Facilities in West Java Using the Analytical Hierarchy Process (AHP) Method,” International Journal of Informatics, Economics, Management and Science, vol. 5, no. 1, p. 95, Jan. 2026, doi: 10.52362/ijiems.v5i1.1600.

I. Herdiana, “Membaca Data Ketimpangan Pendidikan, Ternyata Jawa Barat Bukan Bandung Saja,” BandungBergerak. Accessed: Jul. 23, 2026. [Online]. Available: https://bandungbergerak.id/article/detail/1546036539/membaca-data-ketimpangan-pendidikan-ternyata-jawa-barat-bukan-bandung-saja

CNN Indonesia, “Kemendikbud: Sekolah Kekurangan 1 Juta Guru Hingga 2024 .” Accessed: Feb. 27, 2026. [Online]. Available: https://www.cnnindonesia.com/nasional/20201005180513-20-554645/kemendikbud-sekolah-kekurangan-1-juta-guru-hingga-2024

New Indonesia, “JPPI Ungkap 5 Krisis Pendidikan di Jabar, Aoa Saja?” Accessed: Jul. 23, 2026. [Online]. Available: https://www.new-indonesia.org/2025/07/6916/jppi-ungkap-5-krisis-pendidikan-di-jabar-apa-saja-baca-artikel-detikedu-jppi-ungkap-5-krisis-pendidikan-di-jabar-apa-saja-selengkapnya-https-www-detik-com-edu-edutainment-d-8026729-jppi-un/

M. Walesiak and A. Dudek, “Selecting The Optimal Multidimentional Scaling Procedure for Metric Data with R Environment,” Statistics in Transition, vol. 18, no. 3, pp. 521–540, 2017.

P. Zerzucha and B. Walczak, “Concept of (dis)similarity in data analysis,” Sep. 2012. doi: 10.1016/j.trac.2012.05.005.

U. Von Luxburg, “A Tutorial on Spectral Clustering,” Stat. Comput., vol. 17, no. 4, 2007, [Online]. Available: www.springer.com.

Suliadi and R. R. Marliana, “Unsupervised Learning: Clustering,” in Machine Learning, Program Studi Statistika Universitas Islam Bandung, 2024.

P. J. Rousseeuw, “Silhouettes: a graphical aid to the interpretation and validation of cluster analysis,” J. Comput. Appl. Math., vol. 20, pp. 53–65, 1987, doi: https://doi.org/10.1016/0377-0427(87)90125-7.

Z. Teng, J. Yan, D. Liu, and P. Zhang, “When Does the Silhouette Score Work? A Comprehensive Study in Network Clustering,” Dec. 2025, doi: https://doi.org/10.48550/arXiv.2512.24841.

L. Kaufman and P. J. Rousseeuw, Finding Groups in Data: An Introduction to Cluster Analysis. John Wiley & Sons, Inc., 1990.

Published

2026-08-02