Analisis Segmentasi Mahasiswa Program Studi Statistika Universitas Islam Bandung Tahun 2020-2025 Menggunakan Algoritma K-Modes Clustering
DOI:
https://doi.org/10.29313/bcss.v6i2.24868Keywords:
Data Mining, K-Modes Clustering, Segmentasi MahasiswaAbstract
Abstract. Student segmentation is an important approach for understanding student characteristics and supporting higher education institutions in developing more effective and targeted promotional strategies. This study aims to segment students of the Statistics Study Program at Universitas Islam Bandung using the K-Modes Clustering algorithm. The dataset consisted of 376 student records described by four categorical variables, namely city of origin, high school background, admission pathway, and admission wave. The clustering process was performed by evaluating the number of clusters from k= 2 to k = 5. The clustering performance was assessed using the Silhouette Coefficient based on the Simple Matching Dissimilarity measure. The results indicated that the optimal clustering solution was obtained at k = 2, yielding the highest average Silhouette Coefficient value of 0.2301556. Furthermore, each cluster exhibited distinct characteristics in terms of city of origin, high school background, admission pathway, and admission wave. The findings of this study provide useful insights into student characteristics and may assist the Statistics Study Program at Universitas Islam Bandung in developing more effective and targeted promotional strategies.
Abstrak. Segmentasi mahasiswa merupakan salah satu pendekatan yang penting untuk memahami karakteristik mahasiswa serta mendukung perguruan tinggi dalam menyusun strategi promosi yang lebih efektif dan tepat sasaran. Penelitian ini bertujuan untuk melakukan segmentasi mahasiswa Program Studi Statistika Universitas Islam Bandung menggunakan algoritma K-Modes Clustering. Data yang digunakan terdiri dari 376 data mahasiswa dengan empat variable kategori, yaitu asal kota, asal sekolah, jalur seleksi, dan gelombang seleksi. Proses clustering dilakukan dengan menguji jumlah clustermulai dari k = 2 hingga k = 5. Kinerja hasil clustering dievaluasi menggunakan Silhouette Coefficient berdasarkan ukuran Simple Matching Dissimilarity. Hasil penelitian menunjukkan bahwa jumlah cluster optimal diperoleh pada k = 2 dengan nilai rata-rata Silhouette Coefficient tertinggi sebesar 0.2301556. Selain itu, masing-masing cluster menunjukkan karakteristik yang berbeda berdasarkan asal kota, asal sekolah, jalur seleksi, dan gelombang seleksi. Hasil penelitian ini memberikan informasi mengenai karakteristik mahasiswa dan diharapkan dapat menjadi bahan pertimbangan bagi Program Studi Statistika Universitas Islam Bandung dalam menyusun strategi promosi yang lebih efektif dan tepat sasaran.
References
[2] U. Fayyad, G. Piatetsky-Shapiro, and P. Smyth, “From Data Mining to Knowledge Discovery in Databases,” AI Magazine, vol. 17, no. 3, pp. 37–54, 1996.
[3] J. Han, M. Kamber, and J. Pei, Data Mining Concepts and Techniques, 3rd Edition. Morgan Kaufmann, 2012.
[4] P.-N. Tan, M. Steinbach, and V. Kumar, Introduction to Data Mining. Pearson, 2006.
[5] A. K. Jain, “Data clustering: 50 years beyond K-means,” Pattern Recognit. Lett., vol. 31, no. 8, pp. 651–666, Jun. 2010, doi: 10.1016/j.patrec.2009.09.011.
[6] L. Kaufman and P. J. Rousseeuw, Finding Groups in Data. Wiley, 1990. doi: 10.1002/9780470316801.
[7] Z. Huang, “Clustering Large Data Sets Mixed and Categorical Values,” Proceedings of the First Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), pp. 21–34, 1997.
[8] Z. Huang, “Extensions to the k-Means Algorithm for Clustering Large Data Sets with Categorical Values,” Data Min. Knowl. Discov., vol. 2, no. 3, pp. 283–304, Sep. 1998, doi: 10.1023/A:1009769707641.
[9] L. I. Harisandi, T. Sujaka, and R. Hammad, “Klasterisasi Pemain PUBG Mobile Dengan Algoritma K-Modes Clustering Pada Mayoung Universe,” in Seminar Nasional CORISINDO Universitas Bumigora, 2025.
[10] P. J. Rousseeuw, “Silhouettes: A graphical aid to the interpretation and validation of cluster analysis,” J. Comput. Appl. Math., vol. 20, pp. 53–65, Nov. 1987, doi: 10.1016/0377-0427(87)90125-7.