Pengelompokan Wilayah Pulau Jawa Berdasarkan Kendaraan Bermotor Menggunakan K-Medoids Dengan Euclidean Distance Tahun 2026
DOI:
https://doi.org/10.29313/bcss.v6i2.24634Keywords:
K-Medoids, Clustering, Jumlah Kendaraan Bermotor, Euclidean Distance, Pulau JawaAbstract
Abstract. Java Island recorded the highest number of registered motor vehicles in Indonesia, reaching 103,564,339 units in 2026 across 118 regencies/cities, making it necessary to identify the spatial distribution pattern of vehicle ownership to support well-targeted transportation policy. This study applies the K-Medoids algorithm with Euclidean distance to group regencies/cities in Java Island based on five motor vehicle count variables (motorcycles, passenger cars, trucks, buses, and special vehicles), with the optimal number of clusters determined using the Silhouette Coefficient. K-Medoids was chosen for its robustness to outliers compared to K-Means, since it uses an actual data object (medoid) as the cluster center. Secondary data for 2026 were obtained from the Electronic Registration and Identification (ERI) Dashboard of the Indonesian National Police Traffic Corps and were processed through descriptive analysis, outlier detection (IQR), a multicollinearity test (VIF), Z-score normalization, optimal cluster determination, and K-Medoids clustering using R software. The Silhouette evaluation produced k = 2 as the optimal number of clusters with a score of 0.7656, converging at the third iteration with a total cost of 139.871151 and forming two balanced clusters of 59 regions each. Cluster 1 (high vehicle count) includes DKI Jakarta and major cities, averaging 1,311,029 units per region, while Cluster 2 (low vehicle count) covers semi-urban and rural areas averaging 430,554 units per region.
Abstrak. Pulau Jawa merupakan wilayah dengan jumlah kendaraan bermotor tertinggi di Indonesia, mencapai 103.564.339 unit pada tahun 2026 yang tersebar di 118 Kabupaten/Kota, sehingga diperlukan analisis untuk mengidentifikasi pola persebarannya guna mendukung kebijakan transportasi yang tepat sasaran. Penelitian ini menerapkan algoritma K-Medoids dengan Euclidean Distance untuk mengelompokan wilayah di Pulau Jawa berdasarkan lima variabel jumlah kendaraan bermotor (sepeda motor, mobil penumpang, truk, bus, dan Ransus), dengan jumlah cluster optimal ditentukan menggunakan Silhouette Coefficient. K-Medoids dipilih karena lebih robust terhadap outlier dibandingkan K-Means, sebab menggunakan objek data aktual (medoid) sebagai pusat cluster. Data sekunder tahun 2026 bersumber dari Dashboard Electronic Registration and Identification (ERI) Korlantas Polri, diolah melalui analisis deskriptif, deteksi outlier (IQR), uji multikolinearitas (VIF), normalisasi Z-Score, penentuan cluster optimal, dan penerapan K-Medoids menggunakan software R. Evaluasi Silhouette Coefficient menghasilkan k = 2 sebagai jumlah cluster optimal dengan skor 0,7656, konvergen pada iterasi ke-3 dengan total cost 139,871151, menghasilkan dua cluster seimbang (59 wilayah per cluster). Cluster 1 (Jumlah Kendaraan Tinggi, medoid Klaten, Kab) mencakup DKI Jakarta dan kota-kota besar seperti Surabaya, Bandung, dan Bekasi dengan rata-rata 1.311.029 unit per wilayah, sedangkan Cluster 2 (Jumlah Kendaraan Rendah, medoid Gunung Kidul) mencakup wilayah semi-perkotaan dan perdesaan dengan rata-rata 430.554 unit per wilayah.
References
[2] Kepolisian Negara Republik Indonesia, “Data Registrasi Kendaraan Bermotor per Kabupaten/Kota Tahun 2026,” Dashboard Electronic Registration and Identification (ERI) Korlantas Polri, 2026.
[3] P. N. Tan, M. Steinbach, and V. Kumar, Introduction to Data Mining. Pearson Addison Wesley, 2006.
[4] L. Kaufman and P. J. Rousseeuw, Finding Groups in Data: An Introduction to Cluster Analysis. New York: John Wiley & Sons, 1990.
[5] R. Merdekawati and D. Kumalasari, “Penerapan algoritma K-Medoids untuk pengelompokan wilayah berdasarkan tingkat kepadatan lalu lintas,” Jurnal Statistika dan Komputasi, vol. 7, no. 1, pp. 45–58, 2024.
[6] A. Agresti and B. Finlay, Statistical Methods for the Social Sciences, 4th ed. Pearson Prentice Hall, 2009.
[7] J. F. Hair, W. C. Black, B. J. Babin, and R. E. Anderson, Multivariate Data Analysis, 8th ed. Cengage Learning, 2022.
[8] G. James, D. Witten, T. Hastie, and R. Tibshirani, An Introduction to Statistical Learning: with Applications in R, 2nd ed. Springer, 2023.
[9] P. J. Rousseeuw, “Silhouettes: A graphical aid to the interpretation and validation of cluster analysis,” Journal of Computational and Applied Mathematics, vol. 20, pp. 53–65, 1987.
[10] J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 4th ed. Morgan Kaufmann, 2022.
[11] D. C. Montgomery and G. C. Runger, Applied Statistics and Probability for Engineers, 7th ed. Wiley, 2020.