Pemodelan Topik Skripsi Statistika Universitas Islam Bandung Tahun 2020–2025 Menggunakan BERTopic KernelPCA–K-Means

Authors

  • Muhammad Rivaldi Yusuf Prodi Statistika, Fakultas Farmasi dan Sains, Universitas Islam Bandung, Indonesia
  • Ilham Faishal Mahdy Prodi Statistika, Fakultas Farmasi dan Sains, Universitas Islam Bandung, Indonesia

DOI:

https://doi.org/10.29313/bcss.v6i2.26152

Keywords:

IndoBERT, KernelPCA, K-Means

Abstract

Abstract. This study models and examines the development of undergraduate thesis topics in the Statistics Study Program at Universitas Islam Bandung during 2020–2025 using a modular BERTopic architecture with IndoBERT, Kernel Principal Component Analysis (KernelPCA), and K-Means. The initial corpus contained 491 title-and-abstract documents. After data preparation and text preprocessing, 481 complete documents were retained for analysis. IndoBERT converted each document into a 768-dimensional contextual embedding, which was reduced to five components using KernelPCA with a radial basis function kernel. The number of clusters was determined through the Elbow Method, and K-Means produced six clusters with a within-cluster sum of squares of 8.974795. Topic representation using class-based TF-IDF indicated themes related to process and quality control; insurance, claims, and risk; regional and population data; relationship testing and intervariable modelling; time series and forecasting; and regression, parameter estimation, and probability distributions. The overall  coherence score was 0.546040. Annual distributions fluctuated across the observation period: Cluster 5 had the largest overall membership, while Clusters 2 and 4 gained relative prominence toward the end of the period. The 2025 pattern must be interpreted cautiously because fewer documents were available. The approach provides an interpretable overview of the thematic development of statistics theses.

Abstrak. Penelitian ini memodelkan dan menganalisis perkembangan topik skripsi mahasiswa Program Studi Statistika Universitas Islam Bandung tahun 2020–2025 menggunakan arsitektur BERTopic yang mengintegrasikan IndoBERT, Kernel Principal Component Analysis (KernelPCA), dan K-Means. Korpus awal terdiri atas 491 dokumen judul dan abstrak. Setelah tahap data preparation dan text preprocessing, sebanyak 481 dokumen lengkap dipertahankan untuk analisis. IndoBERT mengubah setiap dokumen menjadi embedding kontekstual berdimensi 768, kemudian KernelPCA dengan kernel radial basis function mereduksinya menjadi lima komponen. Jumlah klaster ditetapkan melalui Elbow Method dan K-Means menghasilkan enam klaster dengan nilai within-cluster sum of squares sebesar 8,974795. Representasi topik menggunakan class-based TF-IDF menunjukkan tema pengendalian proses dan kualitas; asuransi, klaim, dan risiko; data kewilayahan dan kependudukan; pengujian hubungan dan pemodelan antarvariabel; deret waktu dan peramalan; serta regresi, estimasi parameter, dan distribusi probabilitas. Nilai  coherence keseluruhan sebesar 0,546040. Perkembangan tahunan keenam klaster bersifat fluktuatif. Klaster 5 memiliki anggota terbanyak secara keseluruhan, sedangkan Klaster 2 dan Klaster 4 mengalami peningkatan dominasi relatif pada akhir periode. Pola tahun 2025 perlu ditafsirkan secara hati-hati karena jumlah dokumennya lebih sedikit. Pendekatan ini menghasilkan pemetaan tematik yang dapat diinterpretasikan dan menggambarkan perkembangan topik skripsi Statistika Universitas Islam Bandung.

References

Alghamdi, R., & Alfalqi, K. (2015). A Surv e y o f Topic Mode ling i n Text Mining. Nternational Journal of Advanced Computer Science and Applications, 6(1), 7. http://thesai.org/Downloads/Volume6No1/Paper_21-A_Survey_of_Topic_Modeling_in_Text_Mining.pdf
Blei, D. M., Ng, A. Y., & Jordan, M. T. (2002). Latent dirichlet allocation. Advances in Neural Information Processing Systems, (July).
Boyd-Graber, J., Hu, Y., & Mimno, D. (2017). Applications of Topic Models. Applications of Topic Models, XX(Xx), 1–154. https://doi.org/10.1561/9781680833096
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL HLT 2019 - 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies - Proceedings of the Conference, 1, 4171–4186. https://arxiv.org/pdf/1810.04805
Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. http://arxiv.org/abs/2203.05794
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP. COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference, 757–770. https://doi.org/10.18653/v1/2020.coling-main.66
MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. Https://Doi.Org/, 5.1, 281–298. https://projecteuclid.org/ebooks/berkeley-symposium-on-mathematical-statistics-and-probability/Proceedings-of-the-Fifth-Berkeley-Symposium-on-Mathematical-Statistics-and/chapter/Some-methods-for-classification-and-analysis-of-multivariate-observations/bsmsp/1200512992
Nursyahrina, Defit, S., & Sovia, R. (2024). Metode BERTopic dan LDA untuk Analisis Tren Penelitian Bidang Ilmu Komputer. Jurnal KomtekInfo, 11(4), 332–341. https://doi.org/10.35134/komtekinfo.v11i4.580
Ogunleye, B., Maswera, T., Hirsch, L., Gaudoin, J., & Brunsdon, T. (2023). Comparison of Topic Modelling Approaches in the Banking Context. Applied Sciences 2023, Vol. 13, Page 797, 13(2), 797. https://doi.org/10.3390/APP13020797
Pavithra, & Savitha. (2024). Topic Modeling for Evolving Textual Data Using LDA, HDP, NMF, BERTOPIC, and DTM With a Focus on Research Papers. Journal of Technology and Informatics (JoTI), 5(2), 53–63. https://doi.org/10.37802/joti.v5i2.618
Röder, M., Both, A., & Hinneburg, A. (2015). Exploring the space of topic coherence measures. WSDM 2015 - Proceedings of the 8th ACM International Conference on Web Search and Data Mining, 399–408. https://doi.org/10.1145/2684822.2685324;JOURNAL:JOURNAL:ACMCONFERENCES;PAGEGROUP:STRING:PUBLICATION
Tan, P.-Ning., Steinbach, Michael., Karpatne, Anuj., & Kumar, Vipin. (2019). Introduction to data mining. 839. https://books.google.com/books/about/Introduction_to_Data_Mining_EBook_Global.html?hl=id&id=i8AoEAAAQBAJ
Vaswani, A., Brain, G., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention Is All You Need. 1. https://arxiv.org/pdf/1706.03762
Xu, R., Wunsch, D. C., Member, S., & Wunsch, D. I. (2005). Survey of clustering algorithms. IEEE TRANSACTIONS ON NEURAL NETWORKS, 16(3). https://doi.org/10.1109/TNN.2005.845141

Published

2026-08-04