Application of Data Mining for Predicting Bank Customer Credit Eligibility Using the K-Nearest Neighbor (KNN) Algorithm
Penerapan Data Mining Untuk Prediksi Kelayakan Pemberian Kredit Nasabah Bank Menggunakan Algoritma K-Nearest Neighbor (KNN)
DOI:
https://doi.org/10.33050/sensi.v12i2.4535Keywords:
Data Mining, Knowledge Discovery in Databases (KDD), K-Nearest Neighbor (KNN), RapidMiner.Abstract
Credit granting is a core banking service that carries a high risk of non-performing loans if the customer eligibility assessment process is not conducted accurately. Manual evaluation processes tend to be time-consuming and are potentially influenced by the subjectivity of credit analysts. Therefore, a data-driven approach is needed to support decision-making that is more objective, rapid, and consistent. This study aims to apply data mining techniques using the K-Nearest Neighbor (KNN) algorithm to predict customer credit eligibility based on historical loan
application data. The research methodology follows the Knowledge Discovery in Databases (KDD) process, comprising data selection, data preprocessing, data transformation, data mining, and evaluation. The dataset used is the "Loan Prediction Dataset" from Kaggle, consisting of 614 records and 12 attributes, processed using RapidMiner Studio. The transformation stage involved label encoding for categorical attributes and normalization for numerical attributes before splitting the dataset into 80% training data and 20% testing data. The KNN model was constructed using the Euclidean Distance metric and evaluated using a confusion matrix based on accuracy, precision, and recall. The results indicate that the model achieved an accuracy of 74.80%, a precision of 75.31%, and a recall of 74.80%. These findings demonstrate that the KNearest Neighbor algorithm delivers satisfactory classification performance, making it aviable alternative for supporting objective, data-driven decision-making in the credit eligibility assessment process.
Downloads
References
J. Han, M. Kamber, and J. Pei, Data Mining: Concepts and Techniques, 3rd ed. Waltham, MA, USA: Morgan Kaufmann, 2011.
U. Fayyad, G. Piatetsky-Shapiro, and P. Smyth, "From Data Mining to Knowledge
Discovery in Databases," AI Magazine, vol. 17, no. 3, pp. 37–54, 1996.
M. J. A. Berry and G. S. Linoff, Data Mining Techniques: For Marketing, Sales, and
Customer Relationship Management, 2nd ed. Indianapolis, IN, USA: Wiley Publishing,
L. C. Thomas, Consumer Credit Models: Pricing, Profit and Portfolios. Oxford, U.K.:
Oxford University Press, 2009.
D. J. Hand and W. E. Henley, "Statistical Classification Methods in Consumer Credit
Scoring: A Review," Journal of the Royal Statistical Society: Series A (Statistics in
Society), vol. 160, no. 3, pp. 523–541, 1997.
I. Brown and C. Mues, "An Experimental Comparison of Classification Algorithms for
Imbalanced Credit Scoring Data Sets," Expert Systems with Applications, vol. 39, no. 3,
pp. 3446–3453, 2012, doi: 10.1016/j.eswa.2011.09.033.
F. Gorunescu, Data Mining: Concepts, Models and Techniques. Berlin, Germany: Springer, 2011.
S. Lessmann, B. Baesens, H. V. Seow, and L. C. Thomas, "Benchmarking State-of-the-Art Classification Algorithms for Credit Scoring: An Update of Research," European Journal of Operational Research, vol. 247, no. 1, pp. 124–136, 2015, doi:
1016/j.ejor.2015.05.030.
C. Cortes and V. Vapnik, "Support-Vector Networks," Machine Learning, vol. 20, no. 3, pp. 273–297, 1995, doi: 10.1007/BF00994018.
A. Prastyo et al., "Implementasi Algoritma Naïve Bayes pada Klasifikasi Data," Jurnal Teknologi Informasi, 2021.
A. Santoso, Kusrini, and R. Hartanto, "Perbandingan Algoritma Random Forest dan KNearest Neighbor pada Prediksi Kelayakan Kredit," Jurnal Informatika, 2024.
Harlina, "Penerapan K-Nearest Neighbor dengan Forward Selection untuk Prediksi
Kelayakan Kredit," Jurnal Teknologi Informasi, 2018.
M. Faqih, R. Prayoga, M. Dzakiir, and D. Lestari, "Klasifikasi Pengajuan Kartu Kredit
Menggunakan Algoritma K-Nearest Neighbor (KNN)," Jurnal Informatika, 2024.
M. T. Nawawi and A. Suhendar, "Implementation of MSME Credit Loan Determination Using Machine Learning Technology with KNN (K-Nearest Neighbors) Algorithm," JIKO (Jurnal Informatika dan Komputer), vol. 7, no. 3, pp. 217–221, 2024, doi:
33387/jiko.v7i3.9064.
R. A. Sari and E. Simamora, "Penerapan Metode Bootstrap Aggregating K-Nearest
Neighbor dengan Analisis Credit Scoring Berdasarkan Status Pembayaran Kredit Motor," EMASAINS: Jurnal Edukasi Matematika dan Sains, vol. 14, no. 2, 2024, doi:
59672/emasains.v14i2.5019.
Kaggle, "Loan Prediction Dataset." [Online]. Available: https://www.kaggle.com/.
RapidMiner, "RapidMiner Studio Documentation."
[Online]. Available: https://docs.rapidminer.com/.
N. S. Altman, "An Introduction to Kernel and Nearest-Neighbor Nonparametric
Regression," The American Statistician, vol. 46, no. 3, pp. 175–185, 1992.
D. M. W. Powers, "Evaluation: From Precision, Recall and F-Measure to ROC,
Informedness, Markedness and Correlation," Journal of Machine Learning Technologies, vol. 2, no. 1, pp. 37–63, 2011.
A. Zanuardi and H. Suprayitno, "Penerapan Tahapan Knowledge Discovery in Databases (KDD) pada Data Mining," Jurnal Teknologi Informasi, 2018.
