Feldi, Hieronimus (2026) Prediksi Kanker Ovarium Menggunakan Random Forest dan Analisis Explainable AI dengan SHAP. S1 Teknik Informatika thesis, STMIK Widya Cipta Dharma.
|
Text
2243060-S1-Jurnal.pdf Download (410kB) |
|
|
Text
2243060-S1-Teknik Informatika.pdf Restricted to Repository staff only Download (1MB) | Request a copy |
Abstract
Kanker ovarium memiliki tingkat mortalitas yang tinggi akibat keterlambatan deteksi dini sehingga diperlukan model prediksi yang akurat dan mudah diinterpretasikan untuk mendukung pengambilan keputusan. Penelitian ini bertujuan mengembangkan model prediksi dini risiko kanker ovarium menggunakan algoritma Random Forest serta menganalisis interpretabilitas hasil prediksi melalui pendekatan Explainable Artificial Intelligence (XAI) dengan metode SHapley Additive exPlanations (SHAP). Penelitian ini menggunakan metodologi Cross Industry Standard Process for Data Mining (CRISP-DM) yang meliputi tahapan Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, dan Deployment. Dataset yang digunakan merupakan data medis terbuka yang mencakup informasi demografis, klinis, dan biomarker molekuler. Tahap prapengolahan data meliputi filtering, cleaning, encoding, transformasi data, winsorization, normalisasi, data splitting, serta penanganan ketidakseimbangan data menggunakan Synthetic Minority Over-sampling Technique (SMOTE). Model dibangun menggunakan algoritma Random Forest, kemudian dievaluasi menggunakan Confusion Matrix melalui metrik accuracy, precision, recall, dan F1-score. Selanjutnya, metode SHAP digunakan untuk menjelaskan kontribusi setiap fitur terhadap hasil prediksi sehingga model menjadi lebih transparan dan mudah dipahami. Hasil evaluasi pada data latih menunjukkan bahwa model Random Forest memperoleh nilai accuracy sebesar 0,86, precision sebesar 0,87, recall sebesar 0,86, dan F1-score sebesar 0,86. Analisis Feature Importance dan SHAP menunjukkan bahwa fitur Parity, miRNA, DNAMethylation, CA125, BMI, GeneExpression, dan Age merupakan faktor yang paling berpengaruh terhadap hasil prediksi. Hasil penelitian menunjukkan bahwa model Random Forest yang dibangun mampu diimplementasikan dalam aplikasi web serta menghasilkan prediksi risiko kanker ovarium disertai penjelasan kontribusi setiap fitur melalui metode SHAP. Penelitian ini berpotensi membantu dapat menjadi referensi bagi penelitian selanjutnya dalam pengembangan metode prediksi kanker ovarium berbasis machine learning. =========================================================== Ovarian cancer has a high mortality rate due to delayed early detection, highlighting the need for an accurate and interpretable prediction model to support decision-making. This study aims to develop an early prediction model for ovarian cancer risk using the Random Forest algorithm and to analyse the interpretability of prediction results through the Explainable Artificial Intelligence (XAI) approach using the SHapley Additive exPlanations (SHAP) method. This study employed the Cross-Industry Standard Process for Data Mining (CRISP-DM) methodology, which consists of the Business Understanding, Data Understanding, Data Preparation, Modeling, Evaluation, and Deployment phases. The dataset used was an open medical dataset containing demographic, clinical, and molecular biomarker information. The data preprocessing stage included filtering, cleaning, encoding, data transformation, winsorisation, normalisation, data splitting, and handling class imbalance using the Synthetic Minority Over-sampling Technique (SMOTE). The prediction model was developed using the Random Forest algorithm and evaluated using a Confusion Matrix based on accuracy, precision, recall, and F1-score metrics. Furthermore, the SHAP method was applied to explain the contribution of each feature to the prediction results, making the model more transparent and easier to interpret. The evaluation results on the training data showed that the Random Forest model achieved an accuracy of 0.86, a precision of 0.87, a recall of 0.86, and an F1-score of 0.86. Feature Importance and SHAP analyses indicated that Parity, miRNA, DNA Methylation, CA125, BMI, Gene Expression, and Age were the most influential features in the prediction process. The results demonstrate that the developed Random Forest model was successfully implemented in a web-based application capable of predicting ovarian cancer risk while providing feature contribution explanations through the SHAP method. This study may serve as a reference for future research on the development of machine learning-based ovarian cancer prediction methods.
| Item Type: | Thesis (S1 Teknik Informatika) |
|---|---|
| Additional Information: | Pembimbing 1 : Wahyuni, S.Kom., M.Kom Pembimbing 2 : Pitrasacha Adytia, S.T., M.T |
| Uncontrolled Keywords: | Kanker Ovarium, Random Forest, Explainable Artificial Intelligence(XAI), SHAP, Machine Learning. |
| Subjects: | Q Science > QA Mathematics > QA76 Computer software |
| Divisions: | Teknik Informatika |
| Depositing User: | Mr Hieronimus Feldi |
| Date Deposited: | 07 Aug 2026 02:44 |
| Last Modified: | 07 Aug 2026 02:44 |
| URI: | http://repository.wicida.ac.id/id/eprint/6418 |
Actions (login required)
![]() |
View Item |
