Klasifikasi Kompleksitas Gameplay Berbasis Struktur Kalimat pada Deskripsi Game
Original Full-Text Article
Download published version for reading and archivingAbstract
Game descriptions on digital distribution platforms play a crucial role in conveying the characteristics of gameplay to players. However, the language complexity of these descriptions varies and may influence players' understanding of the gameplay being offered. This study aims to classify gameplay complexity based on sentence structure in game descriptions using a Natural Language Processing (NLP) approach. The dataset used is the 10k Most Popular Gaming 2025 dataset obtained from Kaggle, with a focus on the game description column. The description data is grouped into three complexity classes: simple, medium, and complex, based on the linguistic characteristics of the text. The research process includes text preprocessing, sentence-structure-based linguistic feature extraction, and data balancing using the balance rank method. Classification is performed using the Logistic Regression, Random Forest Classifier, and Support Vector Machine algorithms. Evaluation results show that the Random Forest Classifier achieves the highest accuracy of 0.85, while Logistic Regression and Support Vector Machine obtain accuracies of 0.81 each. Feature analysis reveals that word count and average sentence length are the most influential features in determining gameplay complexity. Visualization using Principal Component Analysis shows a clear distribution pattern of complexity classes, although some overlap between classes remains. The results of this study demonstrate that sentence-structure-based linguistic analysis is effective in representing gameplay complexity in game descriptions.
Author Biographies
Program Studi Teknik Informatika, Fakultas Ilmu Komputer, Universitas Lancang Kuning, Kota Pekanbaru, Provinsi Riau, Indonesia.
Program Studi Teknik Informatika, Fakultas Ilmu Komputer, Universitas Lancang Kuning, Kota Pekanbaru, Provinsi Riau, Indonesia.
Program Studi Teknik Informatika, Fakultas Ilmu Komputer, Universitas Lancang Kuning, Kota Pekanbaru, Provinsi Riau, Indonesia.
Program Studi Teknik Informatika, Fakultas Ilmu Komputer, Universitas Lancang Kuning, Kota Pekanbaru, Provinsi Riau, Indonesia.
How to Cite
This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License .
- Share: You are free to copy, distribute, and transmit the work in any medium or format.
- Adapt: You are free to remix, transform, and build upon the work for any purpose, even commercially.
- Attribution: You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.
References
Total: 15 References- Aïdékon, É., Da Silva, W., & Hu, X. (2025). The scaling limit of the volume of loop O(n) quadrangulations. https://doi.org/10.55776/ESP534
- Branco, P., Torgo, L., & Ribeiro, R. (2015). A survey of predictive modelling under imbalanced distributions (pp. 1–48). Retrieved from http://arxiv.org/abs/1505.01658
- Breiman, L. (2001). Random forests. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 12343 LNCS, 503–515. https://doi.org/10.1007/978-3-030-62008-0_35
- Jolliffe, I. T., & Cadima, J. (2016). Principal component analysis: A review and recent developments. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 374(2065), 20150202. https://doi.org/10.1098/rsta.2015.0202
- Liu, F., Jin, T., & Lee, J. S. Y. (2025). Automatic readability assessment for sentences: Neural, hybrid, and large language models. In Language Resources and Evaluation (Springer Netherlands). https://doi.org/10.1007/s10579-024-09800-5