A More Precise Elbow Method For Optimum K-means Clustering

Indra Herdiana (1) , Mokhamad Alfin Kamal (2) , Triyani Triyani (3) , Mutia Nur Estri (4) , Renny Renny (5)
(1) Mathematics Study Program, Jenderal Soedirman University, Indonesia,
(2) Mathematics Study Program, Jenderal Soedirman University, Indonesia,
(3) Mathematics Study Program, Jenderal Soedirman University, Indonesia,
(4) Mathematics Study Program, Jenderal Soedirman University, Indonesia,
(5) Mathematics Study Program, Jenderal Soedirman University, Indonesia

Abstract

K-means clustering is an unsupervised clustering method that requires an initial decision of number of clusters. One method to determine the number of clusters is the elbow method, a heuristic method that relies on a graphical plot of the within-cluster sum of squares against the number of clusters. The method uses the number based on the elbow point, the point closest to $90^\circ$ that indicates the most optimum number of clusters. This research improves the elbow method such that the selection of cluster numbers is unbiased towards visual interpretation. We use the analytical geometric formula to calculate an angle between lines and real analysis principle of derivative to simplify the elbow point determination. We also consider every possibility of the elbow method graph behaviour such that the algorithm is universally applicable. The result is that the elbow point can be measured precisely with a simple algorithm that does not involve complex functions or calculations. This improved method gives an alternative of more reliable cluster determination method that contributes to more optimum k-means clustering, obtained by the fewest clusters with the lowest error.

Full text article

Generated from XML file

References

J. M. Nolin, “Data as oil, infrastructure or asset? three metaphors of data as economic value,” Journal of Information, Communication and Ethics in Society, vol. 18, pp. 28–43, 11 2019. https://doi.org/10.1108/JICES-04-2019-0044.

D. Priyanti and S. Iriani, “Sistem informasi data penduduk pada desa bogoharjo kecamatan ngadirojo kabupaten pacitan,” Indonesian Journal of Network & Security, vol. 2, pp. 55–61, 2013. http://dx.doi.org/10.55181/ijns.v2i4.181.

N. K. Usada and A. Prabawa, “Analisis manajemen pengelolaan data sistem informasi puskesmas di tingkat dinas kesehatan di kabupaten bondowoso,” Jurnal Biostatistik, Kependudukan, dan Informatika Kesehatan, vol. 2, p. 16, 11 2021. https://doi.org/10.7454/bikfokes.v2i1.1020.

Z. Karaca, “The cluster analysis in the manufacturing industry with k-means method: An application for turkey,” Eurasian Journal of Economics and Finance, vol. 6, pp. 1–12, 9 2018. https://doi.org/10.15604/ejef.2018.06.03.001.

N. Negi and G. Chawla, “Clustering algorithms in healthcare,” 2021. https://doi.org/10.1007/978-3-030-67051-1_13.

N. Solikin, B. Hartono, Sugiono, and Linawati, “Farming in kediri indonesia: analysis of cluster k-means,” IOP Conference Series: Earth and Environmental Science, vol. 1041, p. 012015, 6 2022. https://doi.org/10.1088/1755-1315/1041/1/012015.

G. J. Oyewole and G. A. Thopil, “Data clustering: application and trends,” Artificial Intelligence Review, vol. 56, pp. 6439–6475, 7 2023. https://doi.org/10.1007/s10462-022-10325-y.

B. Everitt, Cluster Analysis. Wiley, 2011. https://doi.org/10.1002/9780470977811.

J. Zhao, Y. Bao, D. Li, and X. Guan, “An improved k-means algorithm based on contour similarity,” Mathematics, vol. 12, p. 2211, 7 2024. https://doi.org/10.3390/math12142211.

A. Martino, A. Ghiglietti, F. Ieva, and A. M. Paganoni, “A k-means procedure based on a mahalanobis type distance for clustering multivariate functional data,” arXiv preprint, 8 2017. https://doi.org/10.48550/arXiv.1708.00386.

M. A. Sembiring, “Penerapan metode algoritma k-means clustering untuk pemetaan penyebaran penyakit demam berdarah dengue (dbd),” JOURNAL OF SCIENCE AND SOCIAL RESEARCH, vol. 4, p. 336, 10 2021. https://doi.org/10.54314/jssr.v4i3.712.

Amelia, I. M. Nur, M. Rizky, and S. P. Milasari, “Poverty level grouping in west java province with the k-means clustering method,” Journal Of Data Insights, vol. 1, pp. 51–61, 12 2023. https://doi.org/10.26714/jodi.v1i2.152.

X. Li, L. Yu, L. Hang, and X. Tang, “The parallel implementation and application of an improved k-means algorithm,” J. Univ. Electron. Sci. Technol, vol. 46, pp. 61–68, 2017. https://doi.org/10.3969/j.issn.1001-0548.2017.01.010.

C. Yuan and H. Yang, “Research on k-value selection method of k-means clustering algorithm,” J, vol. 2, pp. 226–235, 6 2019. https://doi.org/10.3390/j2020016.

C. Shi, B. Wei, S. Wei, W. Wang, H. Liu, and J. Liu, “A quantitative discriminant method of elbow point for the optimal number of clusters in clustering algorithm,” EURASIP Journal on Wireless Communications and Networking, vol. 2021, p. 31, 12 2021. https://doi.org/10.1186/s13638-021-01910-w.

K. P. Sinaga and M.-S. Yang, “Unsupervised k-means clustering algorithm,” IEEE Access, vol. 8, pp. 80716–80727, 2020. https://doi.org/10.1109/ACCESS.2020.2988796.

J. Pitt-Francis and J. Whiteley, Guide to Scientific Computing in C++. Springer International Publishing, 2017. https://doi.org/10.1007/978-3-319-73132-2.

G. Fuller and D. Tarwater, Analytic Geometry. Pearson, 1992. https://books.google.co.id/books/about/Analytic_Geometry.html?id=cZbuPwAACAAJ&redir_esc=y

Authors

Indra Herdiana
indra.herdiana@unsoed.ac.id (Primary Contact)
Mokhamad Alfin Kamal
Triyani Triyani
Mutia Nur Estri
Renny Renny
Herdiana, I., Kamal, M. A., Triyani, T., Estri, M. N., & Renny, R. (2026). A More Precise Elbow Method For Optimum K-means Clustering. Journal of the Indonesian Mathematical Society, 32(3), 1862. https://doi.org/10.22342/jims.v32i3.1862

Article Details