<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "https://jats.nlm.nih.gov/publishing/1.3/JATS-journalpublishing1-3.dtd"><article xml:lang="en" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" dtd-version="1.3" article-type="research-article"><front><journal-meta><journal-id journal-id-type="issn">2460-0245</journal-id><journal-title-group><journal-title>Journal of the Indonesian Mathematical Society</journal-title><abbrev-journal-title>JIMS</abbrev-journal-title></journal-title-group><issn pub-type="epub">2460-0245</issn><issn pub-type="ppub">2086-8952</issn><publisher><publisher-name>IndoMS</publisher-name><publisher-loc>Indonesia</publisher-loc></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.22342/jims.v32i3.1862</article-id><article-categories><subj-group><subject>Mathematics Subject Classification</subject></subj-group></article-categories><title-group><article-title>A More Precise Elbow Method For Optimum K-means Clustering</article-title><subtitle>Metode Elbow yang Lebih Presisi untuk Clustering K-means yang Optimal</subtitle></title-group><contrib-group><contrib contrib-type="author"><name><surname>Herdiana</surname><given-names>Indra</given-names></name><address><country country="ID">Indonesia</country><email>indra.herdiana@unsoed.ac.id</email></address><xref ref-type="aff" rid="AFF-1"></xref><xref ref-type="corresp" rid="cor-0"></xref></contrib><contrib contrib-type="author"><name><surname>Kamal</surname><given-names>Mokhamad Alfin</given-names></name><address><country country="ID">Indonesia</country><email>alfin.kamal@mhs.unsoed.ac.id</email></address><xref ref-type="aff" rid="AFF-1"></xref></contrib><contrib contrib-type="author"><name><surname>Triyani</surname><given-names>Triyani</given-names></name><address><country country="ID">Indonesia</country><email>triyani@unsoed.ac.id</email></address><xref ref-type="aff" rid="AFF-1"></xref></contrib><contrib contrib-type="author"><name><surname>Estri</surname><given-names>Mutia Nur</given-names></name><address><country country="ID">Indonesia</country><email>mutia.estri@unsoed.ac.id</email></address><xref ref-type="aff" rid="AFF-1"></xref></contrib><contrib contrib-type="author"><name><surname>Renny</surname><given-names>Renny</given-names></name><address><country country="ID">Indonesia</country><email>renny@unsoed.ac.id</email></address><xref ref-type="aff" rid="AFF-1"></xref></contrib></contrib-group><contrib-group><contrib contrib-type="editor"><name><surname>Rizal</surname><given-names>Jose</given-names></name><address><email>jrizal04@unib.ac.id</email></address></contrib></contrib-group><aff id="AFF-1"><institution content-type="dept">Mathematics Study Program</institution><institution-wrap><institution>Jenderal Soedirman University</institution><institution-id institution-id-type="ror">https://ror.org/02fckb719</institution-id></institution-wrap><country country="ID">Indonesia</country></aff><author-notes><corresp id="cor-0">Corresponding author: Indra Herdiana. Email: <email>indra.herdiana@unsoed.ac.id</email></corresp></author-notes><pub-date date-type="pub" iso-8601-date="2026-09-01" publication-format="electronic"><day>01</day><month>09</month><year>2026</year></pub-date><pub-date date-type="collection" iso-8601-date="2026-09-01" publication-format="electronic"><day>01</day><month>09</month><year>2026</year></pub-date><volume>32</volume><issue>3</issue><issue-title>SEPTEMBER</issue-title><fpage>1</fpage><lpage>21</lpage><history><date date-type="received" iso-8601-date="2024-11-26"><day>26</day><month>11</month><year>2024</year></date><date date-type="accepted" iso-8601-date="2026-03-07"><day>07</day><month>03</month><year>2026</year></date></history><permissions><copyright-statement>Copyright (c) 2026 Journal of the Indonesian Mathematical Society</copyright-statement><copyright-year>2026</copyright-year><copyright-holder>Journal of the Indonesian Mathematical Society</copyright-holder><license xlink:href="https://creativecommons.org/licenses/by-nc-nd/4.0/"><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by-nc-nd/4.0/</ali:license_ref><license-p>This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.</license-p></license></permissions><self-uri xlink:href="https://jims-a.org/index.php/jimsa/article/view/1862" xlink:title="1862"></self-uri><abstract><p>K-means clustering is an unsupervised clustering method that requires an initial decision of number of clusters. One method to determine the number of clusters is the elbow method, a heuristic method that relies on a graphical plot of the within-cluster sum of squares against the number of clusters. The method uses the number based on the elbow point, the point closest to 90 • that indicates the most optimum number of clusters. This research improves the elbow method such that the selection of cluster numbers is unbiased towards visual interpretation. We use the analytical geometric formula to calculate an angle between lines and real analysis principle of derivative to simplify the elbow point determination. We also consider every possibility of the elbow method graph behaviour such that the algorithm is universally applicable. The result is that the elbow point can be measured precisely with a simple algorithm that does not involve complex functions or calculations. This improved method gives an alternative of more reliable cluster determination method that contributes to more optimum k-means clustering, obtained by the fewest clusters with the lowest error.</p></abstract><kwd-group><kwd>k-means clustering</kwd><kwd>elbow method</kwd><kwd>angle between lines</kwd></kwd-group><funding-group><funding-statement>The research is funded by Institute of Research and Community Service (LPPM), Universitas Jenderal Soedirman, under the contract number 26.713 /UN23.35.5/PT.01/II/2024.</funding-statement></funding-group><custom-meta-group><custom-meta><meta-name>File created by JATS Editor</meta-name><meta-value>https://jatseditor.com</meta-value></custom-meta><custom-meta><meta-name>issue-created-year</meta-name><meta-value>2026</meta-value></custom-meta></custom-meta-group></article-meta></front><body><sec id="sec-1"><title>1. INTRODUCTION</title><p>Data are valuable assets in this digital era, inseparable from peoples’ lives. Every decision making process involves data. However, data alone cannot produce valuable information without proper management and interpretation. According to <xref ref-type="bibr" rid="BIBR-1">[1]</xref>, profits are made by data management, not data possession, because data do not have an intrinsic value. Statistics is a branch of mathematics that deals with collection, management, analysis, interpretation, and visualisation of data. It is the main key to data management, providing mathematical and scientific method to managing data. For example, <xref ref-type="bibr" rid="BIBR-2">[2]</xref> and <xref ref-type="bibr" rid="BIBR-3">[3]</xref> develop a data management system for demography and primary health facilities respectively that positively impact the society. These two are just a tiny fraction of the essential roles of statistics for human lives.</p><p>One of widely used data management method is clustering. Clustering is a method of grouping data to several clusters such that data points in the same cluster have a maximum similarity, while data points in diferent clusters have a maximum diference. This method is used in many sectors, such as manufacture, transport, sustainable energy, health, and public policy, since it makes data, that are not necessarily meaningful, become meaningful (see <xref ref-type="bibr" rid="BIBR-4">[4]</xref><xref ref-type="bibr" rid="BIBR-5">[5]</xref><xref ref-type="bibr" rid="BIBR-6">[6]</xref>. Despite the importance, no clustering methods that fit all <xref ref-type="bibr" rid="BIBR-7">[7]</xref>. Interestingly, diferent clustering methods applied to the same data set may result diferently. One of notable clustering methods is k-means clustering.</p><p>K-means clustering is a clustering method which principle is minimising the distance between every object within a cluster and the centre of the cluster called a centroid. The used distance is the Euclidean distance since the data span Euclidean distance, and the centres of clusters are means on Euclidean spaces <xref ref-type="bibr" rid="BIBR-8">[8]</xref><xref ref-type="fn" rid="fn-1">1</xref>. In other words, objects are grouped into one cluster with the nearest centroid, indicating the maximum similarity centred in the centroid. Some examples of k-means use in data management for further application can be seen in <xref ref-type="bibr" rid="BIBR-11">[11]</xref><xref ref-type="bibr" rid="BIBR-12">[12]</xref>. One factor that makes this method widely used is its simplicity <xref ref-type="bibr" rid="BIBR-13">[13]</xref>. K-means clustering uses simple calculations, such as the Euclidean distance, that are easily calculable by both manually and computationally. K-means clusering also converges fast <xref ref-type="bibr" rid="BIBR-13">[13]</xref>. However, although it converges faster than k-medoids clustering, k-means clustering needs an initial determination of number of clusters. This number must not be determined negligently.</p><p>There are several methods to determine such a number of clusters, such as the elbow method, gap statistics, the silhouette method, and canopy <xref ref-type="bibr" rid="BIBR-14">[14]</xref>. These methods have unique advantages, disadvantages, and compatibility. However, despite its simplicity, the elbow method has received criticism for its accuracy. The elbow method determines the number of clusters based on a graph that connects some number of clusters and corresponding sum of squared error (SSE) of each cluster. The optimum number of cluster is determined by the point forming an ”elbow”, the pointiest one, indicating that higher numbers of clusters do not significantly reduce the SSE. The criticism lies on the fact that this method relies on subjective visual interpretation where the elbow point is chosen according to the probably biased plot. The elbow point may be chosen inaccurately, thus afecting the clustering result. Here are some examples of the elbow method implementations.</p><fig id="figure-1"><label>Figure 1</label><caption><p>SSE plot with a clear elbow</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14286" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 1</alt-text></graphic></fig><fig id="figure-2"><label>Figure 2</label><caption><p>SSE plot with a unclear elbow</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14287" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 2</alt-text></graphic></fig><p><xref ref-type="fig" rid="figure-1">Figure 1</xref> illustrates a clear elbow point; we can immediately choose <inline-formula><tex-math id="math-1"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 3 \end{document} ]]></tex-math></inline-formula> as the optimum number of clusters since the pointiest point is apparent. However, there may be a case depicted by <xref ref-type="fig" rid="figure-2">Figure 2</xref> where the graph is almost smooth such that the pointiest point is not clearly shown. The elbow point is not clearly shown.</p><p>The issue extends beyond this point. Although the elbow points appears to be obvious as shown by <xref ref-type="fig" rid="figure-1">1</xref>, the ”obvious elbow” is possible to be ”a false elbow”. It is because the scaling of the horizontal and vertical axis are often disproportional. In the case of <xref ref-type="fig" rid="figure-1">Figure 1</xref>, the horizontal axis scale is 0 to 7, while the vertical axis scale is 0 to 1400. This distortion causes misinterpretations and lead to incorrect choices of optimum number of clusters, thus giving unwanted k-means clustering results.</p><fig id="figure-3"><label>Figure 3</label><caption><p>SSE plot with undistorted scale</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14288" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 3</alt-text></graphic></fig><fig id="figure-4"><label>Figure 4</label><caption><p>SSE plot with distorted scale</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14289" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 4</alt-text></graphic></fig><p><xref ref-type="fig" rid="figure-4">Figure 4</xref> shows a distorted plot (that unfortunately often occurs during elbow method processes) where the vertical axis scale is not 1:1 proportional with the horizontal axis scale. In contrast, <xref ref-type="fig" rid="figure-3">Figure 3</xref> shows the plot with a proportional 1:1 scaling where it indicates absolute SSE decreases. <xref ref-type="fig" rid="figure-4">Figure 4</xref> gives impression that the optimum number of cluster is 3, while <xref ref-type="fig" rid="figure-3">Figure 3</xref> contradicts it.</p><p>Due to the complex subjectivity, k-means clustering often involves another method that is considered less subjective, thus neglecting improvisation of the elbow method. Meanwhile, analytical geometry and real analysis have hidden roles in precising this heuristic method. This paper aims to improve the method by changing the subjectivity into objectivity using some principles of analytical geometry and real analysis.</p><p>Shi etc. in <xref ref-type="bibr" rid="BIBR-15">[15]</xref> provides a method of determining the elbow point precisely using a principle of geometry, namely the cosine rule of a triangle. However, this method involves a long calculation of a triangle side length represented by the distance between two adjacent points. Moreover, this methods uses an inverse trigonometry function that is not a simple standard mathematical operation like addition and multiplication. Sinaga and Yang in <xref ref-type="bibr" rid="BIBR-16">[16]</xref> create an alternative k-means clustering algorithm where it finds the number of clusters independently without using initial cluster method determination. The method’s disadvantage is similar to that of <xref ref-type="bibr" rid="BIBR-15">[15]</xref>. It involves a non-standard arithmetic operation namely natural logarithm.</p><p>The method proposed in this paper ofers an advantage over the methods in <xref ref-type="bibr" rid="BIBR-15">[15]</xref><xref ref-type="bibr" rid="BIBR-16">[16]</xref> that may give similar accuracy but computationally more expensive. It only involves standard arithmetic operations namely addition, subtraction, multiplication, and division, resulting in more eficient algorithm. Not all programming languages have built-in complex mathematical functions like trigonometric inverses and natural logarithm; even C++ needs to add the package cmath in order to use these operations <xref ref-type="bibr" rid="BIBR-17">[17]</xref>. Therefore, using standard arithmetic functions that are included in all programming languages ofer flexibility and simplicity.</p><p>The method uses the principle of measuring an angle between two lines formulated in analytical geometry. The real analysis principle of derivative is also used to optimise the algorithm such that non-standard operations occurring in the algorithm can be omitted. Furthermore, the proposed method also considers every possibility of the elbow method graph behaviour such that the elbow point is chosen more precisely. This method can be an alternative to optimise k-means clustering; results in <xref ref-type="bibr" rid="BIBR-15">[15]</xref> are even shown to be better than the silhouette method. In the research, the silhouette method exhibits inconsistency to diferent data distribution by giving diferent cluster numbers, unlike the elbow method, and is computationally more expensive.</p></sec><sec id="sec-2"><title>2. MAIN RESULTS</title><p>This section discusses the detailed formulation of determining the elbow point using an analytical geometric approach. Data simulation is also provided using Python.</p><p><bold>2.1. Exact Elbow Point Formulation.</bold> The following theories about the elbow method, k-means clustering, and SSE are cited from <xref ref-type="bibr" rid="BIBR-8">[8]</xref><xref ref-type="bibr" rid="BIBR-14">[14]</xref>.</p><p>Let X be a collection of n continuous p-dimensional data. Let X denote the i-th data point on X where <inline-formula><tex-math id="math-2"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf {X} _ {i} = (x _ {1}, x _ {2}, \ldots , x _ {p}) \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-3"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle i = 1 , 2 , \dots , n \end{document} ]]></tex-math></inline-formula> . Assume that X does not contain an outlier such that it is suitable for k-means clustering. Suppose that we cluster the data set into k clusters, and the determination of k utilises the elbow method.</p><p>We initiate the elbow method by plotting the sequence <inline-formula><tex-math id="math-4"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> for <inline-formula><tex-math id="math-5"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = \end{document} ]]></tex-math></inline-formula><inline-formula><tex-math id="math-6"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 1 , 2 , \ldots , n ^ {} \end{document} ]]></tex-math></inline-formula><xref ref-type="fn" rid="fn-2">2 </xref>where <inline-formula><tex-math id="math-7"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k ) \end{document} ]]></tex-math></inline-formula> is the sum of squared errors obtained if the data set is grouped into k clusters by k-means clustering. It is defined by</p><disp-formula id="equation-1"><tex-math id="math-8"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E (k) = \sum_ {j = 1} ^ {k} \sum_ {\mathbf {X} _ {i} \in S _ {j}} \| \mathbf {X} _ {i} - \mathbf {C} _ {S _ {j}} \| ^ {2} \end{document} ]]></tex-math></disp-formula><p>where <inline-formula><tex-math id="math-9"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S _ { j } \end{document} ]]></tex-math></inline-formula> is the j-th cluster, <inline-formula><tex-math id="math-10"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { C } _ { S _ { j } } \end{document} ]]></tex-math></inline-formula> is the centroid of cluster <inline-formula><tex-math id="math-11"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S _ { j } \end{document} ]]></tex-math></inline-formula> , and <inline-formula><tex-math id="math-12"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \| \mathbf { X } _ { i } - \mathbf { C } _ { S _ { j } } \| \end{document} ]]></tex-math></inline-formula> is the Euclidean distance between the data point <inline-formula><tex-math id="math-13"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { X } _ { i } \end{document} ]]></tex-math></inline-formula> and its centroid <inline-formula><tex-math id="math-14"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf { C } _ { S _ { j } } \end{document} ]]></tex-math></inline-formula> . The SSE plotted here is obtained after iterating the k-means clustering such that no data points change their cluster membership. In other words, the plotted SSE is the most optimum SSE indicated by complete k-means clustering, not the initial SSE calculated when centroids are first-chosen randomly.</p><p>K-means clustering groups data points based on the nearest centroid using the Euclidean distance, and centroids change repeatedly by the optimising formula</p><disp-formula id="equation-2"><tex-math id="math-15"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathbf {C} _ {S _ {j}} = \frac {1}{N (S _ {j})} \sum_ {\mathbf {X} _ {i} \in S _ {j}} \mathbf {X} _ {i}, \end{document} ]]></tex-math></disp-formula><p>where <inline-formula><tex-math id="math-16"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle N ( S _ { j } ) \end{document} ]]></tex-math></inline-formula> denotes the number of data points in the cluster <inline-formula><tex-math id="math-17"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S _ { j } \end{document} ]]></tex-math></inline-formula> . The change stops when no data points move to another cluster, thus implying the lowest possible SSE. If the number of clusters increases, the number of centroids also increases. Hence, distances between data points and their nearest centroids tend to shrink, resulting in a smaller SSE. <inline-formula><tex-math id="math-18"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \mathrm { B y } \end{document} ]]></tex-math></inline-formula> this deduction, we conclude that the sequence <inline-formula><tex-math id="math-19"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> is monotonically decreasing as greater clusters imply smaller SSEs.</p><fig id="figure-5"><label>Figure 5</label><caption><p>Initial SSE plotting</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14290" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 5</alt-text></graphic></fig><p>The sequence SSE(k), as illustrated by <xref ref-type="fig" rid="figure-5">Figure 5</xref>, is plotted monotonically decreasing. We connect adjacent points by a straight line such that all point (terms of <inline-formula><tex-math id="math-20"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> form a continuous function consisting of several straight lines. Every corner of the function graph forms an angle that is used to determine the elbow point.</p><p>The elbow point is the point in which there is no significant drop of SSE afterwards. The typical elbow method determines the angle closest to <inline-formula><tex-math id="math-21"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ° \end{document} ]]></tex-math></inline-formula> since it indicates the last point with a significant drop of SSE i.e. the graph starts flattening beyond the point. However, since the elbow method, at the beginning, is a heuristic method, the closest-to-90<sup>◦</sup> choice is rather ambiguous. Two disputes need to be addressed. First, the angle measurement is not mathematical. Second, the method appears to ignore the possibility of facing-downwards corners chosen as the elbow point (illustrated by <xref ref-type="fig" rid="figure-6">Figure 6</xref>). The point may indicate the closest one to 90◦ while there is still a significant SSE drop afterwards.</p><fig id="figure-6"><label>Figure 6</label><caption><p>SSE plot with a facing-downwards corner</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14291" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 6</alt-text></graphic></fig><p>To address the first dispute, we construct a formula to measure the exact angle of the corners. We use the formula of angle between lines that uses line slopes.<target id="anchor-1" target-type="reference-target"/></p><p><bold>Theorem 2.1.</bold><italic>Let </italic><inline-formula><tex-math id="math-22"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { 1 } \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-23"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { 2 } \end{document} ]]></tex-math></inline-formula><italic> be lines with inclination </italic><inline-formula><tex-math id="math-24"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \theta _ { 1 } \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-25"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \theta _ { 2 } \end{document} ]]></tex-math></inline-formula><italic> respectively. Let </italic><inline-formula><tex-math id="math-26"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { 1 } } \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-27"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { 2 } } \end{document} ]]></tex-math></inline-formula><italic> denote the slope of </italic><inline-formula><tex-math id="math-28"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { 1 } \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-29"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { 2 } \end{document} ]]></tex-math></inline-formula><italic> respectively where </italic><inline-formula><tex-math id="math-30"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { 1 } } ~ = ~ \tan \theta _ { 1 } \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-31"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { 2 } } = \tan \theta _ { 2 } \end{document} ]]></tex-math></inline-formula><italic> . The two lines intersect in a point, forming two supplementary angles, </italic><inline-formula><tex-math id="math-32"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \phi \end{document} ]]></tex-math></inline-formula><italic> and ψ as illustrated by </italic><xref ref-type="fig" rid="figure-7">7</xref><italic>. The following holds:</italic></p><p><italic></italic><inline-formula><tex-math id="math-33"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \tan \phi = \frac {m _ {l _ {1}} - m _ {l _ {2}}}{1 + m _ {l _ {1}} m _ {l _ {2}}} \end{document} ]]></tex-math></inline-formula><italic></italic></p><p><italic>and</italic></p><disp-formula id="equation-3"><tex-math id="math-34"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \tan \psi = - \tan \phi . \end{document} ]]></tex-math></disp-formula><p>The <italic>proof</italic> is constructed in <xref ref-type="bibr" rid="BIBR-18">[18]</xref> using the definition of slope, relation among interior and exterior angles of a triangle, and the property of tangent of sum of angles.</p><fig id="figure-7"><label>Figure 7</label><caption><p>Formula of angles between two lines</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14292" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 7</alt-text></graphic></fig><p>Based on <xref ref-type="fig" rid="figure-7">7</xref>, <inline-formula><tex-math id="math-35"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \phi \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-36"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi \end{document} ]]></tex-math></inline-formula> are supplementary. Therefore, unless both are <inline-formula><tex-math id="math-37"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ° \end{document} ]]></tex-math></inline-formula> the former is located in the first quadrant (acute), and the latter is located in the second quadrant (obtuse), or vice versa. The angle located in the first quadrant has a positive tangent, whereas the one located in the second quadrant has a negative tangent.</p><p>We use the modified Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-1">2.1</xref> to construct angle measurement formula of the SSE graph since the line behaviours are diferent. However, in this case, we first assume that the second dispute does not exist i.e. the lines forming the graph get more flattened after each term. The flattening condition is indicated by flattening lines i.e. slopes of the lines get closer to 0 as k increases (the slope of horizontal lines). In other words, we assume that <inline-formula><tex-math id="math-38"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ {l _ {k}} > m _ {l _ {k - 1}} \end{document} ]]></tex-math></inline-formula> for <inline-formula><tex-math id="math-39"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 1 , 2 , . . . , n . \end{document} ]]></tex-math></inline-formula></p><fig id="figure-8"><label>Figure 8</label><caption><p>Implementation of Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-1">2.1</xref> on SSE plotting</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14293" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 8</alt-text></graphic></fig><p><target id="anchor-2" target-type="reference-target"/></p><p><bold>Theorem 2.2.</bold> Let <inline-formula><tex-math id="math-40"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k } \end{document} ]]></tex-math></inline-formula> denote the straight line connecting the point <inline-formula><tex-math id="math-41"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> ) and <inline-formula><tex-math id="math-42"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k + 1 , S S E ( k + 1 ) ) \end{document} ]]></tex-math></inline-formula> , and <inline-formula><tex-math id="math-43"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi _ { k } \end{document} ]]></tex-math></inline-formula> denote the angle of the corner facing upwards. The elbow point is <inline-formula><tex-math id="math-44"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> such that it satisfies</p><disp-formula id="equation-4"><tex-math id="math-45"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \min \left(\tan \psi_ {k} | k = 2, 3,..., n - 1\right)\tag{1} \end{document} ]]></tex-math></disp-formula><p>where</p><disp-formula id="equation-5"><tex-math id="math-46"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \tan (\psi_ {k}) = \frac {- S S E (k + 1) + 2 S S E (k) - S S E (k - 1)}{1 + (S S E (k) - S S E (k - 1)) (S S E (k + 1) - S S E (k))}. \end{document} ]]></tex-math></disp-formula><p>To prove Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-2">2.2</xref>, we first consider the following proposition to prove that <inline-formula><tex-math id="math-47"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi _ { k } \end{document} ]]></tex-math></inline-formula> is always between 90 ◦ ∗and <inline-formula><tex-math id="math-48"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle * 1 8 0 \mathrm { ~ o ~ } * \end{document} ]]></tex-math></inline-formula> . This fact ensures that the angle closest to <inline-formula><tex-math id="math-49"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 \circ * \end{document} ]]></tex-math></inline-formula> has the smallest tangent.<target id="anchor-3" target-type="reference-target"/></p><p><bold>Proposition 2.3.</bold><italic>If ψ is the angle that faces upward, we have </italic><inline-formula><tex-math id="math-50"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ^ { \circ } < \psi < 1 8 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula></p><p><italic>Proof. If </italic><inline-formula><tex-math id="math-51"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k } \end{document} ]]></tex-math></inline-formula><italic> denotes the straight line connecting the point </italic><inline-formula><tex-math id="math-52"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-53"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k + \end{document} ]]></tex-math></inline-formula><italic></italic><inline-formula><tex-math id="math-54"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 1 , S S E ( k + 1 ) ) \end{document} ]]></tex-math></inline-formula> , <italic>since </italic><inline-formula><tex-math id="math-55"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k + 1 ) < S S E ( k ) \end{document} ]]></tex-math></inline-formula><italic> for </italic><inline-formula><tex-math id="math-56"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 1 , 2 , . . . , n \end{document} ]]></tex-math></inline-formula><italic> ,</italic> we have</p><disp-formula id="equation-6"><tex-math id="math-57"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array}{l} m _ {l _ {k}} = \frac {S S E (k + 1) - S S E (k)}{(k + 1) - (k)} \\ \qquad = S S E (k + 1) - S S E (k) \\ \qquad < 0. \end{array} \end{document} ]]></tex-math></disp-formula><p>Moreover, we assume <inline-formula><tex-math id="math-58"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { k } } > m _ { l _ { k } } \end{document} ]]></tex-math></inline-formula> which implies <inline-formula><tex-math id="math-59"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { k } } - m _ { l _ { k - 1 } } > 0 \end{document} ]]></tex-math></inline-formula> . Consequently, since <inline-formula><tex-math id="math-60"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { k } } m _ { l _ { k - 1 } } > 0 \end{document} ]]></tex-math></inline-formula> , we have</p><disp-formula id="equation-7"><tex-math id="math-61"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array}{c} \tan \phi_ {k} = \frac {m _ {l _ {k}} - m _ {l _ {k - 1}}}{1 + m _ {l _ {k}} m _ {l _ {k - 1}}} \\ > 0 \end{array} \end{document} ]]></tex-math></disp-formula><p>and</p><disp-formula id="equation-8"><tex-math id="math-62"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array}{c} \tan \psi = - \tan \phi \\ < 0. \end{array} \end{document} ]]></tex-math></disp-formula><p>Since tan <inline-formula><tex-math id="math-63"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi < 0 \end{document} ]]></tex-math></inline-formula> , the angle <inline-formula><tex-math id="math-64"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi \end{document} ]]></tex-math></inline-formula> must be in quadrant II or III i.e. <inline-formula><tex-math id="math-65"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ^ { \circ } < \psi < 1 8 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula> or <inline-formula><tex-math id="math-66"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 1 8 0 ^ { \circ } < \psi < 2 7 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula> . However, since <italic>ψ</italic> and ϕ are supplementary, it must not exceed 180<sup>◦</sup>. Therefore, we have <inline-formula><tex-math id="math-67"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ^ { \circ } < \psi < 1 8 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula> □</p><p>Here is the <italic>proof</italic> of Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-2">2.2.</xref></p><p><italic>Proof</italic>. Recall that <inline-formula><tex-math id="math-68"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k } \end{document} ]]></tex-math></inline-formula> denote the straight line connecting the point <inline-formula><tex-math id="math-69"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-70"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k { + } 1 , S S E ( k { + } 1 ) ) \end{document} ]]></tex-math></inline-formula> ). As illustrated by <xref ref-type="fig" rid="figure-8">Figure 8</xref> we have <inline-formula><tex-math id="math-71"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \theta _ { 1 } \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-72"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \theta _ { 2 } \end{document} ]]></tex-math></inline-formula> are the inclination of <inline-formula><tex-math id="math-73"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k - 1 } \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-74"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k } \end{document} ]]></tex-math></inline-formula> respectively. Since <inline-formula><tex-math id="math-75"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi + \left( 1 8 0 ^ { \circ } - \theta _ { 1 } \right) + \theta _ { 2 } = 1 8 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula> , we have</p><disp-formula id="equation-9"><tex-math id="math-76"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array}{l} \tan \phi = \tan \left(1 8 0 ^ {\circ} - ((1 8 0 ^ {\circ} - \theta_ {1}) + \theta_ {2})\right) \\ \qquad = \tan (\theta_ {1} - \theta_ {2}) \\ \qquad = \frac {\tan \theta_ {1} - \tan \theta_ {2}}{1 + \tan \theta_ {1} \tan \theta_ {2}} \\ \qquad = \frac {m _ {l _ {k}} - m _ {l _ {k - 1}}}{1 + m _ {l _ {k}} m _ {l _ {k - 1}}}. \end{array} \end{document} ]]></tex-math></disp-formula><p>Elbow point candidates are represented by the angle supplementary to <inline-formula><tex-math id="math-77"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \phi , \end{document} ]]></tex-math></inline-formula> that is <inline-formula><tex-math id="math-78"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi . \end{document} ]]></tex-math></inline-formula> . Since the angles are supplementary, the angles have opposites tangents i.e.</p><disp-formula id="equation-10"><tex-math id="math-79"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array}{l} \tan \psi = - \tan \phi \\ \iff \tan \psi = - \frac {m _ {l _ {k}} - m _ {l _ {k - 1}}}{1 + m _ {l _ {k}} m _ {l _ {k - 1}}} \\ \iff \tan \psi = \frac {m _ {l _ {k - 1}} - m _ {l _ {k}}}{1 + m _ {l _ {k}} m _ {l _ {k - 1}}}. \end{array}\tag{2} \end{document} ]]></tex-math></disp-formula><p>The endpoints, <inline-formula><tex-math id="math-80"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( 1 , S S E ( 1 ) \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-81"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( n , S S E ( n ) ) \end{document} ]]></tex-math></inline-formula> , do not form an angle. Therefore, Equation <xref ref-type="disp-formula" rid="equation-10">2</xref> only holds for <inline-formula><tex-math id="math-82"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 2 , 3 , . . . , n - 1 \end{document} ]]></tex-math></inline-formula></p><p>By the formula of slope of a straight line, since <inline-formula><tex-math id="math-83"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k } \end{document} ]]></tex-math></inline-formula> connects <inline-formula><tex-math id="math-84"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-85"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k + 1 , S S E ( k + 1 ) ) \end{document} ]]></tex-math></inline-formula> , we have <inline-formula><tex-math id="math-86"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ {l _ {k}} = \frac {S S E (k + 1) - S S E (k)}{(k + 1) - k} \end{document} ]]></tex-math></inline-formula> and</p><disp-formula id="equation-11"><tex-math id="math-87"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ {l _ {k - 1}} = \frac {S S E (k) - S S E (k - 1)}{k - (k - 1)}. \end{document} ]]></tex-math></disp-formula><p>Therefore, Equation <xref ref-type="disp-formula" rid="equation-10">2</xref> can be written as</p><disp-formula id="equation-12"><tex-math id="math-88"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array}{l} \tan \psi_ {k} = \frac {m _ {l _ {k - 1}} - m _ {l _ {k}}}{1 + m _ {l _ {k}} m _ {l _ {k - 1}}} \\ \qquad = \frac {\frac {S S E (k) - S S E (k - 1)}{(k + 1) - k} - \frac {S S E (k + 1) - S S E (k)}{k - (k - 1)}}{1 + (\frac {S S E (k + 1) - S S E (k)}{(k + 1) - k}) (\frac {S S E (k) - S S E (k - 1)}{k - (k - 1)})} \\ \qquad = \frac {S S E (k) - S S E (k - 1) - (S S E (k + 1) - S S E (k))}{1 + (S S E (k + 1) - S S E (k)) (S S E (k) - S S E (k - 1))} \\ \qquad = \frac {- S S E (k - 1) + 2 S S E (k) - S S E (k + 1)}{1 + (S S E (k + 1) - S S E (k)) (S S E (k) - S S E (k - 1))}. \end{array} \end{document} ]]></tex-math></disp-formula><p>By Proposition <xref ref-type="custom" custom-type="reference-target" rid="anchor-3">2.3</xref>, we have <inline-formula><tex-math id="math-89"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ^ { \circ } < \psi < 1 8 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula> . Furthermore, since tan <italic>ψ</italic> is monotonically increasing within the interval<xref ref-type="fn" rid="fn-3">3</xref>, the angle <italic>ψ</italic> closest to <inline-formula><tex-math id="math-90"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula> must be satisfied by the smallest (most negative) tan <italic>ψ</italic>. Therefore, we conclude that the elbow point is (k, SSE(k)) that satisfies</p><disp-formula id="equation-13"><tex-math id="math-91"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \min \left(\tan \psi_ {k} | k = 2, 3, \dots , n - 1\right). \end{document} ]]></tex-math></disp-formula><p><bold>2.2. Alternative Method.</bold> Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-2">2.2</xref> holds if the second dispute is assumed to be untrue i.e. a facing downward corner, whose angle is the closest to 90◦, is probably chosen as the elbow point despite a still significant SEE drop afterwards. Some data sets cause the elbow method graph violating the assumption. In other words, there can be a case that <inline-formula><tex-math id="math-92"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { k } } \ge m _ { l _ { k - 1 } } \end{document} ]]></tex-math></inline-formula></p><fig id="figure-9"><label>Figure 9</label><caption><p>Facing-downwards corner impacts on elbow point determination</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14294" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 9</alt-text></graphic></fig><p>In <xref ref-type="fig" rid="figure-9">Figure 9</xref>, the corner that is an intersection of <inline-formula><tex-math id="math-93"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { 1 } \end{document} ]]></tex-math></inline-formula> and <inline-formula><tex-math id="math-94"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { 2 } \end{document} ]]></tex-math></inline-formula> may be chosen if it is the closest angle to <inline-formula><tex-math id="math-95"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle 9 0 ^ { \circ } \end{document} ]]></tex-math></inline-formula> if the usual elbow point criteria is used. However, it is clear that the corner must no be the elbow point as there a larger drop of SSE onwards. It does not necessarily indicate the sign of flattening SSE drops. In this case, <inline-formula><tex-math id="math-96"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 4 \end{document} ]]></tex-math></inline-formula> is the more appropriate elbow point since after <inline-formula><tex-math id="math-97"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 3 \end{document} ]]></tex-math></inline-formula> , there is still a noticeable SSE drop.<target id="anchor-4" target-type="reference-target"/></p><p><bold>Proposition 2.4.</bold><inline-formula><tex-math id="math-98"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle I f m _ { l _ { k } } \le m _ { l _ { k - 1 } } \end{document} ]]></tex-math></inline-formula> , <italic>then </italic><inline-formula><tex-math id="math-99"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k ) - S S E ( k + 1 ) \geq S S E ( k - 1 ) - \end{document} ]]></tex-math></inline-formula><italic></italic><inline-formula><tex-math id="math-100"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k ) \end{document} ]]></tex-math></inline-formula></p><p><italic>Proof. The drop of </italic><inline-formula><tex-math id="math-101"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k ) \end{document} ]]></tex-math></inline-formula><italic> to </italic><inline-formula><tex-math id="math-102"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k + 1 ) { \mathrm { ~ i s ~ } } | S S E ( k ) - S S E ( k + 1 ) | \end{document} ]]></tex-math></inline-formula><italic> |, or for simplicity, since </italic><inline-formula><tex-math id="math-103"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k ) \end{document} ]]></tex-math></inline-formula><italic> is monotonically decreasing, we have </italic><inline-formula><tex-math id="math-104"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k ) - S S E ( k + 1 ) \end{document} ]]></tex-math></inline-formula><italic> for </italic><inline-formula><tex-math id="math-105"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 1 , 2 , . . . , n \end{document} ]]></tex-math></inline-formula></p><p>Suppose that <inline-formula><tex-math id="math-106"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { k } } \le m _ { l _ { k - 1 } } \end{document} ]]></tex-math></inline-formula> . </p><p>We then have</p><disp-formula id="equation-14"><tex-math id="math-107"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \begin{array}{l} m _ {l _ {k}} \leq m _ {l _ {k - 1}} \\ \iff \frac {S S E (k + 1) - S S E (k)}{(k + 1) - k} \leq \frac {S S E (k) - S S E (k - 1)}{k - (k - 1)} \\ \iff S S E (k + 1) - S S E (k) \leq S S E (k) - S S E (k - 1) \\ \iff - (S S E (k + 1) - S S E (k)) \geq - (S S E (k) - S S E (k - 1)) \\ \iff S S E (k) - S S E (k + 1) \geq - S S E (k - 1) - S S E (k) \end{array}\tag{3} \end{document} ]]></tex-math></disp-formula><p>The Equation<xref ref-type="disp-formula" rid="equation-14"> 3</xref> indicates a higher drop after the point <inline-formula><tex-math id="math-108"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> .<xref ref-type="fig" rid="figure-10">Figure 10</xref> illustrates the Proposition <xref ref-type="custom" custom-type="reference-target" rid="anchor-4">2.4</xref>.</p><fig id="figure-10"><label>Figure 10</label><caption><p>Further Effects of Facing-downwards corners</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14277" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 10</alt-text></graphic></fig><p><xref ref-type="fig" rid="figure-10">Figure 10</xref> shows that <inline-formula><tex-math id="math-109"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k } \end{document} ]]></tex-math></inline-formula> has a larger slope that of <inline-formula><tex-math id="math-110"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k - 1 } \end{document} ]]></tex-math></inline-formula> . It indicates a higher SSE drop after the intersection of the two lines as shown by Equation <xref ref-type="disp-formula" rid="equation-14">3</xref>. The corner of the intersection faces downwards.</p><p>By discovering the efect of ignoring the second dispute to the elbow point choice, we reconstruct Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-2">2.2</xref> such that it is more universal. We use the same criteria (the closest-to-90<sup>◦</sup>) with an additional condition: neglecting the point in Proposition <xref ref-type="custom" custom-type="reference-target" rid="anchor-4">2.4</xref>. In this case, <inline-formula><tex-math id="math-111"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 4 \end{document} ]]></tex-math></inline-formula> is the more appropriate elbow point since after <inline-formula><tex-math id="math-112"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 3 \end{document} ]]></tex-math></inline-formula> , there is still a noticeable SSE drop.<target id="anchor-5" target-type="reference-target"/></p><p><bold>Theorem 2.5.</bold><italic>Suppose that </italic><inline-formula><tex-math id="math-113"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle S S E ( k { + } 1 ) \leq S S E ( k ) f o r k = 1 , 2 , . . . , n \end{document} ]]></tex-math></inline-formula><italic> . Let </italic><inline-formula><tex-math id="math-114"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle l _ { k } \end{document} ]]></tex-math></inline-formula><italic> denote the straight line connecting the point </italic><inline-formula><tex-math id="math-115"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula><italic> ) and </italic><inline-formula><tex-math id="math-116"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k + 1 , S S E ( k + 1 ) ) \end{document} ]]></tex-math></inline-formula><italic> and </italic><inline-formula><tex-math id="math-117"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi _ { k } \end{document} ]]></tex-math></inline-formula><italic> denote the angle of the corner facing upwards. The elbow point is </italic><inline-formula><tex-math id="math-118"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula><italic> such that it satisfies </italic><inline-formula><tex-math id="math-119"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \min \left(\tan \psi_ {k} | k = 2, 3,..., n - 1; m _ {l _ {k}} > m _ {l _ {k - 1}}\right) \end{document} ]]></tex-math></inline-formula><italic> where</italic></p><disp-formula id="equation-15"><tex-math id="math-120"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \tan (\psi_ {k}) = \frac {- S S E (k + 1) + 2 S S E (k) - S S E (k - 1)}{1 + (S S E (k) - S S E (k - 1)) (S S E (k + 1) - S S E (k))}.\tag{4} \end{document} ]]></tex-math></disp-formula><p>Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-5">2.5</xref> adds a condition that is the negation of condition in Proposition <xref ref-type="custom" custom-type="reference-target" rid="anchor-4">2.4</xref> into Equation <xref ref-type="disp-formula" rid="equation-4">1</xref>. Hence, points indicating a rise of SSE drop, satisfying the condition in Proposition <xref ref-type="custom" custom-type="reference-target" rid="anchor-4">2.4</xref> is neglected.</p><p><bold>2.3. Program Implementation and Simulation.</bold> K-means clustering is often implemented by software, computationally, especially if it involves a large data set. Hence, we also provide an implementation of the constructed formula into computer programming. Here is the pseudo-code.</p><fig id="figure-11"><label>Algorithm 1</label><caption><p>Elbow Method Pseudocode</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14278" mime-subtype="png" mimetype="image"><alt-text>Algorithm 1</alt-text></graphic></fig><p>The algorithm begins with defining main variables used for k-means clustering and the elbow method namely a data set and an array containing list of tan <italic>ψ</italic>. We define some functions to simplify the algorithm that are the Euclidean distance and the SSE. The Euclidean distance function involves a looping as a representative of consecutive summation of squared diferences; the SSE function also involves a looping of Euclidean distances between every variable and its centroid.</p><p>The next step is creating a looping to put tan <italic>ψ</italic>k consecutively into the defined array. The formula is of Equation <xref ref-type="disp-formula" rid="equation-15">4</xref>. We also put a condition of Theorem <xref ref-type="custom" custom-type="reference-target" rid="anchor-5">2.5</xref> to ignore corners facing upwards. We put tan <inline-formula><tex-math id="math-121"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \psi _ { k } = 0 \end{document} ]]></tex-math></inline-formula> for this kind of points. Hence, the points will not be chosen as the elbow point since they will not become the minimum among negative tangents.</p><p>Lastly, we define the variable ”angle” as the minimum of tan <italic>ψ</italic> . The k-means clustering is then run with the number of clusters.</p><p>Some programming languages already include standard functions used in this algorithm, such as the Euclidean distance, SSE, and k-means itself. Python is one oh such programming languages. Here is the implementation in Python for some sample data set (we can change it to diferent data set).</p><fig id="figure-12"><label>Algorithm 2</label><caption><p>The Elbow Method Implementation on Python</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14279" mime-subtype="png" mimetype="image"><alt-text>Algorithm 2</alt-text></graphic></fig><fig id="figure-13"><label>Here is the output of the given implementation on Python.</label><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14280" mime-subtype="png" mimetype="image"><alt-text>Here is the output of the given implementation on Python.</alt-text></graphic></fig><fig id="figure-14"><label>Figure 11</label><caption><p>SSE plot of Python simulation (distorted)</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14281" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 11</alt-text></graphic></fig><p><xref ref-type="fig" rid="figure-14">Figure 11</xref> shows that the most optimum number of cluster is 3. However, as described in Introduction, it could be misleading. It is answered by <xref ref-type="fig" rid="figure-15">Figure 12</xref> which is the output of the Python code for tan <italic>ψ</italic> calculation.</p><fig id="figure-15"><label>Figure 12</label><caption><p>Values of tan ψk</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14282" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 12</alt-text></graphic></fig><p><xref ref-type="fig" rid="figure-15">Figure 12</xref> shows that the smallest tangent is obtained when <inline-formula><tex-math id="math-122"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 6 . \end{document} ]]></tex-math></inline-formula> . It shows that the distorted SSE plot in <xref ref-type="fig" rid="figure-14">Figure 11</xref> is indeed misleading.</p><fig id="figure-16"><label>Figure 13</label><caption><p>SSE plot of Python simulation (undistorted)</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14283" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 13</alt-text></graphic></fig><p><xref ref-type="fig" rid="figure-16">Figure 13</xref> shows the undistorted plot of SSE. Despite the quite unclear plot, we can conclude that 3 is not the most optimum number of clusters since there is still indeed a significant drop afterwards. The plot starts flattening after k = 6. The result of the k-means clustering is shown in <xref ref-type="fig" rid="figure-17">Figure 14</xref>.</p><fig id="figure-17"><label>Figure 14</label><caption><p>K-means implementation for k = 6</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14284" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 14</alt-text></graphic></fig><p>In comparison to <xref ref-type="fig" rid="figure-17">Figure 14</xref>, <xref ref-type="fig" rid="figure-18">Figure 15</xref> shows the k-means clustering if there are 3 clusters, mistakenly chosen by the biased <xref ref-type="fig" rid="figure-14">Figure 11</xref>. We see that the clustering still leaves a huge distance between data points and their corresponding centroids. <xref ref-type="fig" rid="figure-17">Figure 14</xref> shows a better clustering where the distances are minimum.</p><fig id="figure-18"><label>Figure 15</label><caption><p>Figure example</p></caption><graphic xlink:href="https://jims-a.org/index.php/jimsa/article/download/1862/576/14285" mime-subtype="jpeg" mimetype="image"><alt-text>Figure 15</alt-text></graphic></fig></sec><sec id="sec-3"><title>3. CONCLUDING REMARKS</title><p>The elbow method, known as a heuristic method, is made exact using the formula of angle between lines. Suppose that <inline-formula><tex-math id="math-123"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle ( k , S S E ( k ) ) \end{document} ]]></tex-math></inline-formula> are tuple points of number of clusters and corresponding SSE. The elbow point is the point that satisfy <inline-formula><tex-math id="math-124"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \min \left(\tan \psi_ {k} | k = 2, 3,..., n - 1; m _ {l _ {k}} > m _ {l _ {k - 1}}\right) \end{document} ]]></tex-math></inline-formula></p><p>where</p><disp-formula id="equation-16"><tex-math id="math-125"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle \tan (\psi_ {k}) = \frac {- S S E (k + 1) + 2 S S E (k) - S S E (k - 1)}{1 + (S S E (k) - S S E (k - 1)) (S S E (k + 1) - S S E (k))} \end{document} ]]></tex-math></disp-formula><p>and <inline-formula><tex-math id="math-126"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle m _ { l _ { k } } = S S E ( k + 1 ) - S S E ( k ) \end{document} ]]></tex-math></inline-formula> for <inline-formula><tex-math id="math-127"><![CDATA[ \documentclass{article} \usepackage{amsmath} \begin{document} \displaystyle k = 1 , 2 , . . . , n \end{document} ]]></tex-math></inline-formula> . The algorithm is simple such that it is easily implementable in programming. The precision provided by this method also eliminates the possibility of choosing an incorrect elbow point since it results the same regardless of the visual representation.</p></sec></body><back><sec sec-type="data-availability"><title>Data Availability Statement</title><p>Data used in the research are simulation data as an example of the main result algorithm implementation.</p></sec><sec><title>Declarations.</title><p>The authors declare that this research was conducted in absence of any commercial or financial intervention that can influence the research results.</p></sec><sec sec-type="author-contributions"><title>Author Contributions.</title><p>Indra Herdiana: conceptualisation, methodology, formal analysis, resources, writing - original draft, project administration, funding acquisition. Mokhamad Alfin Kamal: conceptualisation, methodology, software, validation, visualisation. Triyani: validation, supervision, funding acquisition. Mutia Nur Estri: data curation, visualisation, writing - review and editing. Renny: investigation, resources, writing - review and editing. All authors discussed the final results.</p></sec><ack><title>Acknowledgment.</title><p>We would like to express gratitude towards Department of Mathematics of Jenderal Soedirman University and Institute of Research and Community Service (LPPM) of Jenderal Soudirman University for supporting this research work continuously.</p></ack><ref-list><title>REFERENCES</title><ref id="BIBR-1"><element-citation publication-type="journal"><article-title>Data as oil, infrastructure or asset? three metaphors of data as economic value</article-title><source>Journal of Information, Communication and Ethics in Society</source><volume>18</volume><person-group person-group-type="author"><name><surname>Nolin</surname><given-names>J.M.</given-names></name></person-group><year>2019</year><page-range>28-43,</page-range><pub-id pub-id-type="doi">10.1108/JICES-04-2019-0044</pub-id></element-citation></ref><ref id="BIBR-2"><element-citation publication-type="journal"><article-title>Sistem informasi data penduduk pada desa bogoharjo kecamatan ngadirojo kabupaten pacitan</article-title><source>Indonesian Journal of Network &amp; Security</source><volume>2</volume><person-group person-group-type="author"><name><surname>Priyanti</surname><given-names>D.</given-names></name><name><surname>Iriani</surname><given-names>S.</given-names></name></person-group><year>2013</year><page-range>55-61,</page-range><pub-id pub-id-type="doi">10.55181/ijns.v2i4.181</pub-id></element-citation></ref><ref id="BIBR-3"><element-citation publication-type="journal"><article-title>Analisis manajemen pengelolaan data sistem informasi puskesmas di tingkat dinas kesehatan di kabupaten bondowoso</article-title><source>Jurnal Biostatistik, Kependudukan, dan Informatika Kesehatan</source><volume>2</volume><person-group person-group-type="author"><name><surname>Usada</surname><given-names>N.K.</given-names></name><name><surname>Prabawa</surname><given-names>A.</given-names></name></person-group><year>2021</year><page-range>16,</page-range><pub-id pub-id-type="doi">10.7454/bikfokes.v2i1.1020</pub-id></element-citation></ref><ref id="BIBR-4"><element-citation publication-type="journal"><article-title>The cluster analysis in the manufacturing industry with k-means method: An application for turkey</article-title><source>Eurasian Journal of Economics and Finance</source><volume>6</volume><person-group person-group-type="author"><name><surname>Karaca</surname><given-names>Z.</given-names></name></person-group><year>2018</year><page-range>1-12,</page-range><pub-id pub-id-type="doi">10.15604/ejef.2018.06.03.001</pub-id></element-citation></ref><ref id="BIBR-5"><element-citation publication-type="webpage"><article-title>Clustering algorithms in healthcare</article-title><person-group person-group-type="author"><name><surname>Negi</surname><given-names>N.</given-names></name><name><surname>Chawla</surname><given-names>G.</given-names></name></person-group><year>2021</year><pub-id pub-id-type="doi">10.1007/978-3-030-67051-1</pub-id></element-citation></ref><ref id="BIBR-6"><element-citation publication-type="conf-paper"><article-title>Farming in kediri indonesia: analysis of cluster k-means</article-title><source>IOP Conference Series: Earth and Environmental Science</source><volume>1041</volume><person-group person-group-type="author"><name><surname>Solikin</surname><given-names>N.</given-names></name><name><surname>Hartono</surname><given-names>B.</given-names></name><name><surname>Sugiono</surname><given-names>S.</given-names></name><name><surname>Linawati</surname><given-names>L.</given-names></name></person-group><year>2022</year><page-range>012015,</page-range><pub-id pub-id-type="doi">10.1088/1755-1315/1041/1/012015</pub-id></element-citation></ref><ref id="BIBR-7"><element-citation publication-type="journal"><article-title>Data clustering: application and trends</article-title><source>Artificial Intelligence Review</source><volume>56</volume><person-group person-group-type="author"><name><surname>Oyewole</surname><given-names>G.J.</given-names></name><name><surname>Thopil</surname><given-names>G.A.</given-names></name></person-group><year>2023</year><page-range>6439-6475,</page-range><pub-id pub-id-type="doi">10.1007/s10462-022-10325-y</pub-id></element-citation></ref><ref id="BIBR-8"><element-citation publication-type="book"><article-title>Cluster Analysis</article-title><person-group person-group-type="author"><name><surname>Everitt</surname><given-names>B.</given-names></name></person-group><year>2011</year><publisher-name>Wiley</publisher-name><pub-id pub-id-type="doi">10.1002/9780470977811</pub-id></element-citation></ref><ref id="BIBR-9"><element-citation publication-type="journal"><article-title>An improved k-means algorithm based on contour similarity</article-title><source>Mathematics</source><volume>12</volume><person-group person-group-type="author"><name><surname>Zhao</surname><given-names>J.</given-names></name><name><surname>Bao</surname><given-names>Y.</given-names></name><name><surname>Li</surname><given-names>D.</given-names></name><name><surname>Guan</surname><given-names>X.</given-names></name></person-group><year>2024</year><page-range>2211,</page-range><pub-id pub-id-type="doi">10.3390/math12142211</pub-id></element-citation></ref><ref id="BIBR-10"><element-citation publication-type="webpage"><article-title>A k-means procedure based on a mahalanobis type distance for clustering multivariate functional data</article-title><person-group person-group-type="author"><name><surname>Martino</surname><given-names>A.</given-names></name><name><surname>Ghiglietti</surname><given-names>A.</given-names></name><name><surname>Ieva</surname><given-names>F.</given-names></name><name><surname>Paganoni</surname><given-names>A.M.</given-names></name></person-group><comment>arXiv preprint, 8 2017.</comment><pub-id pub-id-type="doi">10.48550/arXiv.1708.00386</pub-id></element-citation></ref><ref id="BIBR-11"><element-citation publication-type="journal"><article-title>Penerapan metode algoritma k-means clustering untuk pemetaan penyebaran penyakit demam berdarah dengue (dbd</article-title><source>JOURNAL OF SCIENCE AND SOCIAL RESEARCH</source><volume>4</volume><person-group person-group-type="author"><name><surname>Sembiring</surname><given-names>M.A.</given-names></name></person-group><year>2021</year><page-range>336,</page-range><pub-id pub-id-type="doi">10.54314/jssr.v4i3.712</pub-id></element-citation></ref><ref id="BIBR-12"><element-citation publication-type="journal"><article-title>Poverty level grouping in west java province with the k-means clustering method</article-title><source>Journal Of Data Insights</source><volume>1</volume><person-group person-group-type="author"><name><surname>Amelia</surname><given-names>A.</given-names></name><name><surname>Nur</surname><given-names>I.M.</given-names></name><name><surname>Rizky</surname><given-names>M.</given-names></name><name><surname>Milasari</surname><given-names>S.P.</given-names></name></person-group><year>2023</year><page-range>51-61,</page-range><pub-id pub-id-type="doi">10.26714/jodi.v1i2.152</pub-id></element-citation></ref><ref id="BIBR-13"><element-citation publication-type="journal"><article-title>The parallel implementation and application of an improved k-means algorithm</article-title><source>J. Univ. Electron. Sci. Technol</source><volume>46</volume><person-group person-group-type="author"><name><surname>Li</surname><given-names>X.</given-names></name><name><surname>Yu</surname><given-names>L.</given-names></name><name><surname>Hang</surname><given-names>L.</given-names></name><name><surname>Tang</surname><given-names>X.</given-names></name></person-group><year>2017</year><page-range>61-68,</page-range><pub-id pub-id-type="doi">10.3969/j.issn.1001-0548.2017.01.010</pub-id></element-citation></ref><ref id="BIBR-14"><element-citation publication-type="journal"><article-title>Research on k-value selection method of k-means clustering algorithm</article-title><source>J</source><volume>2</volume><person-group person-group-type="author"><name><surname>Yuan</surname><given-names>C.</given-names></name><name><surname>Yang</surname><given-names>H.</given-names></name></person-group><year>2019</year><page-range>226-235,</page-range><pub-id pub-id-type="doi">10.3390/j2020016</pub-id></element-citation></ref><ref id="BIBR-15"><element-citation publication-type="journal"><article-title>A quantitative discriminant method of elbow point for the optimal number of clusters in clustering algorithm</article-title><source>EURASIP Journal on Wireless Communications and Networking</source><volume>2021</volume><person-group person-group-type="author"><name><surname>Shi</surname><given-names>C.</given-names></name><name><surname>Wei</surname><given-names>B.</given-names></name><name><surname>Wei</surname><given-names>S.</given-names></name><name><surname>Wang</surname><given-names>W.</given-names></name><name><surname>Liu</surname><given-names>H.</given-names></name><name><surname>Liu</surname><given-names>J.</given-names></name></person-group><year>2021</year><page-range>31,</page-range><pub-id pub-id-type="doi">10.1186/s13638-021-01910-v</pub-id></element-citation></ref><ref id="BIBR-16"><element-citation publication-type="journal"><article-title>Unsupervised k-means clustering algorithm</article-title><source>IEEE Access</source><volume>8</volume><person-group person-group-type="author"><name><surname>Sinaga</surname><given-names>K.P.</given-names></name><name><surname>Yang</surname><given-names>M.-S.</given-names></name></person-group><year>2020</year><page-range>80716-80727,</page-range><publisher-name>// / /</publisher-name></element-citation></ref><ref id="BIBR-17"><element-citation publication-type="book"><article-title>Guide to Scientific Computing in C++</article-title><person-group person-group-type="author"><name><surname>Pitt-Francis</surname><given-names>J.</given-names></name><name><surname>Whiteley</surname><given-names>J.</given-names></name></person-group><year>2017</year><publisher-name>Springer International Publishing</publisher-name><pub-id pub-id-type="doi">10.1007/978-3-319-73132-2</pub-id></element-citation></ref><ref id="BIBR-18"><element-citation publication-type="webpage"><article-title>Analytic Geometry</article-title><person-group person-group-type="author"><name><surname>Fuller</surname><given-names>G.</given-names></name><name><surname>Tarwater</surname><given-names>D.</given-names></name></person-group><year>1992</year><publisher-name>Pearson</publisher-name><comment>id/books/about/Analytic Geometry.html?id=cZbuPwAACAAJ&amp;redir esc=y.</comment><ext-link xlink:href="https://books.google.co" ext-link-type="uri" xlink:title="Website link">Website link</ext-link></element-citation></ref></ref-list><fn-group><fn id="fn-1"><p>Some research use different distance for k-means clustering, such as Mahalanobis distance andcontour similarity with some advantages <xref ref-type="bibr" rid="BIBR-9">[9]</xref><xref ref-type="bibr" rid="BIBR-10">[10]</xref>.  However, this paper mainly focuses on improvingthe elbow method and is not majorly influenced by distance choice</p></fn><fn id="fn-2"><p>The largest possible number of clusters is the number of data points since it is irrational to have more clusters than data points.</p></fn><fn id="fn-3"><p>The derivative of tan <italic>ψ</italic> with respect to <italic>ψ</italic> is sec<xref ref-type="bibr" rid="BIBR-2">[2]</xref><italic>ψ</italic> that is always positive for 90<sup>◦</sup> &lt;<italic> ψ</italic> &lt; 180<sup>◦</sup>. 90°&lt; <italic>ψ</italic> &lt; 180° Therefore, the function increases.</p></fn></fn-group></back></article>