Research Article

K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan

by  Aolia Ikhwanudin, Tubagus Toifur, Muhamad Yusuf
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 125
Published: July 2026
Authors: Aolia Ikhwanudin, Tubagus Toifur, Muhamad Yusuf
10.5120/ijca44904f79e07d
PDF

Aolia Ikhwanudin, Tubagus Toifur, Muhamad Yusuf . K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan. International Journal of Computer Applications. 187, 125 (July 2026), 45-53. DOI=10.5120/ijca44904f79e07d

                        @article{ 10.5120/ijca44904f79e07d,
                        author  = { Aolia Ikhwanudin,Tubagus Toifur,Muhamad Yusuf },
                        title   = { K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 125 },
                        pages   = { 45-53 },
                        doi     = { 10.5120/ijca44904f79e07d },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Aolia Ikhwanudin
                        %A Tubagus Toifur
                        %A Muhamad Yusuf
                        %T K-Means Clustering for Regional Segmentation of Municipal Permit Services: A Case Study of Tangerang Selatan%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 125
                        %P 45-53
                        %R 10.5120/ijca44904f79e07d
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Effective allocation of public service resources in Indonesian cities requires understanding the spatial heterogeneity of permit-type demand at the sub-district level. This paper presents a multi-method unsupervised clustering framework applied to 56 sub-districts in Tangerang Selatan City, using proportion vectors for eight permit categories combined with log-transformed application volume. Four algorithms were evaluated — K-Means, Agglomerative (Ward), Gaussian Mixture Model (GMM), and Spectral Clustering — across k=2..6. Four outlier sub-districts were identified and excluded prior to clustering (kel_0, Alue Bagok, Pondok Benda, Pamulang Timur). On the remaining 52 sub-district, K-Means at k=4 (collapsed to 3 interpretable clusters) achieved the highest silhouette score (0.2456), outperforming Agglomerative (0.2347), GMM (0.1596), and Spectral (0.1401). The three final clusters represent: Cluster 0 (5 sub-districts) — cemetery/land-use permit specialists with high tariffed field-inspection demand (p_Field=0.448, mean 8,106 applications); Cluster 1 (28 sub-districts) — residential service generalists with high free-field inspection demand (p_FreeF=0.413, mean 3,644 applications); and Cluster 2 (19 sub-districts) — high-volume administrative-review centers with elevated inter-agency permit activity (p_Admin=0.440, mean 7,618 applications). These findings provide a data-driven basis for differentiated surveyor allocation, digital service channel design, and spatial planning prioritization at DPMPTSP Tangerang Selatan. Beyond silhouette alone, a comprehensive multi-metric evaluation (Davies–Bouldin and Calinski–Harabasz indices across k=2..6 for all four algorithms), a 500-iteration bootstrap stability analysis, and one-way ANOVA significance tests on the resulting cluster profiles were conducted to validate the robustness of the chosen segmentation.

References
  • MacQueen, J. 1967. Some methods for classification and analysis of multivariate observations. Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability, vol. 1, pp. 281-297. University of California Press, Berkeley.
  • Rousseeuw, P.J. 1987. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. Journal of Computational and Applied Mathematics, 20,53-65. https://doi.org/10.1016/0377-0427(87)90125-7
  • Ikotun, A.M., Ezugwu, A.E., Abualigah, L., Abuhaija, B., and Heming, J. 2023. K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data. Information Sciences, 622,178-210. https://doi.org/10.1016/j.ins.2022.11.139
  • Ezugwu, A.E., Ikotun, A.M., Oyelade, O.O., Abualigah, L., Agushaka, J.O., Eke, C.I., and Akinyelu, A.A. 2022. A comprehensive survey of clustering algorithms: State-of-the-art machine learning applications, taxonomy, challenges, and future research prospects. Engineering Applications of Artificial Intelligence, 110, Article 104743. https://doi.org/10.1016/j.engappai.2022.104743
  • Ward, J.H. 1963. Hierarchical grouping to optimize an objective function. Journal of the American Statistical Association, 58(301), 236-244. https://doi.org/10.1080/01621459.1963.10500845
  • Jain, A.K. 2010. Data clustering: 50 years beyond K-Means. Pattern Recognition Letters, 31(8), 651-666. https://doi.org/10.1016/j.patrec.2009.09.011
  • Halkidi, M., Batistakis, Y., and Vazirgiannis, M. 2001. On clustering validation techniques. Journal of Intelligent Information Systems, 17(2-3), 107-145. https://doi.org/10.1023/A:1012801612483
  • Liu, F.T., Ting, K.M., and Zhou, Z.H. 2012. Isolation-based anomaly detection. ACM Transactions on Knowledge Discovery from Data, 6(1), Article 3, pp. 1-39. https://doi.org/10.1145/2133360.2133363
  • Warrens, M.J. and van der Hoef, H. 2022. Understanding the adjusted Rand index and other partition comparison indices based on counting object pairs. Journal of Classification, 39,487-509. https://doi.org/10.1007/s00357-022-09413-z
  • Pal, S. and Heumann, C. 2022. Clustering compositional data using Dirichlet mixture model. PLOS ONE, 17(5), e0268438. https://doi.org/10.1371/journal.pone.0268438
  • Ng, A.Y., Jordan, M.I., and Weiss, Y. 2001. On spectral clustering: Analysis and an algorithm. Advances in Neural Information Processing Systems (NIPS), 14,849-856.
  • Arthur, D. and Vassilvitskii, S. 2007. k-means++: the advantages of careful seeding. Proceedings of the 18th ACM-SIAM Symposium on Discrete Algorithms (SODA), pp. 1027-1035.
  • Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. 2011. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12,2825-2830.
  • Greenacre, M., Blanco, V., Bove, G., Cannari, M., Carlier, A., Castellan, G., Chavent, M., and Greenacre, L. 2023. Compositional data analysis. Nature Reviews Methods Primers, 3,14. https://doi.org/10.1038/s43586-023-00195-3
  • Xu, D. and Tian, Y. 2015. A comprehensive survey of clustering algorithms. Annals of Data Science, 2(2), 165-193. https://doi.org/10.1007/s40745-015-0040-1
  • Soffan, S., Bramantoro, A., and Alzahrani, A.A. 2025. Combination of machine learning and data envelopment analysis to measure the efficiency of the Tax Service Office. PeerJ Computer Science, 11, e2672. https://doi.org/10.7717/peerj-cs.2672
  • Batty, M. 2013. The New Science of Cities. MIT Press, Cambridge, MA.
  • Anselin, L. 1995. Local indicators of spatial association: LISA. Geographical Analysis, 27(2), 93-115. https://doi.org/10.1111/j.1538-4632.1995.tb00338.x
  • Fadhel, M.A., Duhaim, A.M., Saihood, A., Sewify, A., Al-Hamadani, M.N.A., and colleagues. 2024. Assessing urban vulnerability to emergencies: a spatiotemporal approach using K-Means clustering. Land, 13(11), 1744. https://doi.org/10.3390/land13111744
  • Lenssen, L. and Schubert, E. 2023. Medoid silhouette clustering with automatic cluster number selection. Data Mining and Knowledge Discovery, 38,1523-1564. https://doi.org/10.1007/s10618-023-00979-9
  • Shao, C., Du, X., Yu, J., and Chen, J. 2022. Cluster-based improved Isolation Forest. Entropy, 24(5), 611. https://doi.org/10.3390/e24050611
  • Schubert, E. 2023. Stop using the elbow criterion for k-means and how to choose the number of clusters instead. ACM SIGKDD Explorations Newsletter, 25(1), 36-42. https://doi.org/10.1145/3606274.3606278
  • Gracias, J.S., Parnell, G.S., Specking, E., Pohl, E.A., and Buchanan, R. 2023. Smart cities: A structured literature review. Smart Cities, 6(4), 1719-1743. https://doi.org/10.3390/smartcities6040079
  • Alamsyah, A. and Muhammad, M.A. 2024. Leveraging generative AI for public service innovation: a path to smart government in Indonesia. Digital Government: Research and Practice, 5(4), Article 32. https://doi.org/10.1145/3761820
  • Xia, J., Zhang, Y., Song, J., Chen, Y., Wang, Y., and Liu, S. 2022. Revisiting dimensionality reduction techniques for visual cluster analysis: An empirical study. IEEE Transactions on Visualization and Computer Graphics, 28(1), 529-539. https://doi.org/10.1109/TVCG.2021.3114694
  • Rousseeuw, P.J. and Kaufman, L. 1990. Finding Groups in Data: An Introduction to Cluster Analysis. Wiley, New York.
  • [AUTHOR TO COMPLETE] - Reference on clustering for Indonesian regional development or fiscal capacity analysis.
  • [AUTHOR TO COMPLETE] - Reference on data-driven public service analytics in Southeast Asian municipalities.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

K-Means clustering silhouette coefficient permit segmentation sub-district Tangerang Selatan DPMPTSP outlier detection method comparison StandardScaler

Powered by PhDFocusTM