Performance-Storage Trade-Off Analysis of MongoDB Indexing Strategies Using Large-Scale E-Commerce Data
DOI:
https://doi.org/10.47794/jesica.v3i2.45Keywords:
Benchmarking, Compound Index, Index Optimization, MongoDB, NoSQL Database, Query PerformanceAbstract
Indexing strategy in document-oriented databases represents a fundamental design decision that directly shapes both query performance and storage consumption, yet empirical evidence quantifying this trade-off at production-grade scale remains limited in existing literature. This study evaluated three indexing configurations in MongoDB, namely no index, single-field index on product identifier, and compound index on product identifier together with order status, across four collection sizes of 1 million, 5 million, 10 million, and 13 million documents derived from a synthetic large-scale e-commerce dataset obtained from Kaggle containing approximately 13 million order-line records. A controlled benchmarking procedure was employed in which each indexing condition was tested under identical query workloads repeated 30 times per configuration, with query execution time, index creation time, and index storage size recorded as evaluation metrics. Results showed that unindexed collections produced full collection scans with mean execution times scaling from 1,247.70 ms at 1 million documents to 33,300.07 ms at 13 million documents, while both indexed conditions reduced execution times to single-digit milliseconds by activating index scan paths. For multi-predicate queries, the compound index outperformed the single-field index by a factor of 6.8 at full scale, recording 1.33 ms against 9.03 ms, while incurring only 1 MB of additional storage overhead. These findings indicate that compound indexing represents the most balanced strategy for high-volume e-commerce query workloads, delivering substantial performance gains at negligible additional storage cost.
Downloads
References
[1] A. Makris, K. Tserpes, G. Spiliopoulos, D. Zissis, and D. Anagnostopoulos, “MongoDB Vs PostgreSQL: A comparative study on performance aspects,” Geoinformatica, vol. 25, no. 2, pp. 243–268, Apr. 2021, doi: 10.1007/s10707-020-00407-w.
[2] D. A. Saputra, Mohammad Adi Farich, Muhammad Naafi’an Anugerah, Imam Prayogo Pujiono, and Dicky Anggriawan Nugroho, “Perbandingan Kinerja MongoDB dan MySQL pada Aplikasi dengan Beban Data Besar,” SAINSTECH: JURNAL PENELITIAN DAN PENGKAJIAN SAINS DAN TEKNOLOGI, vol. 35, no. 4, pp. 71–78, Dec. 2025, doi: 10.37277/stch.v35i4.2543.
[3] A. Hamadou et al., “A COMPARATIVE EXPERIMENTAL STUDY OF INDEX PERFORMANCE IN MONGODB AND POSTGRESQL,” Int. J. Eng. Sci. Res. Technol., vol. 15, no. 4, 2026, [Online]. Available: http://www.ijesrt.com.
[4] A. Juan Syahwali, J. Jenderal Ahmad Yani No, P. Palembang, and S. Selatan, “Pemanfaatan MongoDB dalam Sistem Informasi Akademik untuk Pengelolaan Data Mahasiswa Pada Universitas Bina Darma,” Jurnal Sains Student Research, vol. 3, no. 2, pp. 466–469, 2025, doi: 10.61722/jssr.v3i2.4335.
[5] A. A. Muhamad, D. Adhi Kusuma, I. Renaldi, and A. Wibowo, “Studi Literatur Komparasi SQL dan NoSQL dalam Pemilihan Basis Data Ideal untuk Skalabilitas Tinggi,” 2025. [Online]. Available: https://jurnal.unw.ac.id/index.php/IKN.
[6] M. Nuriev, R. Zaripova, O. Yanova, I. Koshkina, and A. Chupaev, “Enhancing MongoDB query performance through index optimization,” in E3S Web of Conferences, EDP Sciences, Jun. 2024, doi: 10.1051/e3sconf/202453103022.
[7] D. Tao, E. Liu, S. R. Kadupitige, M. Cahill, A. Fekete, and U. Röhm, “First Past the Post: Evaluating Query Optimization in MongoDB,” Preprint, Sep. 2024, [Online]. Available: http://arxiv.org/abs/2409.16544.
[8] R. Ivanov, “MongoDB Aggregation Pipeline Performance: Analysis of Query Plan Selection and Optimizer Behavior Across Versions and Collection Scales,” Information (Switzerland), vol. 17, no. 5, May 2026, doi: 10.3390/info17050488.
[9] V. Sumalatha and S. Pabboju, “Optimal Index Selection using Optimized Deep Deterministic Policy Gradient for NoSQL Database,” Engineering, Technology and Applied Science Research, vol. 14, no. 6, pp. 18125–18130, Dec. 2024, doi: 10.48084/etasr.8832.
[10] G. Raphaela Mei Lanny Br Aritonang, M. Arief Hasan, M. Khasbulla Ridwan, M. Nur Iman, R. Carlo Pratama Silalahi, and R. Suranta Sipayung, “Optimasi Query SQL Server dengan Teknik Indexing dan Performance Monitoring,” JATI : Jurnal Mahasiswa Teknik Informatika, vol. 9 No. 2, 2025, doi: 10.36040/jati.v9i2.13179
[11] K. S. Lestari, F. A. Riansyah, A. H. Taqyudin, and J. I. Saputro, “Optimalisasi Kinerja Indexing Dalam Pencarian Data Di Database Mongodb,” Journal Sensi:Strategic Of Education in Information System , vol. 12 No.01, 2026, doi:10.33050/sensi.v12i1.4351
[12] T. P. Avrylya and Y. A. Susetyo, “Perbandingan Response Time Pencarian Menggunakan Text Indexing Pada MongoDB dan ArangoDB Berbasis Web,” MALCOM: Indonesian Journal of Machine Learning and Computer Science, vol. 4, no. 3, pp. 777–785, May 2024, doi: 10.57152/malcom.v4i3.1308.
[13] Imam Thoib, Beda Puspita Candra, Nafis Sururi, and Danang Satya Nugraha, “Perbandingan Performa Pencarian Data Berbasis Teks dengan Dan Tanpa Full-text Index pada Basis Data MySQL,” INSOLOGI: Jurnal Sains dan Teknologi, vol. 3, no. 6, pp. 663–673, Dec. 2024, doi: 10.55123/insologi.v3i6.4596.
[14] Swain, “Synthetic E-Commerce Dataset (Very Large),” 2025. Accessed: Jun. 01, 2026. [Online]. Available: https://www.kaggle.com/datasets/swainproject/synthetic-e-commerce-dataset-very-large.
[15] A. Hodijah, “Experimental Evaluation of Indexing Between HBase and MongoDB,” Atlantis Press Proceedings of the 2nd International Seminar of Science and Applied Technology (ISSAT 2021), 2021, doi: 10.2991/aer.k.211106.002
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Kunti Inayati, Umi Chotijah

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
