BIG DATA ANALYSIS OF RAINFALL USING APACHE SPARK: A CASE STUDY OF INDONESIA

Authors

  • Arnold Paulinus Partogi Sinaga Universitas Prima Indonesia, Indonesia Author
  • Jhon Nobel Helsingki Turnip Universitas Prima Indonesia, Indonesia Author
  • Mohammad Irfan Fahmi Universitas Prima Indonesia, Indonesia Author

DOI:

https://doi.org/10.54314/jssr.v9i4.7078

Keywords:

Rainfall Analysis, Apache Spark, Big Data Analytics

Abstract

Rainfall is one of the most important climatological parameters that plays a significant role in supporting agriculture, water resource management, and hydrometeorological disaster mitigation in Indonesia. As the volume of meteorological data continues to increase, technologies capable of processing large-scale data efficiently are required. This study aims to analyze rainfall patterns in Indonesia using a Big Data approach by utilizing Apache Spark as the primary data processing framework. The data used in this research consist of daily rainfall data from 34 provinces in Indonesia obtained from the Meteorology, Climatology, and Geophysics Agency (BMKG) in CSV format. This research employs a descriptive quantitative method consisting of data ingestion, data cleaning, data transformation, and descriptive statistical analysis using Apache Spark with PySpark. The analysis was conducted to obtain information regarding the number of valid records, average rainfall, maximum rainfall, and minimum rainfall values for each province. The results indicate that Apache Spark is capable of processing rainfall data in a structured, efficient, and automated manner. Based on the analysis, West Sumatra Province recorded the highest average rainfall at 16.01 mm/day, while Central Sulawesi Province recorded the lowest average rainfall at 0.30 mm/day. In addition, the highest maximum rainfall was observed in Bengkulu Province at 186.4 mm. The use of Spark DataFrame and aggregation functions such as count(), avg(), max(), and min() proved effective in supporting data processing and analysis activities. The findings demonstrate that Apache Spark can serve as an effective solution for large-scale climatological data processing and provide valuable information to support decision-making in agriculture, water resource management, and disaster mitigation in Indonesia.

Downloads

Download data is not yet available.

References

World Meteorological Organization (WMO), State of Climate Services 2021: Water. Geneva, Switzerland: WMO, 2021

Intergovernmental Panel on Climate Change (IPCC), Climate Change 2021: The Physical Science Basis. Cambridge, U.K.: Cambridge University Press, 2021

World Meteorological Organization (WMO), State of the Global Climate 2023. Geneva, Switzerland: WMO, 2024.

United Nations, Global Water and Climate Assessment Report. New York, NY, USA: United Nations, 2023.

Food and Agriculture Organization (FAO), The State of Food and Agriculture 2023. Rome, Italy: FAO, 2023.

World Meteorological Organization (WMO), Atlas of Mortality and Economic Losses from Weather, Climate and Water Extremes (1970–2021). Geneva, Switzerland: WMO, 2023

United Nations Office for Disaster Risk Reduction (UNDRR), Global Assessment Report on Disaster Risk Reduction 2022. Geneva, Switzerland: UNDRR, 2022.

Intergovernmental Panel on Climate Change (IPCC), Climate Change 2022: Impacts, Adaptation and Vulnerability. Cambridge, U.K.: Cambridge University Press, 2022.

World Meteorological Organization (WMO), State of Climate Services 2022. Geneva, Switzerland: WMO, 2022

A. Pratama, D. S. Wibowo, and L. Kurniawan, “Big data characteristics in meteorological datasets for climate monitoring,” Journal of Environmental Informatics, vol. 39, no. 2, pp. 88–97, 2023.

X. Liu, Y. Sun, and M. Chen, “Distributed computing frameworks for large-scale climate data analytics,” Future Generation Computer Systems, vol. 143, pp. 275–288, 2023.

R. Singh and P. Gupta, “Statistical analysis of climate data using distributed computing frameworks,” Big Data Research, vol. 31, 2024.

M. Zaharia et al., “Apache Spark: A Unified Engine for Large-Scale Data Processing,” Communications of the ACM, vol. 64, no. 6, pp. 56–65, 2021.

T. Brown and R. Wilson, “Climate data processing using Apache Spark: A scalable framework for meteorological analysis,” Journal of Big Data, vol. 10, no. 1, 2023

M. Rodriguez and A. Torres, “Efficient meteorological data processing using Apache Spark and cloud computing,” IEEE Access, vol. 13, pp. 11245–11260, 2025

Apache Software Foundation, Apache Spark Documentation: PySpark API Reference, 2025

C. M. Liyew and H. A. Melese, “Machine learning techniques to predict daily rainfall amount,” Journal of Big Data, vol. 8, no. 1, 2021

Downloads

Published

2026-09-02

How to Cite

BIG DATA ANALYSIS OF RAINFALL USING APACHE SPARK: A CASE STUDY OF INDONESIA. (2026). JOURNAL OF SCIENCE AND SOCIAL RESEARCH, 9(4), 6400-6409. https://doi.org/10.54314/jssr.v9i4.7078