A massive amount of data contains useful information known as Big Data. To analyze the real time data clustering algorithms are proposed. In this paper, a framework called Apache Spark for implementing the partition based clustering Algorithms (for e.g.: fuzzy c-means) is used. It works better for iterative algorithms by means of in-memory computations and scalability as the data chosen for analysis occurs at real time. The spark gives low computational requirements to cluster large data set comparative with other frameworks. We propose Incremental Scalable Random Sampling with Iterative Optimization Fuzzy c-means Algorithm (ISRSIO-FCM) to cluster the large data by using Apache spark. This result is compared with Scalable Random sampling plus Extension Fuzzy c-means (Srse-FCM) and Scalable Literal Fuzzy c-means (SLFCM) which proves that proposed work produces good results than others.
Volume 11 | 04-Special Issue
Pages: 940-950