The Web has become a large ecosystem that reaches billions of users through information processing and sharing, and most of this information resides in pixels. Web based services like YouTube and Flickr, and social networks such as Facebook have become increasingly popular, enabling users to eas- ily upload, share and annotate massive numbers of images and videos. Therefore, there is a critical need for novel algorithms able to understand big visual data and exploit noisy user annotations. Despite the recent success in visual recognition using a fully supervised setting, learning with weak labels and transferring knowledge to novel domains is still very challenging. This is a fundamental task in the open world, where the distribution of visual concepts follows a long tail that might change over time. Thus, we need a indexing algorithm that will index large scale data in order to reduce the searching and retrieval of image data withing massive data. This paper describes former techniques for the same and proposed fusion technique using k-means and residual index algorithm. k-means has ability to extract centroid from large scale data. It can be used with other algorithm to reduce the clustering time and accuracy like residual k-means inverted index. We conclude our paper by providing insights into the proposed indexing scheme and role of it in searching the data, use of it for indexing and tagging of data. Finally we discus future work to get more accuracy within the k-means domain.
Volume 11 | 01-Special Issue
Pages: 760-764