A Survey: Feature Selection and Machine Learning Methods for Tamil Text Classification

  • N. Rajkumar, T. S. Subashini, K. Rajan, V. Ramalingam

Abstract

The objective and aim of this survey is to issue a sketch for Machine Learning algorithm of Tamil Text classification. Since the past half decade, almost all regional language increased in the size of its computerized repository, Tamil language also has captivated their all type of information to online digital repositories. The text data in digital format both in offline and online mode is increased in very huge, so there is requirement for a document classification system to classify the document according to its class name. Document classification is a method that examines the text data assumed in the document and classify the categories. Text being occurred in Tamil language requires the challenges of Natural Language processing (NLP). In this paper gives a survey of Tamil Text Classification works completed on Tamil Language content. The automated document classification implies at the crossroads of Machine Learning (ML), NLP and Information Retrieval (IR). This study mainly give readers fruitfully acquire the necessary important information about the needed algorithm and its associated techniques. Thus, i believe that through this study it is useful to other professionals and researcher to propose new technique in the province of Tamil text classification. This study shows that supervised learning algorithms the decision tree (DT), Artificial Neural Network (ANN), k-nearest neighbor (kNN) classifier, N-gram, Naive Bayes (NB) and Deep Learning performed optimal to Text Classification task.

Published
2020-10-30
How to Cite
N. Rajkumar, T. S. Subashini, K. Rajan, V. Ramalingam. (2020). A Survey: Feature Selection and Machine Learning Methods for Tamil Text Classification. International Journal of Advanced Science and Technology, 29(05), 13917 - 13922. Retrieved from http://sersc.org/journals/index.php/IJAST/article/view/33194