TY - GEN
T1 - Keyword Extraction Performance Analysis
AU - Kumbhar, A.
AU - Savargaonkar, M.
AU - Nalwaya, A.
AU - Bian, C.
AU - Abouelenien, M.
N1 - DBLP License: DBLP's bibliographic metadata records provided through http://dblp.org/ are distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Although the bibliographic metadata records are provided consistent with CC0 1.0 Dedication, the content described by the metadata records is not. Content may be subject to copyright, rights of privacy, rights of publicity and other restrictions.
PY - 2019
Y1 - 2019
N2 - This paper presents a survey-cum-evaluation of methods for the comprehensive comparison of the task of keyword extraction using datasets of various sizes, forms, and genre. We use four different datasets which includes Amazon product data-Automotive, SemEval 2010, TMDB and Stack Exchange. Moreover, a subset of 100 Amazon product reviews is annotated and utilized for evaluation in this paper, to our knowledge, for the first time. Datasets are evaluated by five Natural Language Processing approaches (3 unsupervised and 2 supervised), which include TF-IDF, RAKE, TextRank, LDA and Shallow Neural Network. We use a ten-fold cross-validation scheme and evaluate the performance of the aforementioned approaches using recall, precision and F-score. Our analysis and results provide guidelines on the proper approaches to use for different types of datasets. Furthermore, our results indicate that certain approaches achieve improved performance with certain datasets due to inherent characteristics of the data.
AB - This paper presents a survey-cum-evaluation of methods for the comprehensive comparison of the task of keyword extraction using datasets of various sizes, forms, and genre. We use four different datasets which includes Amazon product data-Automotive, SemEval 2010, TMDB and Stack Exchange. Moreover, a subset of 100 Amazon product reviews is annotated and utilized for evaluation in this paper, to our knowledge, for the first time. Datasets are evaluated by five Natural Language Processing approaches (3 unsupervised and 2 supervised), which include TF-IDF, RAKE, TextRank, LDA and Shallow Neural Network. We use a ten-fold cross-validation scheme and evaluate the performance of the aforementioned approaches using recall, precision and F-score. Our analysis and results provide guidelines on the proper approaches to use for different types of datasets. Furthermore, our results indicate that certain approaches achieve improved performance with certain datasets due to inherent characteristics of the data.
KW - NLP
KW - Text Mining
UR - https://www.scopus.com/pages/publications/85065625031
UR - https://www.scopus.com/pages/publications/85065625031
UR - https://www.mendeley.com/catalogue/e5ed4a71-db51-3559-b523-f7bae8b77e8d/
U2 - 10.1109/MIPR.2019.00111
DO - 10.1109/MIPR.2019.00111
M3 - Conference contribution
SN - 9781728111988
T3 - Proceedings - 2nd International Conference on Multimedia Information Processing and Retrieval, MIPR 2019
SP - 550
EP - 553
BT - Proceedings - 2nd International Conference on Multimedia Information Processing and Retrieval, MIPR 2019
ER -