Author Affiliations
[2] Associate Professor, Dept. of CSE Er. Perumal Manimekalai College of Engineering, Hosur, Krishnagiri, Tamil Nadu.
[1] [3] [4] [5] Dept. of CSE Er. Perumal Manimekalai College of Engineering, Hosur, Krishnagiri, Tamil Nadu.
Abstract
Manual literature review techniques consume approximately 60–70% of research time in academic workflows. With the exponential growth of scientific publications, researchers face increasing challenges in extracting structured knowledge—such as methodologies, datasets, and contributions—from unstructured PDF documents. Existing tools provide only fragmented solutions and lack end-to-end automation. This project proposes an advanced AI-driven pipeline that integrates Apache Tika, DeBERTa, SimCSE, KMeans clustering, BART summarization, and CRF+LSTM+BERT-based Named Entity Recognition (NER) for automated knowledge extraction from research papers. The ensemble pipeline demonstrates superior performance with a structured extraction accuracy of 98% F1-score, clustering silhouette score of 92%, and ROUGE-2 summarization score of 0.45. Model validation is conducted using cross-validation techniques, including F1-score, silhouette coefficient, ROUGE metrics, extraction accuracy, and semantic coherence analysis. The proposed system significantly reduces manual effort, enabling researchers to focus on innovation by automating literature review processes.
Keywords: Knowledge extraction, Literature review automation, Apache Tika, DeBERTa, SimCSE, BART, CRF, LSTM, F1-score, ROUGE metrics.
How to Cite This Article
Ravishree M M, S Suganya, Lithin Kumar M, Pavan C, Prasanth C (2026). Automated Knowledge Extraction from Research Papers for Literature Surveys. International Journal of Innovative Research in Multidisciplinary Education & Technology (IJIRMET), 11(3).