Evolution and Research Dynamics of Metadata Extraction in Academic Documents

Autores/as

  • Anderson Joseph Ochoa-Trujillo
  • Cristian Andres Villarreal-Orozco
  • Juan Pablo Hoyos-Sanchez

DOI:

https://doi.org/10.18041/2619-4465/interfaces.2.14060

Palabras clave:

Extracción de metadatos, Análisis cienciométrico, procesamiento de lenguaje natural, Tree of Science, aprendizaje profundo

Resumen

Metadata extraction is a key technique for organizing, retrieving, and analyzing unstructured information in academic, clinical, and legal documents. It enables the transformation of raw text into structured knowledge, supporting applications such as text mining, recommendation systems, and semantic ontologies. This field has seen substantial growth, driven by advancements in Natural Language Processing (NLP). This paper analyzes the scientific production on metadata extraction from 2001 to 2025, highlighting three phases of development: a slow emergence (2001–2009, with a 2.25% annual growth rate), consolidation (2010–2016, with a 12.25% annual growth rate), and exponential expansion (2017–2024, with a 24.98% annual growth rate). The highest publication output occurred in 2023 (141 articles), while citation impact peaked in 2020 (2,922 citations), partly due to health-related research during the COVID-19 pandemic. A systematic search in Web of Science and Scopus yielded 1,177 unique articles. The analysis, conducted following the PRISMA framework, employed scientometric methods in conjunction with the Tree of Science approach to map the structural dynamics of knowledge in the field. The United States emerged as the leading contributor, accounting for 272 publications (20.25%) and 8,075 citations (30.81%). Notable researchers such as Zhang J. (14 publications, 974 citations, H-index = 10) demonstrate significant individual influence. The findings further highlight a methodological shift toward deep learning and transformer-based models, underscoring the field’s strategic importance in contemporary digital ecosystems.

Descargas

Los datos de descarga aún no están disponibles.

Referencias

[1]M. Lipinski, K. Yao, C. Breitinger, J. Beel, and B. Gipp, “Evaluation of header metadata extraction approaches and tools for scientific PDF documents,” in Proceedings of the 13th ACM/IEEE-CS joint conference on Digital libraries, in JCDL ’13. New York, NY, USA: Association for Computing Machinery, Jul. 2013, pp. 385–386. doi: 10.1145/2467696.2467753.

[2]R. Abascal, B. Rumpler, and J.-M. Pinon, “Information Retrieval in Digital Theses Based on Natural Language Processing Tools,” in Advances in Natural Language Processing, J. L. Vicedo, P. Martínez-Barco, R. Muńoz, and M. Saiz Noeda, Eds., Berlin, Heidelberg: Springer, 2004, pp. 172–182. doi: 10.1007/978-3-540-30228-5_16.

[3]Z. Boukhers and A. Bouabdallah, “Vision and natural language for metadata extraction from scientific PDF documents: a multimodal approach,” in Proceedings of the 22nd ACM/IEEE Joint Conference on Digital Libraries, in JCDL ’22. New York, NY, USA: Association for Computing Machinery, Jun. 2022, pp. 1–5. doi: 10.1145/3529372.3533295.

[4]J. Du et al., “Use of deep learning-based NLP models for full-text data elements extraction for systematic literature review tasks,” Sci. Rep., vol. 15, no. 1, p. 19379, Jun. 2025, doi: 10.1038/s41598-025-03979-5.

[5]P. Raval and H. Bhaidasna, “A Review of Extracting Metadata from Scholarly Articles using Natural Language Processing (NLP),” in 2024 3rd International Conference on Automation, Computing and Renewable Systems (ICACRS), Dec. 2024, pp. 1355–1359. doi: 10.1109/ICACRS62842.2024.10841656.

[6]A. Vaswani et al., “Attention is all you need,” Adv. Neural Inf. Process. Syst., vol. 30, 2017.

[7]J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” presented at the Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 2019, pp. 4171–4186.

[8]M. J. Page et al., “PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviews,” BMJ, vol. 372, p. n160, Mar. 2021, doi: 10.1136/bmj.n160.

[9]M. Zuluaga, S. Robledo, O. Arbelaez-Echeverri, G. A. Osorio-Zuluaga, and N. Duque-Méndez, “Tree of Science - ToS: A Web-Based Tool for Scientific Literature Recommendation. Search Less, Research More!,” Issues Sci. Technol. Librariansh., no. 100, Art. no. 100, Aug. 2022, doi: 10.29173/istl2696.

[10]S. Robledo, L. Valencia, M. Zuluaga, O. A. Echeverri, and J. W. A. Valencia, “tosr: Create the Tree of Science from WoS and Scopus,” J. Scientometr. Res., vol. 13, no. 2, pp. 459–465, Jul. 2025, doi: 10.5530/jscires.13.2.36.

[11]L. Ganti, N. A. Persaud, and T. S. Stead, “Bibliometric analysis methods for the medical literature,” Acad. Med. Surg., Jan. 2025, doi: 10.62186/001c.129134.

[12]W. M. To and B. T. W. Yu, “Artificial Intelligence Research in Tourism and Hospitality Journals: Trends, Emerging Themes, and the Rise of Generative AI,” Tour. Hosp., vol. 6, no. 2, Art. no. 2, Jun. 2025, doi: 10.3390/tourhosp6020063.

[13]S. Valencia, M. Zuluaga, M. C. Florian Pérez, K. F. Montoya-Quintero, M. S. Candamil-Cortés, and S. Robledo, “Human Gut Microbiome: A Connecting Organ Between Nutrition, Metabolism, and Health,” Int. J. Mol. Sci., vol. 26, no. 9, Art. no. 9, Jan. 2025, doi: 10.3390/ijms26094112.

[14]A. Vargas-Hernández, S. Robledo, and G. R. Quiceno, “Virtual Teaching for Online Learning from the Perspective of Higher Education: A Bibliometric Analysis,” J. Scientometr. Res., vol. 13, no. 2, pp. 406–418, Jul. 2025, doi: 10.5530/jscires.13.2.32.

[15]I. Irwanto and T. D. S. Rini, “Research trends in blended learning in chemistry: A bibliometric analysis of scopus indexed publications (2012–2022),” J. Turk. Sci. Educ., vol. 21, no. 3, pp. 566–578, Sep. 2024, doi: 10.36681/tused.2024.030.

[16]H. Ma and L. Ismail, “Bibliometric analysis and systematic review of digital competence in education,” Humanit. Soc. Sci. Commun., vol. 12, no. 1, p. 185, Feb. 2025, doi: 10.1057/s41599-025-04401-1.

[17]S. Pyysalo et al., “BioInfer: a corpus for information extraction in the biomedical domain,” BMC Bioinformatics, vol. 8, pp. 1–24, 2007.

[18]E. Pasolli, D. T. Truong, F. Malik, L. Waldron, and N. Segata, “Machine Learning Meta-analysis of Large Metagenomic Datasets: Tools and Biological Insights,” PLOS Comput. Biol., vol. 12, no. 7, p. e1004977, Jul. 2016, doi: 10.1371/journal.pcbi.1004977.

[19]M. Trotzek, S. Koitka, and C. M. Friedrich, “Utilizing Neural Networks and Linguistic Metadata for Early Detection of Depression Indications in Text Sequences,” IEEE Trans. Knowl. Data Eng., vol. 32, no. 3, pp. 588–601, Mar. 2020, doi: 10.1109/TKDE.2018.2885515.

[20]T. C. Rindflesch and M. Fiszman, “The interaction of domain knowledge and linguistic structure in natural language processing: interpreting hypernymic propositions in biomedical text,” Unified Med. Lang. Syst., vol. 36, no. 6, pp. 462–477, Dec. 2003, doi: 10.1016/j.jbi.2003.11.003.

[21]L. Luo et al., “An attention-based BiLSTM-CRF approach to document-level chemical named entity recognition,” Bioinformatics, vol. 34, no. 8, pp. 1381–1388, 2018.

[22]R. Janani and S. Vijayarani, “Text document clustering using spectral clustering algorithm with particle swarm optimization,” Expert Syst. Appl., vol. 134, pp. 192–200, 2019, doi: 10.1016/j.eswa.2019.05.030.

[23]N. Asadova et al., “T-matrix representation of optical scattering response: Suggestion for a data format,” J. Quant. Spectrosc. Radiat. Transf., vol. 333, p. 109310, Mar. 2025, doi: 10.1016/j.jqsrt.2024.109310.

[24]H. Turki et al., “MeSH2Matrix: combining MeSH keywords and machine learning for biomedical relation classification based on PubMed,” J. Biomed. Semant., vol. 15, no. 1, p. 18, Oct. 2024, doi: 10.1186/s13326-024-00319-w.

[25]C. A. Primiero et al., “A protocol for annotation of total body photography for machine learning to analyze skin phenotype and lesion classification,” Front. Med., vol. 11, p. 1380984, 2024, doi: 10.3389/fmed.2024.1380984.

[26]L. Akhtyamova, P. Martínez, K. Verspoor, and J. Cardiff, “Testing contextualized word embeddings to improve NER in Spanish clinical case narratives,” IEEE Access, vol. 8, pp. 164717–164726, 2020.

[27]S. Novichkova, S. Egorov, and N. Daraselia, “MedScan, a natural language processing engine for MEDLINE abstracts,” Bioinformatics, vol. 19, no. 13, pp. 1699–1706, 2003, doi: 10.1093/bioinformatics/btg207.

[28]L. Shi, C. Jianping, and X. Jie, “Prospecting Information Extraction by Text Mining Based on Convolutional Neural Networks-A Case Study of the Lala Copper Deposit, China,” IEEE Access, vol. 6, pp. 52286–52297, 2018, doi: 10.1109/ACCESS.2018.2870203.

[29]L. Chen, S. Xu, L. Zhu, J. Zhang, X. Lei, and G. Yang, “A deep learning based method for extracting semantic information from patent documents,” Scientometrics, vol. 125, no. 1, pp. 289–312, 2020, doi: 10.1007/s11192-020-03634-y.

[30]J.-L. Seng and J. T. Lai, “An intelligent information segmentation approach to extract financial data for business valuation,” Expert Syst. Appl., vol. 37, no. 9, pp. 6515–6530, 2010, doi: 10.1016/j.eswa.2010.02.134.

[31]I. Malashin, I. Masich, V. Tynchenko, A. Gantimurov, V. Nelyub, and A. Borodulin, “Image Text Extraction and Natural Language Processing of Unstructured Data from Medical Reports,” Mach. Learn. Knowl. Extr., vol. 6, no. 2, pp. 1361–1377, 2024, doi: 10.3390/make6020064.

[32]D. Tkaczyk, P. Szostek, M. Fedoryszak, P. J. Dendek, and Ł. Bolikowski, “CERMINE: Automatic extraction of structured metadata from scientific literature,” Int. J. Doc. Anal. Recognit., vol. 18, no. 4, pp. 317–335, 2015, doi: 10.1007/s10032-015-0249-8.

[33]R. K. Srihari, W. Li, T. Cornell, and C. Niu, “InfoXtract: A customizable intermediate level information extraction engine,” Nat. Lang. Eng., vol. 14, no. 1, pp. 33–69, 2008, doi: 10.1017/S1351324906004116.

[34]N. Xu et al., “Compare the performance of multiple binary classification models in microbial high-throughput sequencing datasets,” Sci. Total Environ., vol. 837, p. 155807, 2022.

[35]G. Mustafa et al., “Optimizing document classification: Unleashing the power of genetic algorithms,” IEEE Access, vol. 11, pp. 83136–83149, 2023.

[36]T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” ArXiv Prepr. ArXiv13013781, 2013.

[37]J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” presented at the Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, pp. 1532–1543.

[38]F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.

[39]A. Kumar and B. Starly, “‘FabNER’: information extraction from manufacturing process science domain literature using named entity recognition,” J. Intell. Manuf., vol. 33, no. 8, pp. 2393–2407, 2022.

[40]J. Wang, M. Li, Q. Diao, H. Lin, Z. Yang, and Y. Zhang, “Biomedical document triage using a hierarchical attention-based capsule network,” BMC Bioinformatics, vol. 21, pp. 1–20, 2020.

[41]Y. Wang et al., “A comparison of word embeddings for the biomedical natural language processing,” J. Biomed. Inform., vol. 87, pp. 12–20, Nov. 2018, doi: 10.1016/j.jbi.2018.09.008.

[42]D. Vukadin, A. S. Kurdija, G. Delač, and M. Šilić, “Information extraction from free-form CV documents in multiple languages,” IEEE Access, vol. 9, pp. 84559–84575, 2021.

[43]A. Deloose, G. Gysels, B. De Baets, and J. Verwaeren, “Combining natural language processing and multidimensional classifiers to predict and correct CMMS metadata,” Comput. Ind., vol. 145, p. 103830, 2023.

[44]Y. Zhang, X. Chang, Y. Lin, J. Mišić, and V. B. Mišić, “Exploring function call graph vectorization and file statistical features in malicious PE file classification,” IEEE Access, vol. 8, pp. 44652–44660, 2020.

[45]S. Tajjour, S. Garg, S. S. Chandel, and D. Sharma, “A novel hybrid artificial neural network technique for the early skin cancer diagnosis using color space conversions of original images,” Int. J. Imaging Syst. Technol., vol. 33, no. 1, pp. 276–286, 2023.

[46]C. Liu, D. Q. Huynh, Y. Sun, M. Reynolds, and S. Atkinson, “A vision-based pipeline for vehicle counting, speed estimation, and classification,” IEEE Trans. Intell. Transp. Syst., vol. 22, no. 12, pp. 7547–7560, 2020.

[47]S. Luo and J. Yu, “A semantic enhancement-based multimodal network model for extracting information from evidence lists,” Neural Netw., vol. 187, p. 107387, 2025.

[48]R. Benfer and J. Müller, “Semantic digital twin creation of building systems through time series based metadata inference–A review,” Energy Build., p. 114637, 2024.

[49]H. Gbada, K. Kalti, and M. A. Mahjoub, “Deep learning approaches for information extraction from visually rich documents: datasets, challenges and methods,” Int. J. Doc. Anal. Recognit. IJDAR, pp. 1–22, 2024.

[50]S. Yuan, “Design of Intelligent Document Categorization System for Office Software Combined with Neural Networks,” Appl. Math. Nonlinear Sci., vol. 9, no. 1, 2024, doi: 10.2478/amns-2024-3357.

[51]H.-C. Shin et al., “Deep convolutional neural networks for computer-aided detection: CNN architectures, dataset characteristics and transfer learning,” IEEE Trans. Med. Imaging, vol. 35, no. 5, pp. 1285–1298, 2016.

[52]N. Viani et al., “Supervised methods to extract clinical events from cardiology reports in Italian,” J. Biomed. Inform., vol. 95, p. 103219, 2019.

[53]S. Zhang and N. Elhadad, “Unsupervised biomedical named entity recognition: Experiments with clinical and biological texts,” J. Biomed. Inform., vol. 46, no. 6, pp. 1088–1098, 2013.

[54]R. A. Jonker, T. Almeida, R. Antunes, J. R. Almeida, and S. Matos, “Multi-head CRF classifier for biomedical multi-class named entity recognition on Spanish clinical notes,” Database, vol. 2024, p. baae068, 2024.

[55]S. Zhao and L. Li, “Temporal information extraction with the scalable cross-sentence context for electronic health records,” J. Biomed. Inform., vol. 128, p. 104052, 2022, doi: 10.1016/j.jbi.2022.104052.

[56]R. K. Zulkarneev, N. I. Yusupova, O. N. Smetanina, M. M. Gayanova, and A. M. Vulfin, “Method and models of extraction of knowledge from medical documents,” Inform. Autom., vol. 21, no. 6, pp. 1169–1210, 2022.

[57]D. C. Moreno and M. Vargas-Lombardo, “Design and construction of a NLP based knowledge extraction methodology in the medical domain applied to clinical information,” Healthc. Inform. Res., vol. 24, no. 4, pp. 376–380, 2018.

Descargas

Publicado

2026-09-11

Número

Sección

Artículos