The second Maria de Maeztu Strategic Research Program (CEX2021-001195-M) of the Department of Information and Communication Technologies (DTIC) takes place between 2023 and 2026. The website for this program is under construction. You can find some details in this news.

The first María de Maeztu Strategic Research Program (MDM-2015-0502) took place between January 2016 and June 2020. It was focused on data-driven knowledge extraction, boosting synergistic research initiatives across our different research areas.

Back PDF Digest - free online tool to parse PDF files

You can use our freely online tool to parse yours PDF files. Our approach is based on the PDFdigest tool, a PDF textual content extraction system specially designed to extract scientific articles' headings and logical structure (title, authors, abstract and so on) and its textual content to. The result is provided in a XML file. Furthermore, PDFdigest also provides a structured HTML file as a clone of the original PDF file. 

In addition, the pre-processing step implemented in DrInventor (link) is applied to the previous XML file in order to mark off tokens and sentences. As a result, we also provide an additional GATE document.

Details at http://scientmin.taln.upf.edu/pdfdigest/pdfparser.php

Department of Information and Communication Technologies, UPF

Grant CEX2021-001195-M funded by MCIN/AEI /10.13039/501100011033


 


Department of Information and Communication Technologies, UPF

[email protected]

  • Àngel Lozano - Scientific director
  • Aurelio Ruiz - Program management