Thesis linked to the implementation of the María de Maeztu Strategic Research Program.

Open access to PhD thesis carried out at the Department can be found at TDX

Please visit these pages for information on our PhD, MSc and BSc programs.

 

Back PDF Digest - free online tool to parse PDF files

You can use our freely online tool to parse yours PDF files. Our approach is based on the PDFdigest tool, a PDF textual content extraction system specially designed to extract scientific articles' headings and logical structure (title, authors, abstract and so on) and its textual content to. The result is provided in a XML file. Furthermore, PDFdigest also provides a structured HTML file as a clone of the original PDF file. 

In addition, the pre-processing step implemented in DrInventor (link) is applied to the previous XML file in order to mark off tokens and sentences. As a result, we also provide an additional GATE document.

Details at http://scientmin.taln.upf.edu/pdfdigest/pdfparser.php