PDF Digest - free online tool to parse PDF files
We develop a large number of software tools and hosting infrastructures to support the research developed at the Department. We will be detailing in this section the different tools available. You can take a look for the moment at the offer available within the UPF Knowledge Portal, the innovations created in the context of EU projects in the Innovation Radar and the software sections of some of our research groups:
Artificial Intelligence |
Nonlinear Time Series Analysis |
Web Research |
Music Technology |
Interactive Technologies |
Barcelona MedTech |
Natural Language Processing |
Nonlinear Time Series Analysis |
UbicaLab |
Wireless Networking |
Educational Technologies |
You can use our freely online tool to parse yours PDF files. Our approach is based on the PDFdigest tool, a PDF textual content extraction system specially designed to extract scientific articles' headings and logical structure (title, authors, abstract and so on) and its textual content to. The result is provided in a XML file. Furthermore, PDFdigest also provides a structured HTML file as a clone of the original PDF file.
In addition, the pre-processing step implemented in DrInventor (link) is applied to the previous XML file in order to mark off tokens and sentences. As a result, we also provide an additional GATE document.
Details at http://scientmin.taln.upf.edu/