Extract text

Copy all the text from the PDF

No database: AIMRAN doesn't store your files, text or results

What is Extract text?

Copying text from a PDF page by page is slow and the layout often gets jumbled. This tool pulls all the text out of the document at once so you can paste it into Word, an email or wherever you need it.

You can add separators with each page number and see how many words and characters the document has. If the PDF is a scan with no text, the tool tells you and points you to OCR.

How to use it

  1. Drag in the PDF.
  2. Turn on “Separate by pages” if you want markers between pages.
  3. Click “Extract text”.
  4. Copy the text or download it as a .txt file.

Advantages

Technical details

Extraction uses pdf.js to read the text fragments on each page and rebuilds the lines following the content order: it inserts a line break when the baseline changes and a space when it detects a horizontal gap between fragments. With the separate-by-pages option, “--- Page N ---” is added before each page. Only the PDF's real text is obtained; images and scanned documents contain no text and require OCR. Complex tables and columns may lose their original layout. Accepts PDFs up to 200 MB.

Frequently asked questions

How do I copy all the text from a PDF?

Drag in the PDF, click “Extract text” and then copy the result or download it as .txt.

Why doesn't any text come out of my PDF?

It is probably a scanned document made of images. Use the “Scanned PDF OCR” tool to recognize the text.

Is the formatting kept when extracting text?

Lines and paragraphs are kept, but not bold, sizes or tables. It is plain text.

How do I convert a PDF to TXT?

Extract the text and click download: you will get a .txt file with the PDF's content.

More PDF tools