Skip to main content

Add support for reading plain text from PDFs

Currently there’s no built-in command line tool on macOS for reading plain text from PDFs (for summarisation etc), as far as I know. Textutil can be used for reading from RTFs, Word Docs, HTML and more, but not PDF. We’d have to embed some helper binary in the same way that we do for ffmpeg.

Status: Completed2 comments

Log in to comment and vote

Comments2

  • joethephish

    Team•

    Nov 21, 2025

    This is now implemented via support for Homebrew and `pdftotext`. For example, you can say “Extract the text from this PDF”.

    (Note that “organising documents by content” is significantly more complicated. I’m mulling it over though!)

  • Thierry Weber

    •

    Apr 7, 2025

    Good idea, this option would allow you, for example, to scan entire folders and thus ‘organise’ documents by type of content (e.g. invoices). I find the idea very useful when it comes to scanned documents all stored in the same folder.