7 ms·
MarkItDown: Python tool for converting files and office documents to Markdown
- ccbikai 2y agoI made a version that can run entirely within the browser https://www.html.zone/markitdown/ https://www.html.zone/markitdown/
- ezxs 2y agoit would be cool if Word just had that implemented inside the product like Google Docs does.
- benatkin 2y agoNary a mention of LLMs in the readme. That was an unexpected but pleasant surprise, when the idea of converting something to markdown for LLMs is floated as if it's new and the greatest thing since sliced bread. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=LLM%20markdown&sort=byDate&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... It's interesting to read the code. It's mostly glue code, and most of it is in single 1101 line file. But it does indeed say what the README says it does. Here is the special handling for Wikipedia: https://github.com/microsoft/markitdown/blob/main/src/markitdown/_markitdown.py#L216 https://github.com/microsoft/markitdown/blob/main/src/markit... Edit: good to see the one from yesterday flagged. I tried to assume good intent, but also wondered if it was a place to draw a line in the sand. https://news.ycombinator.com/item?id=42405758 https://news.ycombinator.com/item?id=42405758 Edit 2: ah, it came down to simple violation of the Show HN rules. I didn't notice, but yeah, that's definitely the case.
- zamadatix 2y ago> Nary a mention of LLMs in the readme. That was an unexpected but pleasant surprise No surprise it has still managed to come up in the comments in spite of that!
- benatkin 2y agoYep, and that’s fine. It’s just that there is a lot of false assumptions and magical thinking going around about LLMs and Markdown and I was glad to not find any in the README.
- irskep 2y agoI worked on an in-house version of this feature for my employer (turning files into LLM-friendly text). After reading the source code, I can say this is a pretty reasonable implementation of this type of thing. But I would avoid using it for images, since the LLM providers let you just pass images directly, and I would also avoid using it for spreadsheets, since LLMs are very bad at interpreting Markdown tables. There are a lot of random startups and open source projects who try to make this space sound fancy, but I really hope the end state is a simple project like this, easy to understand and easy to deploy. I do wish it had a knob to turn for "how much processing do you want me to do." For PDF specifically, you either have to get a crappy version of the plain text using heuristics in a way that is very sensitive to how the PDF is exported, or you have to go full OCR, and it's annoying when a project locks you into one or the other. I'm also not sure I'd want to use the speech-to-text features here since they might have very different performance characteristics than the text-to-text stuff.
- cosmie 2y agoFrom your experience, what would be the best way to handle spreadsheets?
- simonw 2y agoI don't think tabular data of any sort is a particularly good fit for LLMs at the moment. What are you trying to do with it? If you want to answer questions like "how many students does Everglade High School have?" and you have a spreadsheet of schools where one of the columns is "number of students" I guess you could feed that into an LLM, but it doesn't feel like a great tool for the job. I'd instead use systems like ChatGPT Code Interpreter where the LLM gets to load up that data programatically and answer questions by running code against it. Text-to-SQL systems could work well for that too.
- btown 2y agoThis is an active area of research: https://github.com/SpursGoZmy/Awesome-Tabular-LLMs https://github.com/SpursGoZmy/Awesome-Tabular-LLMs is a good starting point!
- fritzo 2y agoConverters like this are much more useful if they are bi-directional, even if the two directions aren't exactly inverses.
- theanonymousone 2y agoWhy is the repository 95% "HTML" code?
- sphars 2y agoThere's some very large HTML files in the test directory, including an offline version of the Microsoft Wikipedia page
- caterama 2y agoAnd the core code mostly calls other libraries for heavy lifting -- eg `mammoth`: https://github.com/mwilliamson/python-mammoth https://github.com/mwilliamson/python-mammoth
- valbaca 2y agotests
- markhneedham 2y agoQuite curious how this compares to docling - https://github.com/DS4SD/docling https://github.com/DS4SD/docling docling uses an LLM IIRC, so that's already a difference in approach
- phren0logy 2y agoIn my use, docling has not involved an LLM. There are a few choices for OCR, but I don't think a vision model is one of them. It's certainly touted as a solution to digest documents into plain text for LLM use, but (unless I just haven't run into that part of it) it does not employe an LLM for its functions.
- ekianjo 2y agodocling does not use LLMs...
- simonw 2y agoIf you have uv installed you can run this against a file without first installing anything like this: uvx markitdown path-to-file.pdf (This will cache the necessary packages the first time you run it, then reuse those cached packages on future invocations.) I've tried it against HTML and PDFs so far and it seems pretty decent.
- wrboyce 2y agoIs uvx just part of uv? I keep a few python packages around via pipx (itself via homebrew) but am a big fan of uv for python projects… Do I just need to install uv globally (via brew?) to do this? Is there a mechanism to also have the installed utils available in my PATH (so I can invoke them without a uvx prefix)?
- buibuibui 2y agoWow that is magic! I just installed uv because of your comment.
- figomore 2y agoPandoc (https://pandoc.org https://pandoc.org) can be used to convert a .docx file to markdown and other file formats like djot and typst. I don't think pandoc can convert powerpoint and excel files.
- disgruntledphd2 2y agoYeah that was the interesting part to me, at least. Plus, it's Microsoft so hopefully it will work for their files.
- LordDragonfang 2y ago...I did not catch that it was from Microsoft. I was wondering why a random markdown converter was so notable.
- _rs 2y agoThat was the first thing I checked, and it looks like they’re using some existing python package to parse docx files. I wonder if they contributed to it or vetted it strongly
- disgruntledphd2 2y agoWow, I dunno if that's good or bad, certainly it's not what I expected.
- wis 2y agoLooking at the code, it looks like they used existing Python packages to read and parse MS Office formats, not what I expected, seeing that the repo is in Microsoft's org on GitHub I expected them to have used Microsoft's "official" libraries for parsing these formats, through Component Object Model (COM). They used Mammoth for docx (Word) [1][2] Python-pptx for ppt (PowerPoint) [3][4] and Pandas for XSLX (Excel) [5] [1] https://github.com/microsoft/markitdown/blob/70ab149ff1657c327ebd6ca940988f5d3c5d80d0/src/markitdown/_markitdown.py#L495 https://github.com/microsoft/markitdown/blob/70ab149ff1657c3... [2] https://pypi.org/project/mammoth/ https://pypi.org/project/mammoth/ [3] https://github.com/microsoft/markitdown/blob/70ab149ff1657c327ebd6ca940988f5d3c5d80d0/src/markitdown/_markitdown.py#L539 https://github.com/microsoft/markitdown/blob/70ab149ff1657c3... [4] https://pypi.org/project/python-pptx/ https://pypi.org/project/python-pptx/ [5] https://github.com/microsoft/markitdown/blob/70ab149ff1657c327ebd6ca940988f5d3c5d80d0/src/markitdown/_markitdown.py#L513 https://github.com/microsoft/markitdown/blob/70ab149ff1657c3...
- LittleTimothy 2y agoThis is... interesting. From my understanding - and people can correct me if I'm wrong, but didn't Microsoft spend an extremely large amount of effort essentially trying to screw people who made things like this in the 2000s? Interoperability and the Open Office movement were prety hard fought. It's kind of crazy to see MSFT do this today. Did I just misunderstand and the underlying formats (docx etc) were actually pretty friendly, or have the formats evolved a lot since then? Or is it more a case of "It doesn't matter if it looks terrible because we're feeding it to the AI beast anyway" A cynic might say it became suddenly easy when MSFT had a reason to allow you to genereate markdown to feed into it's AI?
- dmonitor 2y agoI don't think that's a cynical take considering the description > (e.g., for indexing, text analysis, etc.)
- badlibrarian 2y agoMicrosoft filed a covenant not to sue and made all the formats open ~20 years ago. A lot of people bitched at the time but there's a long list of software that supports the format now. It is complicated because the apps themselves are complicated and decades old, and imperfect because the format or app you're converting to likely doesn't support all of the features and certainly none of the quirks. https://en.wikipedia.org/wiki/Office_Open_XML https://en.wikipedia.org/wiki/Office_Open_XML It took browsers 15 years just to render HTML whitespace nearly consistently, so keep that in mind as you read that history.
- btown 2y agoFor PDFs it's entirely a wrapper around https://pdfminersix.readthedocs.io/en/latest/tutorial/highlevel.html https://pdfminersix.readthedocs.io/en/latest/tutorial/highle... - https://github.com/microsoft/markitdown/blob/main/src/markitdown/_markitdown.py#L478 https://github.com/microsoft/markitdown/blob/main/src/markit... So if that's your use case, PDFMiner might be better to integrate with directly!
- kepano 2y agoNever thought I'd see the day. Yet... not surprising because plain text is the ideal format for analysis, LLM training, etc. The question businesses will start to ask is why are we putting our data into .docx files in the first place?
- mdaniel 2y agoI can't tell if you're trolling or what but the idea of most business users (a) knowing markdown (b) reverting to html for the damn near infinite layout and/or styling things that markdown doesn't support (c) ignoring mail merge (d) wanting change tracking ... makes your comment laughable
- throwaway81523 2y agoWhy not Pandoc?
- johannesrexx 2y agoPandoc does not have a PDF reader.
- ulrischa 2y agoI wonder how a powerpoint can be converted to markdown
- poidos 2y agoVery timely, thanks! Was just yesterday working on chaining together `xlsx` and `tablemark` to accomplish this. I found `uvx markitdown my-excel.XLSX | sed 's/ NaN/ /g' my-markdown.md` to be just what I needed to get my spreadsheet into a reasonably-legible markdown table when rendered by GitLab.
- constantinum 2y agoI will try it with some complex layout PDFs or documents with tables. These documents have real business use cases for automation — insurance, banking, etc. Anyone here who wants to convert PDF documents or scanned images as it is preserving the layout, do try LLMWhisperer - https://unstract.com/llmwhisperer/ https://unstract.com/llmwhisperer/
- starkparker 2y agoI index a lot of tabletop RPG books in PDF format, which often have complex visual layouts and many tables that parsers typically have difficulty with. If this is just a wrapper around PDFMiner, as noted in another comment, I don't see any value added by this tool. This handles them... fine. It either doesn't recognize or never attempts to handle tables, which makes it fundamentally a non-starter for my typical usage, but to its credit it seems to have at least some sense of table cells; it organizes columns in a manner that isn't fully readable but isn't as broken as some other solutions, either. It otherwise handles text that's in variable-width columns or wrapped in complex ways around art work rather well. It inserts extraneous spaces on fully justified text, which is frustrating but not unusual, and sometimes adds extraneous line breaks on mid-sentence column breaks. The biggest miss, though, is how it completely misses headings! This seems fundamental for any use case, including grooming sources for LLM training. It doesn't identify a single heading in any PDF I've thrown at it so far.
- hks0 2y agoThis is amazing and really useful, love the idea; but let me tell you a story, it's a bit of a tangent but relevant enough: In an online language class we were sending the assignments to our teacher via slack, the teacher would then mark our mistakes and send it back. I, as a true hater of all the heavy weight text formats for everyday communications, autonomously fired up the terminal, wrote my assignment in my_name.md and happily sent it without giving it any thought. This is what I hear the next session: "... and everybody did a great job! Although someone just sent me their assignment in a stupid format. I don't know what it was! I could neither highlight it or make the text bold or anything. Don't do that to me again please". Before that I never dreamed of meeting someone who preferred a word document _after_ opening a .md file, and I also learned if I had chosen product design as a career, everyone would've suffered immensely (or maybe not, I would've just ended up jobless).
- EasyMark 2y agoIf you are talking about an online language class as in "I'm learning Yiddish" then I don't understand why it would confuse that that someone who isn't a coder or writer (and they're a big if) who doesn't know what the heck markdown is and hence wouldn't want to deal with it since they're used to MS Word or other word processor app. that's probably like 95% of the population at least.
- hks0 2y agoIt doesn't confuse anyone, quite the opposite. The irony for me was my own isolation with the non-tech folks.
- powersnail 2y ago> Before that I never dreamed of meeting someone who preferred a word document _after_ opening a .md file That's like 90% of the people I know outside of computer/engineering circle. Most of people probably have never opened a plaintext file in their life. They would have no idea what to do with a `.md` file. In fact, some older engineers would not know what markdown is either, since it's only been around for two decades or so, but they can probably work with it anyway (the strength of plain text).
- yawnxyz 2y agoanyone get the Bing search DocumentConverter working? It keeps getting me null results
- sneak 2y agoI wish we had a markdown equivalent for spreadsheets. Markdown tables ain’t it.
- acrophiliac 2y agoYour comment has me very curious what exactly you are looking for in a "markdown equivalent" for spreadsheets. Do you want Excel to be able to export the spreadsheet in a Markdown-like format (including formulas, etc)? Or do you want to build the spreadsheet in a text editor using Markdown++ syntax and then use some GUI application to render it? Or do you simply want an ASCII version of Excel that works in a terminal?
- sneak 2y agothe second one. I want a human editable plain text spreadsheet format. CSV and TSV ain’t it, not the least of which is because they don’t have formulae.
- zelphirkalt 2y agoOrg-mode. Emacs Org mode has tables with formulas, being able to make use of many programming languages. By default Calc (I believe GNU Calc) and Elisp. However using sbe you can make it use code blocks written in any language that you have support for using org-babel. For example I have time tracking spreadsheets using source blocks of GNU Guile code for time calculation. Of course you can put that under version control easily, since it is just a text document.
- einpoklum 2y agoThis is BS, it doesn't support Office documents, it supports only Microsoft's broken office documents which don't obey their own custom specs. Why doesn't this work on ODF files?
- lbrunson 2y agoAre there any good libraries for the opposite, going from markdown to pdf or docx? Pandoc gets most of the way there but struggles with certain things like tables.
- roamerz 2y agoSince it’s Microsoft maybe it will do a half decent job on Outlook HTML and .docx. I have evaluated most of them out there, paid included and haven’t found one that I thought was good enough to run in production. Definitely will be giving this a try.
- be_erik 2y agoOh thank god. I can finally retire my docx to pandoc to markdown tool chain. I can’t believe M$ was the big one to go first. Good on ya.
- toastal 2y agoSo we convert from rich formats with metadata & advanced features to a format without the former & severely lacking at the latter.
- konfekt 2y agoThough it promises to convert everything to Markdown, it seems to be a worse version of what the already existing tools such as PDFtotext, docx2txt, pptx2md, ... collected [here] do without even pretending to export to Markdown. Looking at its [source], it indeed seems to be a wrapper to python variants of those. Making the pool smaller can hardly improve the output. [here] https://github.com/Konfekt/vim-office https://github.com/Konfekt/vim-office [source] htps://github.com/microsoft/markitdown/blob/main/src/markitdown/_markitdown.py
- SuperHeavy256 2y agoI don't think it works if you try installing it using pip. Can anyone confirm? I ended up downloading it manually, making a venv, and then running it.
- ekianjo 2y agoany idea how it compares to Docling?
- zelphirkalt 2y agoIf the source document is anything half decent, this would serve to lose information, as markdown is far from flexible and powerful enough to represent all kinds of formatting and layout present in source documents. If all you need is the text information, then that might be just what you want, lossily compressing documents.