4 ms·
Show HN: ArXiv-txt, LLM-friendly ArXiv papers
Just change arxiv.org to arxiv-txt.org in the URL to get the paper info in markdown
Example:
Original URL:
https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762
Change to:
https://arxiv-txt.org/abs/1706.03762 https://arxiv-txt.org/abs/1706.03762
To fetch the raw text directly, use https://arxiv-txt.org/raw/abs/1706.03762 https://arxiv-txt.org/raw/abs/1706.03762, this will be particularly useful for APIs and agents
- lgas 2y agoIt just extracts the abstracts?
- jmartin2683 2y agoThis would be awesome wrapped in an MCP server/tool call :)
- jerpint 2y agowhoa - i haven't yet played with MCP - might be a good first project!
- westurner 2y agoIf you train an LLM on only formally verified code, it should not be expected to generate formally verified code. Similarly, if you train an LLM on only published ScholarlyArticles ['s abstracts], it should not be expected to generate publishable or true text. Traceability for Retraction would be necessary to prevent lossy feedback.
- sbpost 2y agoThe example you give doesn't seem to work - the raw txt does not have authors.
- jerpint 2y agoyou're right - I hadn't noticed! I fixed it now, thanks for pointing it out
- cchance 2y agoWas super excited that it was going to be the actual papers, kinda cool but just being abstracts doesn't go very far, good luck getting the papers working thats gonna be pretty cool once working, then to feed it all into a vector db XD
- owalerys 2y agoReally clean API design, I'm a fan!