23 ms·
ArXiv now offers papers in HTML format
- wolverine876 3y agoMany here say they prefer html documents. How do you annotate them? How do you make local copies? Also, how will you read them in the decades to come? I love PDF.
- injuly 3y agoFor anyone who needs it, arxiv-vanity is amazing: https://www.arxiv-vanity.com/ https://www.arxiv-vanity.com/
- westurner 3y agoarxiv-sanity-lite: https://github.com/karpathy/arxiv-sanity-lite https://github.com/karpathy/arxiv-sanity-lite
- 101008 3y agoIs there an open source tool to convert any PDF to something like this?
- mcpherrinm 3y agoIt sounds like (from the shout-out in the post) they're using https://math.nist.gov/~BMiller/LaTeXML/ https://math.nist.gov/~BMiller/LaTeXML/ to convert the paper's LaTeX into HTML, not from PDF. The most versatile tool I know of for converting various document formats, including PDF to HTML, is the oss ebook tool Calibre: https://manual.calibre-ebook.com/conversion.html https://manual.calibre-ebook.com/conversion.html I have seen https://pdfbox.apache.org/ https://pdfbox.apache.org/ used for extracting text from PDFs for analysis, but you won't get HTML output.
- hk-senokr 3y agoGive it to the United States 2 minutes you're open and your hack smoker Hancock minutes and even this your combustion area of monument time cuz you said looking on the baseball miserable I didn't want to buy it I desktop and your current events are my not me his not he I took it in the garage your prime minister 70 or my event your lucky alone haircut at Josephine alone hacker smoker king Kong young under hackers no car orange county Joseph Adidas adorius avenue I got a new Nissan I thought you need something f*** at Robert Omaha Fernandez Serbia Yunnan i England England Britannia English
- shrimpx 3y agoSince the article doesn't link to any example HTML article, here's a random link: https://browse.arxiv.org/html/2312.12451v1 https://browse.arxiv.org/html/2312.12451v1 It's cool that it has a dark mode. Didn't see a toggle but renders in the system mode. Overall will make arXiv a lot more accessible on mobile.
- burkaman 3y agoAnd here's the PDF of the same paper for comparison: https://arxiv.org/pdf/2312.12451.pdf https://arxiv.org/pdf/2312.12451.pdf
- FredPret 3y agoThe contrast is massive. I'm much more likely to read the html version; that PDF is deeply off-putting in some hard to define way. Maybe it's the two columns, or the font, or the fact that the format doesn't adjust to fit different screen sizes.
- lemper 3y agodefo concur. will read the html version when on mobile from now on.
- ForkMeOnTinder 3y agoDefinitely the two columns for me. It's super annoying skimming a paper and having to scroll down and back up again in a zig-zag pattern.
- mmis1000 3y agoI think the consuming device matters. A ipad or computer have much wider screen width. One column layout is too wide for them for average people to scan text lines quickly. While it looks perfectly fine on a phone. Two columns layout looks terrible on a smartphone, the text is too tiny to read comfortably. It would probably be even better if you can flip it left and right like a ebook instead of scrolling to allocate the content faster. But current design is good enough IMO. (Compare to reading a pdf on cellphone)
- alephnerd 3y agoThis is a great UX addition. Why did it take them so long?
- gwern 3y agoThe conversion is still very error-prone. It can't convert a lot of packages, and the last paper I read, StarVector, half the HTML version is just missing. (I think it hit an error at a figure of some sort.) I reported an error, but I've been reporting errors against the ar5iv and abstracts for years now and the long tail of problems just seems like an incredible slog.
- KRAKRISMOTT 3y agoWhere are the computer vision people? This is the perfect type of problem for multi modal LLMs
- IlliOnato 3y agoExcept that the errors made by an LLM might be harder to spot then converter errors that typically are very blatant, and don't usually alter text (perhaps just drop parts of it). Also, a bug in a converter is conceptually much easier to fix than to re-train your LLM. I am not sure that AI in it's current state is useful when "high fidelity" is required.
- dginev 3y agoCan confirm. From an ar5iv standpoint, 2.56% articles currently fail to convert entirely, and 22.9% have known errors to the converter. That leaves 74.5% of nominally usable articles. This success rate is noticeably lower for the newest batches of arXiv submissions, as the converter hasn't caught up with the most recent package innovations. We have a plan in place to meaningfully fall back for unknown packages, but that will take at least another year to put in place, and likely another couple of years to stabilize. Meanwhile, there is some hope that with arXiv launching the HTML Beta we will get more contributions for package support (LaTeXML is an open source project, with public domain licensing, everybody benefits). But again the original point is spot on. Coverage will be hit-or-miss for a while longer yet, for an arbitrary arXiv submission. The good news is that authors could work towards better support for their articles, if they wanted to.
- binarymax 3y agoNice! Now I don’t need to manually replace arxiv with ar5iv. Congrats to the team.
- imjonse 3y ago"Our ultimate goal is to backfill arXiv’s entire corpus so that every paper will have an HTML version, but for now this feature is reserved for new papers." For now it only works for papers submitted this month. But it's great to have this feature, makes it so much easier to read on phones.
- eviks 3y agoFinally a modern format you can copy&paste from and read on one of the most popular computing platforms!!!
- shusaku 3y agoSeems like the references aren’t working very well. I really want journals to have two way links in a paper. I get google scholar alerts about certain papers being cited, and I want to skip to “why did they cite this? Did they use it, improve it, it just mention it?”
- r3trohack3r 3y agoI’d never considered setting up citation alerts like this. Thank you for the idea!
- shrimpx 3y agoLooks like clicking a reference adds the hash to the URL but doesn't scroll to the reference. If you load the hash URL directly in the browser you get a 404 page...
- burkaman 3y agohttps://browse.arxiv.org/html/2312.12451v1#bib.bib1 https://browse.arxiv.org/html/2312.12451v1#bib.bib1 works, but https://browse.arxiv.org/html/2312.12451v1/#bib.bib1 https://browse.arxiv.org/html/2312.12451v1/#bib.bib1 doesn't.
- pushfoo 3y agoPreviously discussed: https://news.ycombinator.com/item?id=38713215 https://news.ycombinator.com/item?id=38713215
- carlosjobim 3y agoWith the 2024 browser update, this means I can read these articles on my ancient Kindle perfectly fine.
- ChrisArchitect 3y ago[dupe] from yesterday More here: https://news.ycombinator.com/item?id=38713215 https://news.ycombinator.com/item?id=38713215
- winwang 3y agoProbably more accessible in general. (PDF) Papers are psychologically scary.
- mmis1000 3y agoPdf is by design a image format that can also embed text. It just don't have the primitives to properly retain the article structure.
- PaulHoule 3y agoNah, it's a super-complex system that creates a graph of components, can draw vectors like PostScript, can embed 3-d models, etc. The spec is here https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandards/PDF32000_2008.pdf https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandard... if you look at sections 14.6 through 14.10 you will find quite baroque facilities for representing the structure of documents in great detail, making documents with accessibility data, making documents that can reflow with HTML, etc. Note to mention the 14.11 stuff which addresses problems with high end printing (say you want to make litho plates for a book.) For that matter sections 14.4 and 14.5 describe facilities that can be used to add additional private data to PDF files for particular applications. For instance Adobe Illustrator's files are PDF files with some extra private data, and https://en.wikipedia.org/wiki/GeoPDF https://en.wikipedia.org/wiki/GeoPDF I like to complain that PDF has no facility to draw a circle but instead makes you approximate a circle with (accursed) Bézier curves but other than that the main complaint people make about PDF is that it is too complicated not that it is lacking this feature or that feature. Contrast that to a highly opinionated document format like DjVu https://en.wikipedia.org/wiki/DjVu https://en.wikipedia.org/wiki/DjVu which came out around the same time as PDF and is specialized for the problem of scanned documents and works by decomposing the document into three layers, one of which is a bilevel layer intended to represent text. All three layers have specialized coding schemes, the text layer in particular tries to identify that every copy of (say) the letter "e" or the character "漢" is the same and reuse s the same bitmap for them.
- anonimo37 3y ago
- ZeroCool2u 3y agoWow, this is _so_ much better!
- choppaface 3y agoHope they benefit from CDN caching now too. Edit: aaaand they got Fastly https://news.ycombinator.com/item?id=38723373 https://news.ycombinator.com/item?id=38723373
- cozzyd 3y agodoesn't work great with long author lists... https://browse.arxiv.org/html/2312.12907v1 https://browse.arxiv.org/html/2312.12907v1
- degenerate 3y agoThe PDF is worse, so there is no simple answer to this: https://arxiv.org/pdf/2312.12907v1.pdf https://arxiv.org/pdf/2312.12907v1.pdf At least the HTML version pairs each author with their affiliations, instead of the PDF which has all the names on page 1, and all the affiliations on page 2. That's completely unreadable.
- cozzyd 3y agoThe PDF is better because I'm trained to scroll past the author list. That takes forever on the html version .
- mattigames 3y agoYou can click the "Introduction" anchor on the left side and it scrolls for you past the author list
- cozzyd 3y agowell it skips the abstract too, but yes, you can scroll back up to see it.
- mattigames 3y agoYeah, its a bit weird that the abstract doesn't have a link on the left
- cozzyd 3y agoProbably because \abstract{ } is treated differently than \section{ }, I guess...
- Al-Khwarizmi 3y agoNice! It would be even better if they offered authors of previous papers the option of converting to HTML, as the latex sources are already in the system.
- fprog 3y agoThe article states they're going to backfill all, or nearly all, previously submitted papers!
- FredPret 3y agoThis is brilliant. I don't share academia's love of LateX multi-column PDFs.
- tiagod 3y agoI like multi-column text on paper (literally), but it's awkward in digital where you can just shape text on the fly to whatever column size you want
- golol 3y agoThe oroblem is that gaining this responsiveness fundamentally makes your task much more difficult. Instead of just creating a picture you're now writing code that has to be maintained. In my philosophy arxiv is for documents which are set in granite - pictures.
- deleted 3y ago[deleted]
- leoncaet 3y agoI just hope they don't stop to offer the papers in PDF. Even when I'm on a computer, I still prefer to read PDFs.
- creatonez 3y agoThere is a taste component to it of course, but the history of PDF shows that it's the wrong format for reading on a computer. It was originally meant to be the end result of a publishing process before printing, a layer that sits right between the publishing software and the postscript that gets sent to the printer. This makes the PDF format quite inflexible for reading on a computer, with it being impossible to properly zoom or adjust the reading experience. Unfortunately many institutions and businesses have ignored its limitation because PDF turned out to be an obvious-but-naive to put a 'sheets of paper' metaphor into a computer system, which in the 1990s appealed to tech illiterate folks doing bare-bones computerization of existing paper systems. So later we got complicated and error-prone tools for editing PDFs, and many random additions to the spec to allow for unusual use cases.
- impendia 3y ago> This makes the PDF format quite inflexible for reading on a computer, with it being impossible to properly zoom or adjust the reading experience. As an academic researcher, generally speaking I also prefer PDF, and the inflexibility and static nature is a feature, not a bug. I appreciate the fact that a paper will appear the same everywhere, that I can refer to "the top of page 7", etc. The exception is if I wanted to just skim a paper; in this case, I think I'd prefer HTML. I'm a huge fan of what arXiv is doing here. It effectively preserves the status quo, while adding an additional option on the side. The HTML option might prove a little bit useful for me, and it is likely to prove extremely useful for people with disabilities.
- creatonez 3y ago> I appreciate the fact that a paper will appear the same everywhere, that I can refer to "the top of page 7", etc. There are many great solutions to this problem, including ones that don't require Javascript at all. This website (https://gwern.net/silk-road https://gwern.net/silk-road) presents a really good example -- every header and sub-header is a clickable anchor. If more granularity is needed, on newer articles most of the paragraphs start with an italicized margin note -- though for technical writing, paragraph anchors might be better. The page also pays careful attention to print CSS and has a 'reader mode' to convert all links to footnotes when printed. Some websites will also preserve the text you select in a URL anchor, but more often than not this is just cumbersome. It also has a greater risk of not surviving changes to the webpage.
- sylware 3y agoLike the maths noscript/basic (x)html wikipedia generator: The magic of inline images at a known DPI, of course you can provide images for different DPIs. Reading maths/science noscript/basic (x)html documents on my 100 DPI monitor, on wikipedia. Not yet fully ready on arxiv.
- gms7777 3y agoAbout time. Biorxiv and medrxiv have been doing this for probably half a decade at this point?
- dginev 3y agoWrong, arXiv was first. Check this HTML paper from 1997: https://arxiv.org/html/astro-ph/9708066 https://arxiv.org/html/astro-ph/9708066
- cbf66 3y agomedRxiv and bioRxiv get most of their submissions as Word files. It's a much easier conversion, and if necessary they have manual touch-up. Not feasible for arXiv's volume.
- jez 3y agoIt would be neat if they offered submitters the chance to upload their own HTML version alongside the PDF version, instead of always relying on an automatic conversion process. - I can imagine authors feeling frustrated if someone reaches out about a problem in the HTML version of their paper, but they have no way to correct it except by hoping that a change to the PDF fixes a change to the generated HTML. Easier to just fix the formatting problem in the PDF outright. - It would be neat to allow people to experiment with alternative formatting for their papers. For example, imagine a paper about a programming language that embeds a sandbox you can use to play around with the language under discussion. Or a paper about multivariable calculus and you can interact with a three dimensional plot of some function.
- layer8 3y agoThey’d have to define and document a “safe” subset of HTML, and implement a filter/checker for it. Otherwise we’d end up with papers containing ads and tracking and XSS vulnerabilities and whatnot.
- digging 3y agoThose are issues with JavaScript, not HTML. Wouldn't filtering out iframes pretty much keep us in the clear?
- layer8 3y agoThe parent wanted interactive 3D plots, which means JavaScript embedded in or linked from the HTML. Then there‘s stuff like JavaScript embedded in SVG.
- CaptainOfCoit 3y ago> Those are issues with JavaScript, not HTML What about various HTML tags that remote load resources? From script, link, to things like img or CSS `background-image` attribute, added in a `style` attribute. There is a bunch of ways to do remote requests even without HTML.
- endergen 3y agoI was hoping this meant that html native submissions would be possible, so that people made interactive explanations.
- lucidrains 3y agonice! will make reading papers on the phone so much more pleasant!
- tarboreus 3y agoOne of the reasons is to make the papers more accessible to people with disabilities, especially the blind. I participated in a conference they hosted on this a few months ago, I recommend taking a look at the recordings if you're interested in thinking on this. https://accessibility2023.arxiv.org/ https://accessibility2023.arxiv.org/
- miki123211 3y agoBlind person here, can confirm this. Reading PDFs with a screen reader is bad, reading PDFs that come from LaTeX is worse, reading LaTeX math is pretty much impossible. All the semantic info you need is just thrown away. You can make decently accessible PDFs but it's lots of work, you need Acrobat on the producer' side and might also need it on the consumer's side. Free tools don't even come close. There's also the fact that the process of making accessible PDFs in Acrobat isn't itself accessible. With that said, the way screen readers treat HTML math certainly isn't perfect, it's geared more towards school children than anything above calculus. I'm probably going to stay with my LaTeX source files for now. At least ArXiv offers those, not many sites do. To be fair, that approach also has its own set of problems (particularly when people use some extra fancy formatting in their math equations, making the markup hard to read), but I find this to be the best approach for me so far, at least on AI/ML papers.
- saurik 3y agoHuh. It would seem like, of all the things which should make it easy to generate the correct accessibility information, the pipeline of compiling a paper from source code in LaTeX should nail it... maybe we should all pitch in to some pool to pay someone to put in the required effort to connect all the dots?
- semi-extrinsic 3y agoKind of tangential, but it's also kind of surprising how difficult it is in LaTeX to make a plot of an equation. Say I have Equation \ref{eq}. Why can't I just say "plot \ref{eq} for x from -6 to 11" and get my graph? And yes, I know about pgfplots, PSTricks, TikZ etc. But in all those cases, I need to define the same equation twice, in different syntax to boot. It's kind of unsatisfying.
- odyssey7 3y agoarticle { text-justify: Knuth-Plass; }
- SushiHippie 3y agoMind explaining?
- odyssey7 3y agoThe comment is invalid CSS to apply the Knuth-Plass algorithm in rendering an HTML article. Knuth being a perfectionist’s perfectionist, TeX uses this algorithm to determine optimal line breaks to provide for better text justification. Here’s a discussion of hacks to achieve the algorithm’s results on web pages and an upcoming CSS feature as of 2020. https://mpetroff.net/2020/05/pre-calculated-line-breaks-for-html-css/ https://mpetroff.net/2020/05/pre-calculated-line-breaks-for-...
- SushiHippie 3y agoThank you!
- computerfriend 3y agoIf only.
- matt1 3y agoFor anyone interested in staying informed about important new AI/ML papers on arXiv, check out https://www.emergentmind.com https://www.emergentmind.com, a site I'm building that should help. Emergent Mind works by checking social media for arXiv paper mentions (HackerNews, Reddit, X, YouTube, and GitHub), then ranks the papers based on how much social media activity there has been and how long since the paper was published (similar to how HN and Reddit work, except using social media activity, not upvotes, for the ranking). Then, for each paper, it summarizes it using GPT-4, links to the social media discussions, paper references, and related papers. It's a fairly new site and I haven't shared it much yet. Would love any feedback or requests you all have for improving it.
- raccoonDivider 3y agoThat looks great. No real feedback yet, but it's the kind of thing I've always been looking for as a better alternative to Twitter.
- matt1 3y agoThanks! I've got a lot more planned for it too. If anyone has any feedback that doesn't make sense to share here, or if you're a researcher who is open to some questions about how you currently follow arXiv papers, drop me a note at matt@emergentmind.com.
- CodeCube 3y agoLove to see Energent Mind continuing to innovate!
- sureglymop 3y agoLove the clean design of the website! Looks amazing on mobile.
- matt1 3y agoThanks! If you ever run into any issues or have any suggestions for improving the site, drop me a note: matt@emergentmind.com.
- apstats 3y agoI wonder if this could be used to train an LLM to convert PDFs with rich charts into HTML?
- reqo 3y agoA lot of AI/ML papers these days have an accompanying interactive page like [0], will we see anything like these now directly in arXive? [0] https://voyager.minedojo.org/ https://voyager.minedojo.org/
- z2h-a6n 3y agoI think then arXiv would have to deal with mantaining the tech stack and providing the presumably much higher server capacity to serve the more varied web pages that would result, so it seems like a tall order. arXiv already has an experimental integration with Papers with Code [0], which I guess provides similar results for the reader, though the authors have to figure out their own web hosting. [0] https://info.arxiv.org/labs/showcase.html#arxiv-links-to-code-data https://info.arxiv.org/labs/showcase.html#arxiv-links-to-cod...
- MahiShafiullah 3y agoSecond that. Something I put out recently had an (admittedly video heavy) webpage that had 1TB of traffic over the past month. Cloudflare handled it for free for me, but at ArXiv’s scale it’s bound to be a problem.
- ansk 3y agoWhen I open a large pdf on arxiv (100+ MB, not uncommon for ML papers focused on hi-res image generation), there is a significant load time (10+ seconds) before anything is rendered at all other than a loading bar. Does anyone know what the source of this delay is? Is it network-bound or is Chrome just really slow to render large PDFs? Do PDFs have to be fully downloaded to begin rendering? In any case, this delay is my only gripe with arxiv and a progressively rendered HTML doc that instantly loads the document text would be a huge improvement.
- IlliOnato 3y agoIt may be even that the time is taken to generate a PDF. The format in which articles are submitted and stored in arXive is LaTeX. PDF is automatically generated from it. Probably arXiv does some caching of PDFs so they don't have to be generated anew every time they are requested, but I don't know how this caching works.
- upbeat_general 3y agoI have the same issue. From what I can tell it’s just network-bound and the Arxiv servers are slow. They theoretically allow for you to setup a caching server but after spending a while trying to get it setup, I haven’t been able to get it to work. https://info.arxiv.org/help/faq/cache.html https://info.arxiv.org/help/faq/cache.html
- arccy 3y agomaybe it'll be faster now with fastly https://news.ycombinator.com/item?id=38723373 https://news.ycombinator.com/item?id=38723373
- 10000truths 3y ago> Does anyone know what the source of this delay is? Is it network-bound or is Chrome just really slow to render large PDFs? Do PDFs have to be fully downloaded to begin rendering? In any case, this delay is my only gripe with arxiv and a progressively rendered HTML doc that instantly loads the document text would be a huge improvement. The default PDF format puts the xref table at the end of the file, forcing a full download before rendering can take place. PDF-1.2 onwards supports linearized PDFs, and most PDF export tools have some way of enabling it (usually an option like "optimize for web").
- ww520 3y agoThat's great. Now I can read the papers on my phone.
- svag 3y agoThe tool that it's being used for this offering is this one, https://github.com/arXiv/arxiv-readability https://github.com/arXiv/arxiv-readability, just to save a few clicks :)
- IshKebab 3y agoWow I did not know they have the LaTeX for all the papers and compile it themselves! That's pretty crazy. What if they don't have packages you need? What if your paper isn't written with LaTeX?
- deleted 3y ago[deleted]
- r4indeer 3y ago> What if they don't have packages you need? Unlikely. But if so, you can provide the packages yourself: https://info.arxiv.org/help/submit_tex.html#wegotem https://info.arxiv.org/help/submit_tex.html#wegotem > What if your paper isn't written with LaTeX? Then they still accept PDF or HTML. See: https://info.arxiv.org/help/submit/index.html#formats-for-text-of-submission https://info.arxiv.org/help/submit/index.html#formats-for-te...
- aragilar 3y agoThey specify what version of texlive they use. This is significantly better than what publishers offer (usually a really old latex version, not even pdflatex).
- deleted 3y ago[deleted]
- ofou 3y agoI wonder how better is this compared to Pandoc's
- dginev 3y agoThat's it in spirit, but in practice it's refreshed: https://github.com/arXiv/arxiv-view-as-html https://github.com/arXiv/arxiv-view-as-html
- WendyTheWillow 3y agoI’m so far left wanting for an app that gives me a way to easily track and consume newly published work of a given topic. The existing apps are not great, and maybe this change will make it easier to provide better “reader” views, and possibly even tts (I like to listen+read).
- codethief 3y agoUgh. I don't belong to the target audience (people with disabilities) but the typesetting doesn't exactly look pleasant on my machine (Chrome on Linux).
- aragonite 3y agoA lot of academic journals (say from Springer) also offer HTML formats for papers published in the past decade or so, which I personally often find more convenient for reading purposes than PDFs. For example, I parse text a lot faster if I use a regex to split each paragraph into sentences and place a linebreak after each sentence, or if I do natural language "syntax highlighting" by assigning a distinctive color to functional words indicating logical structure like 'if/then', 'and', 'or', 'not', 'because', and 'is'. And sometimes it really improves readability to be able to do "semantic highlighting", in the sense of say assigning a different hashed color to each proper name (or each labeled thesis, etc) that occurs in the paper. Such manipulations are basically impossible with PDFs. It makes me wish sci-hub would start archiving HTML versions in addition to PDFs!
- johnsillings 3y agohttps://www.arxiv-vanity.com/ https://www.arxiv-vanity.com/
- jakderrida 3y agoAnd, of course, https://ar5iv.labs.arxiv.org/html https://ar5iv.labs.arxiv.org/html However, ar5iv isn't a la carte like arxiv-vanity. They pretty much do last month's papers every month or so. Something like that.
- dginev 3y agoHi, ar5iv creator here. You can think of both arxiv-vanity and ar5iv as the "alpha" experiments that lead into the official arXiv "beta" HTML announced today. Once a few rounds of feedback and improvements are integrated, and the full collection of articles acquires HTML in the main arXiv site, ar5iv will be decommissioned. The plan is to turn all existing ar5iv links into redirects to the official HTML, and free up the resources for maintaining it. I am not sure what are the plans for maintaining arxiv-vanity, but I suspect they may head down a similar path some time later.
- jakderrida 3y agolmao! The actual creator of ar5iv? Sometimes I forget this isn't reddit and legit accomplished people comment here. Reminds of Burning Man when people kept telling me, "Never talk trash on the art at the main landmarks. The artists are frequently within listening distance." So, of course, I'd walk around talking about buying the art for $50K-$60k, knowing it's already scheduled to be burned with the landmark.
- philipashlock 3y ago30 years after HTML was invented to support accessibility and collaboration for research and academia and the same day the White House released their new accessibility guidance which happens to be the first time they've published formal new policy natively has HTML rather than PDF - https://www.whitehouse.gov/omb/management/ofcio/m-24-08-strengthening-digital-accessibility-and-the-management-of-section-508-of-the-rehabilitation-act/ https://www.whitehouse.gov/omb/management/ofcio/m-24-08-stre...
- murphyslab 3y agoI feel surprised by how succinct, easy-to-understand, and sensible the policy (M-23-22) is: > Default to HTML: HyperText Markup Language (HTML) is the standard for publishing documents designed to be displayed in a web browser. HTML provides numerous advantages (e.g., easier to make accessible, friendlier to assistive technology, more dynamic and responsive, easier to maintain). When developing information for the web, agencies should default to creating and publishing content in an HTML format in lieu of publishing content in other electronic document formats that are designed for printing or preserving and protecting the content and layout of the document (e.g., PDF and DOCX formats). An agency should develop online content in a non-HTML format only if necessitated by a specific user need. https://www.whitehouse.gov/omb/management/ofcio/delivering-a-digital-first-public-experience/ https://www.whitehouse.gov/omb/management/ofcio/delivering-a...
- wolverine876 3y agoHmmm ... accessibility is essential, but PDF is far better for static documents: There's no straightfoward, standard way to read an html document on another platform. Also, the html document may not be readable in 10+ years (unlike most PDFs), and updates are too fluid and hard to track. I think the general problem is that the end-user doesn't control an html document, e.g., for annotation, as a local record, etc.
- shakow 3y ago> There's no straightfoward, standard way to read an html document on another platform. What do you think of the epub format?
- jll29 3y agoIt's a cool feature because it makes the papers more finable, more easily navigatable, easier to read online and faster to scroll through. I am also happy for blind people that they can more easily use ArXive with Braille readers now. (I'm still a fan of printing the PDFs, because I annotate on paper and refer to page numbers, but the HTML feature is in addition to PDF download, not a replacement.) One thing that still sucks (not ArXiv related though) is reading mathematical formulae on the Kindle - wonder if someone with rendering expertise could have a look into the MOBI format.
- isaacfung 3y agoThis would never happen but in an ideal world, we should be able to click on a citation to jump to the part of the paper that is being referenced and each paper page should have a discussion board so we can easily communicate with the authors and group the discussion in one place instead of us having to google to see if there is relevant discussion on twitter/reddit. We can even put links to talks, tutorials, blogs, github repo, demo, paperswithcode/google scholar/open review, background material, a timeline of citations in tree form on the same page(actually I am seeing more machine learning papers that have a project page that does some of these) or even turn it into a mini wiki. I just think html has so much more potential(especially now with LLM we can do semantic search). I wonder if there would be interest in such a chrom extension overlay. Related projects: https://github.com/ahrm/sioyek https://github.com/ahrm/sioyek https://github.com/arxiv-vanity/engrafo https://github.com/arxiv-vanity/engrafo https://github.com/dginev/ar5iv https://github.com/dginev/ar5iv https://academ.us/article/2111.15588/ https://academ.us/article/2111.15588/ (powered by https://github.com/jgm/pandoc https://github.com/jgm/pandoc I believe)
- me_jumper 3y agoI think https://web.hypothes.is/ https://web.hypothes.is/ would be of interest to you.
- golol 3y agoIMO pdf and HTML optimize for different things. pdf is easy and pretty. HTML is easy and responsive. But making pdf responsive is impossible and making HTML pretty is not easy. I think having arxiv for well-polished pretty documents, not responsive ugly documents. Most researchers don't have time to make an HTML responsive and pretty.
- querez 3y agoAm researcher, care about responsiveness way more than pretty. I am super glad for the option. Downloading PDFs is super annoying. I'm stoked.
- mmis1000 3y agoWell... download html is even harder nowadays, because many pages are dynamically generated. Although there are surely some browser extensions that can help you to finish it in a few clicks..
- radicalriddler 3y agoFUCK YES (excuse my profanity). I have a tool that converts HTML to Neural Speech and I always wanted to push arXiv papers through it, but couldn't be bothered with a PDF implementation.
- topicseed 3y agoWhat do they use to convert a PDF document to a clean, correct HTML document? It's a difficult space, especially with the variety of layouts you may find in PDF documents...
- blackbear_ 3y agoArxiv encourages users to submit the latex source of their papers rather than the PDF
- SushiHippie 3y ago> The tool that it's being used for this offering is this one, https://github.com/arXiv/arxiv-readability https://github.com/arXiv/arxiv-readability, just to save a few clicks :) https://news.ycombinator.com/item?id=38726582 https://news.ycombinator.com/item?id=38726582
- vegabook 3y agoPDF is objectively much better than HTML at rendering text documents. And it's not even close. This could easily have been done 10, even 15-20 years ago. That it didn't is not just inertia. Latex and PDF have enormously better text rendering, and the static format locks a state-commit in time that is much easier to go back to and reference/critique. Unlike the intrinsically fluid nature of HTML. For academic work, milestone-like formats, that lock state in time, are useful for those who later build on them. And again, the rendering just doesn't compare and that imparts [sub]conscious quality signals.
- imranq 3y agoAt this point are academic papers simply peer-reviewed blog posts?
- acjohnson55 3y agoThis is great! I browse papers on mobile, and PDF is so bad for that use case.
- alecsm 3y agoI don't read many papers but this makes it easier for me to save them in Joplin.
- hollerith 3y agoI'm sad that the best they can do is HTML format. HTML is a mess.
- nojvek 3y agoOMG. This is amazing. I legit hated reading two column pdfs on a smartphone.
- deleted 3y ago[deleted]
- wildpeaks 3y agoVery good decision, always bet on the web.
- sicariusnoctis 3y agoPersonally, I would prefer the conventional Latin Modern math font instead of Palatino math. Latin Modern is used by: - Wikipedia. - Math.StackExchange. - Nearly all papers, including the ones hosted on arxiv in PDF format. - Nearly any math videos, slides/presentations, notes. - Almost everything, really. Palatino just looks weird. Also, I imagine that authors might do math formatting hacks that were only tested on Latin Modern, and might end up breaking on Palatino. TL;DR: Palatino :( Latin Modern :)
- IHLayman 3y agoFun fact: if seems that if you use Lockdown mode on Apple devices you can't open PDFs from a browser (no official documentation says it but there is anecdotal evidence). This would allow people with Lockdown mode to open Arxiv papers more easily.
- matrix2596 3y agothats great news. I was using arxiv vanity to read on mobile phones. I am not seeing it on all articles, is it only for new papers?
- therealmarv 3y agoThis is the reason I've never liked LaTeX from a data point view. It's made to be printed out or get to look beautiful on a PDF but was never designed to get you to a HTML file or a Word file. I've written my thesis in Markdown in the past because of this (best for humans) which can be easily transformed to HTML, Word, PDF and even LaTeX https://github.com/tompollard/phd_thesis_markdown https://github.com/tompollard/phd_thesis_markdown And I think that XML is the best format for machines.
- delhanty 3y ago> If you are familiar with ar5iv, an arXivLabs collaboration, our HTML offering is essentially bringing this impactful project fully “in-house”. Our ultimate goal is to backfill arXiv’s entire corpus so that every paper will have an HTML version, but for now this feature is reserved for new papers. IIRC, ar5iv was created on his own initiative by Deynan Ginev https://twitter.com/dginev/status/1736792316675825981 https://twitter.com/dginev/status/1736792316675825981 and it seems that he has worked tirelessly to fix nearly all of the edge cases during the collaboration. This project creates huge value to humanity so Deynan is to be heartily thanked.
- dginev 3y agoThanks for the kind words, but some corrections: 1. My name is Deyan (hi!) 2. ar5iv was the latest frontend incarnation, but our actual work on converting LaTeX to HTML goes back nearly 20 years behind the scenes. 3. I was an undergraduate student when I was introduced to the project back in 2007. It was started "in spirit" by 3 senior co-conspirators back then: Michael Kohlhase, Bruce Miller and Robert Miner. And I am by no means a solitary actor today, even if I may be the chief online presence of the people involved. Bruce is doing the bulk of the hard work on LaTeXML to this day. I documented some of the history in an invited talk for CICM 2022, which you can find on youtube, or see the slides at: https://prodg.org/talks/welcome_to_ar5iv https://prodg.org/talks/welcome_to_ar5iv It's really great that the HTML has now reached "home base" in arXiv, and I hope their team gets a lot more of the positive attention going forward - today's achievement is entirely theirs!
- ngcc_hk 3y agoWent through it and may I ask whether there is any “personal” level of this ar5iv converter or just one of few mentioned parser. Btw given we are into quotation academic world, I wonder whether you may have mention Gartner Group to invent that technology curve. To be honest there is a variation I like more which deal with the chasm issue.
- indrora 3y agoI remember stumbling upon your work long ago when I was working on a project to have "e-zines" that consumed a series of `article` class files and rendered them out into PDF and HTML as a series package. I had come across latex2html, Dan Gildea's project, and found myself unpleasantly dissatisfied with how it worked. As I understand it, it's more a "half implementation of lots of packages" rather than what ar5iv seems to be, which is "enough of the core LaTeX engine producing HTML instead of DVI"? I'd love to know more about the nitty gritty of how the engine does its thing. I'm curious: How has modern web tech (e.g. WebAssembly, Canvas, etc) helped or gotten in the way of getting good LaTeX rendering in the browser?
- trostaft 3y agoTaking a look at a paper I have that went up this month and another that went up before the dec cutoff on ar5iv, they look 90% OK! Figures with side-by-side plots and algorithm environments are the common culprit for being broken though. Particularly in figures, it seems like the width argument isn't being interpreted correctly. Interestingly this review paper seems to have their side by side figures intact (e.g. fig 2 fig 4). Maybe it's because he used a subfigure like environment (judging by the subcaptions)? https://ar5iv.labs.arxiv.org/html/1609.04747 https://ar5iv.labs.arxiv.org/html/1609.04747
- dginev 3y agoFor the image widths, there is some CSS fine-tuning that is still needed on the arXiv HTML side. I think that will get fixed soon, just needs the right height directive set. Getting subfigures emulated via flexbox is one of our more recent LaTeXML enhancements, and still has some ongoing work (working on it today actually). It can be a bit finicky to test - there are easily 20 different ways people can write LaTeX for subfigures in arXiv.
- blackoil 3y ago> Didn't see a toggle you can run toggleColorScheme() twice in console to switch to light theme or dark theme.
- charleshan 3y agoThis is awesome! Push to Kindle (HTML to EPUB) isn't converting the page properly but I'm sure it's coming soon
- zerop 3y agoThey should also add commenting capabilities under the paper.. a good discussion will lead to more research and information discovery
- krick 3y agoCurious to see how well it will work. Does anybody here know a robust and not crazy computationally expensive solution to extract tables from fairly clean PDF files (especially non-english)?
- llamaInSouth 3y agoNice.... a website that offers even more web pages.
- happyyalda 3y agoUnfortunately, I am from Iran so I can't use this new feature. I got '403 Forbidden' message from the arXiv server. Worse than that, I totally lost my access to arXiv since they changed their CDN to fastly, because fucking mullahs don't like fastly!
- forgingahead 3y agoWhat I would like is for ArXiv to have an LLM to rewrite all papers away from the stodgy, stilted language prevalent in every paper. Just write clearly gang, use proper paragraph breaks and stop with the run-on sentences.
- Erratic6576 3y ago[dead]
- creatonez 3y agoI am glad to see a sans font being used, rather than trying to replicate the serif font from the original papers. It's a bit narrow and fuzzy on low resolutions, but a massive improvement just by switching to sans.
- dang 3y agoWe detached this comment from https://news.ycombinator.com/item?id=38724925 https://news.ycombinator.com/item?id=38724925.
- SallyThinks 3y agoSaw it last night ! I was sooo happy ! Reading papers on phone is a nightmare. Well done guys !
- deleted 3y ago[deleted]
- quickthrower2 3y agoReading papers on mobile now considered sane!
- astrolx 3y agoThis is excellent news. Their HTML formatting is also more pleasant than the HTML articles offered by most journals in my field (e.g arXiv HTML footnotes displayed as sidenotes on large displays!)
- amai 3y agoThis will be on of the most popular applications written in Perl, because this is based on 20 year old https://en.wikipedia.org/wiki/LaTeXML https://en.wikipedia.org/wiki/LaTeXML.
- jcq3 3y agoIt will ease data scraping, automated meta analysis...
- alexmolas 3y agoThis makes downloading and parsing paper data easily, which is pretty handy in the LLM era.
- HeavyStorm 3y agoThank God. Maybe we can now adapt those for mobile?
- killjoywashere 3y agoSo, I'm seeing a lot of chatter in the thread about LaTeX and converting that to HTML and PDF, so LaTeX should be the superior single source of truth. Please keep in mind that many areas of science think of latex as an allergy. I even have a colleague, a plasma physicist, who strongly encourages his team to not use LaTeX because a) collaborators get confused and b) it can be a massive time suck.
- clircle 3y agoI agree with your colleague. At my institution, all of the lowest quality drafts I read are made with latex. I think it's because the programs people use to write latex do not have spelling and grammar checking. Also, the people that prefer latex, are the same types of people that are more interested in technical things, than spelling and grammar.
- deleted 3y ago[deleted]