4 ms·
What about clean, semantic HTML? It was already optimized for bots and search engines (which are bots) and it has been used for decades. Why we need to serve i
by collimarco 1mo ago
What about clean, semantic HTML?
It was already optimized for bots and search engines (which are bots) and it has been used for decades. Why we need to serve in markdown now?
There are also many parts of the HTML, like navs, that are useful for bots and AI and may be removed in the markdown version.
- slowin 1mo agoPresumably markdown uses far fewer tokens.
- simonw 1mo agoThat used to matter to me back in the days when the best models still only accepted ~32,000 tokens, but these days even the models that run on my laptop are happy with ~100,000 and the hosted models I use take ~200,000 or more.
- slowin 1mo agoIf it's one of many tool calls, I'd assume that less is more.
- simonw 1mo agoThe trick there is to use a subagent to read the HTML page and extract the relevant information, than dumping all that HTML into your top-level session. That's effectively using an LLM as an HTML to markdown converter, which is both absurdly wasteful and also surprisingly inexpensive (if you use a model like GPT-5.6 Luna.)
- honr 1mo agoI found HTML itself works [slightly] better with some LLMs (that I happen to frequent). So, when scraping, I parse the resulting html, simplify it, and produce simpler html contents (closer to semantics of what I guessed the real content was). Given processing tools for html are far more mature, I tend to keep content in semantic, low structure html.
- selcuka 1mo agoThey accept more tokens these days, but they are still more accurate with a shorter context [1]. [1] https://arxiv.org/abs/2307.03172 https://arxiv.org/abs/2307.03172
- honr 1mo agoIs that even true? I most often use HTML. HTML is about 5%-20% more tokens than a similar Markdown. As a rule of thumb, the number of tags/structural tokens doubles, when going from markdown to html, while the rest don't change much. On the other hand, I can view HTML without any extra/unusual tools. And composing HTML when I need a bit of structure is far easier than composing markdown.
- slowin 1mo ago> On the other hand, I can view HTML without any extra/unusual tools. And composing HTML when I need a bit of structure is far easier than composing markdown. This is kind of the opposite of reality no? Markdown is just plain text and meant to be human readable. You don't need XML tags to read and write it, opposed to html where you do and you need a browser to properly view it.
- honr 1mo agoNo, it's just that I wasn't very clear. HTML I can view in any browser / webview / etc. Good markdown viewers are fewer / more special, or end up translating md to html for display. And by composing, I didn't mean writing by hand. We are talking about prompting, right? Or that is what I thought we are talking about. Composing HTML "components" into a final prompt HTML is easier than composing markdown snippets into the final prompt. That is because with HTML there are several ergonomic libraries to parse HTML to AST and to format AST back to HTML. The libraries (for parsing to AST and back to strings) are more limited with markdown.
- sebastiennight 1mo agoA "prompt" is what goes into the LLM, so I'm not sure what you mean by "a final prompt HTML".
- honr 1mo agoI often compose prompts from various sources (my "AGENTS.md", "CURRENT_TASK", "CURRENT_PHASE", ...). Last year I was composing in markdown and that got very tedious very quickly. I tried various text and text-like formats, and it turned out HTML is already one of the best formats for this, including for an "AGENTS.md" (which I keep in HTML despite the required ".md" extension). Of course, if you build up your prompt entirely by hand (copy pasting snippets, etc.) you can get away with markdown or plain text. That works okay for some harnesses, but it leaves too much control (or more accurately, opportunity to misunderstand) to the harness.
- o_m 1mo agoMarkdown isn't as expressive. Not all HTML content can be converted to Markdown without losing some of the semantics
- Zardoz84 1mo agoNot my problem. Currently my problem is the high traffic of bots that kills our web apps and don't respect robots.txt or meta tags . We are currently deploying Anubis.
- dubcanada 1mo agoThen convert from HTML to markdown before you convert to tokens? It's not rocket science, if markdown is "better" stripping all possible HTML tags and leaving just text, images also works. That is more then likely what any automated "serve markdown on the fly" would end up doing.
- meindnoch 1mo ago>What about clean, semantic HTML? Which React package is this?
- deleted 1mo ago[deleted]
- k1m 1mo agoI agree. I think a lot of people here are assuming that the full HTML retrieved has to go into the LLM eating up tokens. But why wouldn't the agent try to clean up first and remove bloat and convert to markdown itself, before feeding into LLM. Semantic HTML would make that easier.
- troupo 1mo ago> But why wouldn't the agent try to clean up first and remove bloat and convert to markdown itself, before feeding into LLM. There's no "agent". It's a few wrappers around API calls in a trenchcoat.