4 ms·
For as much as I would love for this to work, I'm not getting great results trying out the 1.5b model in their example notebook on Colab. It is impressively fa
by sippeangelo 2y ago
For as much as I would love for this to work, I'm not getting great results trying out the 1.5b model in their example notebook on Colab.
It is impressively fast, but testing it on an arxiv.org page (specifically https://arxiv.org/abs/2306.03872 https://arxiv.org/abs/2306.03872) only gives me a short markdown file containing the abstract, the "View PDF" link and the submission history. It completely leaves out the title (!), authors and other links, which are definitely present in the HTML in multiple places!
I'd argue that Arxiv.org is a reasonable example in the age of webapps, so what gives?
- faangguyindia 2y agoQuestion is why even use these small models? When you've Google Flash which is lightening fast and cheap. My brother implemented it in option-k : https://github.com/zerocorebeta/Option-K https://github.com/zerocorebeta/Option-K It's near instant. So why waste time on small models? It's going to cost more than Google flash.
- FL33TW00D 2y agoPrivacy, Cost, Latency, Connectivity.
- oezi 2y agoWhat is Google Flash? Do you mean Gemini Flash? If so, then the article talks about that general purpose LLMs are worse than this specialized LLM for Markdown conversion.
- sippeangelo 2y agoIn this case it is not, though. As much as I'd like a self-hostable, cheap and lean model for this specific task, instead we have a completely inflexible model that I can't just prompt tweak to behave better in even not-so-special cases like above. I'm sure there are good examples of specialised LLMs that do work well (like ones that are trained on specific sciences), but here the model doesn't have enough language comprehension to understand plain English instructions. How do I tweak it without fine-tuning? With a traditional approach to scraping this is trivial, but here it's unfeasible to the end user.
- dartos 2y agoSometimes you don’t want to share all your data with the largest corporations on the planet.
- deleted 2y ago[deleted]
- randomdata 2y agoSmall models often do a much better job when you have a well-defined task.