3 ms·
These OCR improvements will almost certainly be brought to google books, which is great. Long term it can enable compressing all non-digital rare books into a m
by fngjdflmdflg 10mo ago
These OCR improvements will almost certainly be brought to google books, which is great. Long term it can enable compressing all non-digital rare books into a manageable size that can be stored for less than $5,000.[0] It would also be great for archive.org to move to this from Tesseract. I wonder what the cost would be, both in raw cost to run, and via a paid API, to do that.
[0] https://annas-archive.org/blog/critical-window.html https://annas-archive.org/blog/critical-window.html
- kridsdale3 10mo agoMore Data for the Data Gods!
- levocardia 10mo agoThis is a really interesting "data flywheel" -- better model >> more usable data >> even better model
- tills13 10mo agosurely there's an upper limit to this though with models literally eating themselves.
- jeffbee 10mo agoWhen a human students learns to read more carefully we don't consider that a negative.
- Choco31415 10mo agoWe can wait for that to start appearing in tests or benchmarks first.
- Workaccount2 10mo agoThey already purposely train them on their own output, it's called synthetic training data.
- visarga 10mo agoNot always, you can improve the loop by putting something real inside, like, a code execution tool, a search engine, a human, other AIs or an API. As long as the model can make use of that external environment its data can improve. By the same logic a human isolated from other humans for a long time might also be in a situation of going crazy. Practical example - using LLMs to create deep research reports. It pulls over 500 sources into a complex analysis, and after all that compiling and contrasting it generates an article with references, like a wiki page. That text is probably superior to most of its sources in quality. It does not trust any one source completely, it does not even pretend to present the truth, it only summarizes the distribution of information it found on the topic. Imagine scaling wikipedia 1000x by deep-reporting every conceivable topic.