Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
raunakchowdhuri
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
raunakchowdhuri
6mo ago
The big one is that LLMs get lazy on repetitive tasks. They'll skip rows or consolidate entries instead of grinding through every last one. So you need verify-and-re-extract loops rather than single-pass processing. Breaking work into
2.
▲
by
raunakchowdhuri
6mo ago
We've made a lot of changes in the past few months that make our standard extract much, much better, as well as Deep Extract for documents even longer than that. We'd love for you to give it a try!
3.
▲
Reducto releases Deep Extract
(reducto.ai)
50 points
by
raunakchowdhuri
6mo ago
|
9 comments
4.
▲
Show HN: A SOTA chart-extraction system combining traditional CV and LVMs
(reducto.ai)
1 points
by
raunakchowdhuri
10mo ago
|
0 comments
5.
▲
by
raunakchowdhuri
11mo ago
Have a slack channel with them, these are the versions they mentioned: posthog-node 4.18.1 posthog-js 1.297.3 posthog-react-native 4.11.1 posthog-docusaurus 2.0.6
6.
▲
We did a DB migration without logical replication – with zero downtime
(reducto.ai)
4 points
by
raunakchowdhuri
1y ago
|
0 comments
7.
▲
by
raunakchowdhuri
1y ago
We're fixing it! This for some reason happens on only _some_ phones in our office so was hard to repro. I think has to do with Safari rendering. Will tone down our WebGPU usage
8.
▲
by
raunakchowdhuri
1y ago
dang no way! we were both in boston too
9.
▲
by
raunakchowdhuri
1y ago
this is exactly where we're going with this! glad you see the vision :)
10.
▲
by
raunakchowdhuri
1y ago
yep!
11.
▲
by
raunakchowdhuri
2y ago
comparisons to more outputs coming soon!
12.
▲
by
raunakchowdhuri
2y ago
We ran some benchmarks comparing against Gemini Flash 2.0. You can find the full writeup here: https://reducto.ai/blog/lvm-ocr-accuracy-mistral-gemini A high level summary is that while this is an impressive model, it
13.
▲
Evaluating Mistral OCR Against Gemini 2.0 Flash
(reducto.ai)
15 points
by
raunakchowdhuri
2y ago
|
0 comments
14.
▲
by
raunakchowdhuri
2y ago
CTO of Reducto here. Love this writeup! We’ve generally found that Gemini 2.0 is a great model and have tested this (and nearly every VLM) very extensively. A big part of our research focus is incorporating the best of what new VLMs offer w
15.
▲
by
raunakchowdhuri
2y ago
would encourage you to take a look at some of the real data here! https://huggingface.co/spaces/reducto/rd_table_bench you'll find that most of the errors here are structural issues with the table or inabilit
16.
▲
by
raunakchowdhuri
2y ago
Love the Pubtables work! It's a really useful dataset. Their data comes from existing annotations from scientific papers, so in our experience it doesn't include a lot of the hardest cases that a lot of methods fail at today. The
17.
▲
Rd-TableBench – Accurately evaluating table extraction
(reducto.ai)
29 points
by
raunakchowdhuri
2y ago
|
6 comments
18.
▲
by
raunakchowdhuri
2y ago
hmmm idk how I would feel about giving an llm cluster access from a security pov
19.
▲
by
raunakchowdhuri
2y ago
Interesting... how did you do the scraping of the documentation?
20.
▲
Show HN: Reducto – A vision based document ingestion API for LLMs
(reducto.ai)
3 points
by
raunakchowdhuri
3y ago
|
0 comments
21.
▲
Open Sourcing Remembrall: A Long-Term Memory Proxy for LLMs
(github.com)
3 points
by
raunakchowdhuri
3y ago
|
1 comments
22.
▲
by
raunakchowdhuri
3y ago
Hey HN, A few weeks ago I shared a beta of Remembrall here and got a lot of great feedback from people in this community, with one of the most common requests being to open source the project. Excited to share that we’re doing exactly that!
23.
▲
by
raunakchowdhuri
3y ago
Yep - full export will always be supported for your data. No customer # restriction. I have some optimizations in the works such that the secondary gpt 3.5 call only gets triggered when one iterative conversation thread ends (determined via
24.
▲
by
raunakchowdhuri
3y ago
Not right now, but it's a good idea. I'll add that later this week.
25.
▲
by
raunakchowdhuri
3y ago
So cool! How do tools like pgvector do it? https://github.com/pgvector/pgvector They have a WHERE clause built in - no? And then you can additional sort by semantic similarity? Or is this a bit different than that...
26.
▲
by
raunakchowdhuri
3y ago
For sure! When you send an OpenAI request, after a delay (to ensure the user doesn't keep chatting in the same session), a secondary GPT 3.5 call is made to "autosave" the result. This GPT call gets the information from the c
27.
▲
Show HN: Remembrall – Long-term memory proxy for LLMs
(remembrall.dev)
4 points
by
raunakchowdhuri
3y ago
|
6 comments
28.
▲
by
raunakchowdhuri
3y ago
Neil posted the paper in our fraternity's ML group chat (MIT things lol), and I expressed some skepticism at the results. Initially we started looking into it more for curiosity's sake, but as we started digging we kept finding mo
29.
▲
by
raunakchowdhuri
3y ago
Can attest that the distribution is odd from the test set that we sampled. We've already run the compute to run the zero-shot GPT model on all of the datapoints in the provided test set. We're going through the process now of grad
30.
▲
by
raunakchowdhuri
3y ago
You're right that our post doesn't quite show that GPT4 cannot perform well on MIT curriculum. We try to be up front about this in the conclusion: > Our critiques are largely of the methodology and rigor of this study, not abou
More ›