Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
_peregrine_
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Training a Specialist Code Search Agent with Turbopuffer
(appliedcompute.com)
1 points
by
_peregrine_
27d ago
|
0 comments
2.
▲
The Shape and Feel of the Post-AI Data Stack
(iandmacomber.com)
1 points
by
_peregrine_
1mo ago
|
0 comments
3.
▲
by
_peregrine_
1mo ago
instant bookmark!
4.
▲
A graph-theoretic approach to building reliable LLM judges for retrieval
(georgianailab.substack.com)
3 points
by
_peregrine_
4mo ago
|
0 comments
5.
▲
Mixing numeric attributes into text search for better first-stage relevance
(turbopuffer.com)
3 points
by
_peregrine_
5mo ago
|
0 comments
6.
▲
How to build a distributed queue in a single JSON file on object storage
(turbopuffer.com)
2 points
by
_peregrine_
8mo ago
|
0 comments
7.
▲
by
_peregrine_
8mo ago
this is actually not how cursor uses turbopuffer, as they index per codebase and thus need many mid-sizes indexes as opposed to one massive index as this post describes
8.
▲
by
_peregrine_
8mo ago
unfortunately i'm not able to share the customer or use case :( but the metrics that you see in the first charts in the post are from a production cluster
9.
▲
by
_peregrine_
8mo ago
the solution described in the blog post is currently in production at 100B vectors
10.
▲
by
_peregrine_
8mo ago
seems like a good rule of thumb to me! though i would perhaps lump "cost" into the "until it breaks" equation. even with decent perf, pg_vector's economics can be much worse, especially in multi-tenant scenarios whe
11.
▲
ANN v3: 200ms p99 query latency over 100B vectors
(turbopuffer.com)
109 points
by
_peregrine_
8mo ago
|
47 comments
12.
▲
Designing inverted indexes in a KV-store on object storage
(turbopuffer.com)
3 points
by
_peregrine_
9mo ago
|
0 comments
13.
▲
Why BM25 queries with more terms can be faster (and other scaling surprises)
(turbopuffer.com)
53 points
by
_peregrine_
9mo ago
|
1 comments
14.
▲
Vectorized MAXSCORE over WAND, especially for long LLM-generated queries
(turbopuffer.com)
2 points
by
_peregrine_
10mo ago
|
0 comments
15.
▲
Flink's 95% Problem
(tinybird.co)
2 points
by
_peregrine_
11mo ago
|
0 comments
16.
▲
by
_peregrine_
1y ago
yeah I mean that's basically what Javi talks about in the post... if you can throw hardware at it you can scale it (ingestion scales linearly with shards) but the post has some interesting thoughts on how you do the high-scale ingestio
17.
▲
by
_peregrine_
1y ago
definitely interesting and related
18.
▲
by
_peregrine_
1y ago
nice one
19.
▲
How to ingest 1B rows/s in ClickHouse
(tinybird.co)
32 points
by
_peregrine_
1y ago
|
10 comments
20.
▲
How we built our own Claude Code (for data)
(tinybird.co)
3 points
by
_peregrine_
1y ago
|
0 comments
21.
▲
Show HN: Click the Circle
(tb-peregrine.github.io)
2 points
by
_peregrine_
1y ago
|
0 comments
22.
▲
by
_peregrine_
1y ago
Pretty solid at SQL generation, too. Just tested in our generation benchmark: https://llm-benchmark.tinybird.live/ Not quite as good as Claude but by the best Qwen model so far and 2x as fast as qwen3-235b-a22b-07-25 Specif
23.
▲
Tinybird made a ClickHouse CLI agent
(tinybird.co)
1 points
by
_peregrine_
1y ago
|
1 comments
24.
▲
Why LLMs Struggle with Analytics
(tinybird.co)
1 points
by
_peregrine_
1y ago
|
0 comments
25.
▲
by
_peregrine_
1y ago
Yeah we're always looking for new models to add
26.
▲
by
_peregrine_
1y ago
Yeah I mean SQL is pretty nuanced - one of the things we want to improve in the benchmark is how we measure "success", in the sense that multiple correct SQL results can look structurally dissimilar while semantically answering th
27.
▲
by
_peregrine_
1y ago
We need to add it
28.
▲
by
_peregrine_
1y ago
Actually no, we have it up to 3 attempts. In fact, Opus 4 failed on 36/50 tests on the first attempt, but it was REALLY good at nailing the second attempt after receiving error feedback.
29.
▲
by
_peregrine_
1y ago
We should definitely add o3 - probably will soon. Also looking at testing the Qwen models
30.
▲
by
_peregrine_
1y ago
Noted, also feel free to add an issue to the GitHub repo: https://github.com/tinybirdco/llm-benchmark
More ›