Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kgeist
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
91.
▲
by
kgeist
5mo ago
Generic robots are potentially faster to configure and set up in many different environments, and they're more accessible, no? If I wanted a specialized robot right now for my small business, I'm not even sure where to begin. But
92.
▲
by
kgeist
5mo ago
Did someone compare DeepSeek 4 Flash to Qwen3.6-27B on real tasks (quality + speed)? According to the benchmarks at artificialanalysis.ai, Qwen3.6-27B is better at agentic tasks, and DS4 is only 2 points better at coding (both with max reas
93.
▲
by
kgeist
5mo ago
I experimented with chunking strategies (how to split, what size chunks should be, how much they should overlap, etc.), query rewriting (one query produces several subqueries to explore different possibilities/search paths in parallel)
94.
▲
by
kgeist
5mo ago
>Experiments at Cactus showed that MLPs can be completely dropped from transformer networks, as long as the model relies on external knowledge source. Heh, what a coincidence, just today one of my students presented research results whic
95.
▲
by
kgeist
5mo ago
When I implemented retrieval in our production system a few months ago, one of the most important benchmarks was cross-language retrieval (query in one language, documents in another), which is a common situation in large enterprises (headq
96.
▲
by
kgeist
5mo ago
Lots of comments here already, just my two cents. I work in R&D and I prefer prototyping things in Python with AI (although we're a 100% Go shop) because: 1) Python is expressive and has packages for everything => faster iterati
97.
▲
by
kgeist
5mo ago
It depends a lot on how you run those models. I think a lot of disagreement is because of that. A lot of people run local models with incredibly small context windows (makes an agentic LLM circle in loops), use very small quants (like 4 bit
98.
▲
by
kgeist
5mo ago
If you look at the earliest versions of Cyrillic, it's basically identical, in shape and form, to the variant of the Greek alphabet used in the Byzantine Empire at the time, they just added letters for the sounds not found in Greek, li
99.
▲
by
kgeist
5mo ago
>Russia has always hated the fact that small Bulgaria gave them their alphabet/culture As someone who lived in Russia for 36 years (also studied linguistics there), it's my first time hearing this. >Most recent rage bait is
100.
▲
by
kgeist
5mo ago
Heh, I made something very similar for the Qwen3 models a while back. It only runs Qwen3, supports only some quants, loads from GGUF, and has inference optimized by Claude (in a loop). The whole thing is compact (just a couple of files) and
101.
▲
by
kgeist
5mo ago
Like 2/3 posts on HN now have this "No X. No Y. No Z." pattern. It's one of strong signals for me that the author didn't bother and just copy pasted their LLM's output as is. And the LLM mostly likely was point
102.
▲
by
kgeist
5mo ago
Electron uses Chromium and nothing prevents them from disabling it, if it ever ends up there.
103.
▲
by
kgeist
5mo ago
>and if you grow the stack, you might have to move it Most stacks are tiny and have bounded growth. Really large stacks usually happen with deep recursion, but it's not a very common pattern in non-functional languages (and function
104.
▲
by
kgeist
5mo ago
Interesting how times have changed. Back in 2015, the entire Go runtime (already a mature codebase) was rewritten from C to Go semi-automatically: one of the maintainers wrote a C-to-Go conversion tool (for a subset of C they used) so that
105.
▲
by
kgeist
5mo ago
Interesting case: - a project manager vibe-coded the change without thinking it through at all - the PR was reviewed by an LLM - an actual engineer gave LGTM without really reviewing the changes, trusting the LLM Did I get this right?
106.
▲
by
kgeist
5mo ago
Maybe running additional inference on all sessions to detect OpenClaw usage would require spending more money than they would save with that detection in the first place (which is the original goal). I also suspect the Claude Code team is j
107.
▲
by
kgeist
5mo ago
The benchmark is strange: single-run results (the author acknowledges it's unreliable) and uses older models like GPT-4o or Opus 4 (although the benchmark is from 2026).
108.
▲
by
kgeist
5mo ago
>The short answer is that variable names are one of the things that confuses LLMs rather than helps them. Unlike with humans, names undermine a model's efforts to keep track of state over larger scales. Models confuse similarly name
109.
▲
by
kgeist
5mo ago
We're planning to do the same thing - buy something like 8xH100 and run all coding there. The CTO almost agreed to find the budget for it but I need to make sure there are no risks before we buy (i.e. it's a viable/usable set
110.
▲
by
kgeist
6mo ago
How do you validate that the reports are correct? What if an executive makes a wrong business decision because the LLM wrote a wrong SQL query?
111.
▲
by
kgeist
6mo ago
>Stash makes your AI remember you. Every session. Forever. How does it fight context pollution?
112.
▲
by
kgeist
6mo ago
Custom constrained decoding could have solved this. Penalize comment tokens :)
113.
▲
by
kgeist
6mo ago
Interesting, my assumption used to be that models over-edit when they're run with optimizations in attention blocks (quantization, Gated DeltaNet, sliding window etc.). I.e. they can't always reconstruct the original code precisel
114.
▲
by
kgeist
6mo ago
From what I understand, ~30b is enough "intelligence" to make coding/reasoning etc. work, in general. Above ~30b, it's less about intelligence, and more about memorization. Larger models fail less and one-shot more often
115.
▲
by
kgeist
6mo ago
>Latency, throughput, and routes don't matter here. When it's 10 seconds for the first token and then a 1KB/sec streamed response, whatever is fine. You can serve Australia from the US and it'll barely matter. This ma
116.
▲
by
kgeist
6mo ago
Discussed 10 months ago here: https://news.ycombinator.com/item?id=44125598 Back then the consensus was that the idea was absurd, I'm surprised they're now trying to make it into a product
117.
▲
by
kgeist
6mo ago
Llama.cpp already uses an idea from it internally for the KV cache [0] So a quantized KV cache now must see less degradation [0] https://github.com/ggml-org/llama.cpp/pull/21038
118.
▲
by
kgeist
6mo ago
>No mention of the fact that Ollama is about 1000x easier to use I remember changing the context size from the default unusable 2k to something bigger the model actually supports required creating a new model file in Ollama if you wanted
119.
▲
by
kgeist
6mo ago
I wonder why it's so bad. Do they just paste a CSV into the raw model? Because in my experience, even small local models can handle it reasonably well if the harness forces them to write & run a Python script that parses the table
120.
▲
by
kgeist
6mo ago
I agree the original poster exaggerated it. But generally models indeed have stopped growing at around 1-1.5 trillion parameters, at least for the last couple of years. >Even now, I don't know if parameter count stopped mattering or
More ›