Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gnulinux
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
gnulinux
1y ago
Side effects
32.
▲
by
gnulinux
1y ago
There is no moving goal post. The license isn't an open source license, which by definition means the code is not open source. When you have access to source code of a program, but don't necessarily have the legal rights to distri
33.
▲
by
gnulinux
1y ago
Sure, quarter of a 1B, the point was a generalization about <<1B models.
34.
▲
by
gnulinux
1y ago
Well, this is a 270M model which is like 1/3 of 1B parameters. In the grand scheme of things, it's basically a few matrix multiplications, barely anything more than that. I don't think it's meant to have a lot of knowled
35.
▲
by
gnulinux
1y ago
Rerankers are used downstream from an embedding model. Embedding models are "coarse" so they give false positives for things that may not be as relevant as contender text. Re-ranker, ranks bunch of text based on a query in order t
36.
▲
by
gnulinux
1y ago
This is true, but I still avoid using examples. Any example biases the output to an unacceptable degree even in best LLMS like Gemini Pro 2.5 or Claude Opus. If I write "try to do X, for example you can do A, B, or C" LLM will do
37.
▲
by
gnulinux
1y ago
Maybe, ever since I graduated from college I learned again and again that pretty much anything worth thinking about in life boils down to math for me. I'd maybe/probably study CS, as a minor or double major, but Pure/Applied
38.
▲
by
gnulinux
1y ago
My first impressions: not impressed at all. I tried using this for my daily tasks today and for writing it was very poor. For this task o3 was much better. I'm not planning on using this model in the upcoming days, I'll keep using
39.
▲
by
gnulinux
1y ago
Imho chatterbox is the current open weight SOTA model in terms of quality: https://huggingface.co/ResembleAI/chatterbox
40.
▲
by
gnulinux
1y ago
No, unfortunately, I haven't used Qwen3-coder yet. I do like Claude 4 Sonnet, but my favorite programming LLM is Gemini 2.5 Pro at the moment, I think it's the smartest model (Claude and o3 do print better code though). I have exp
41.
▲
by
gnulinux
1y ago
Sure you're right, but if I can squeeze out o4-mini level utility out of it, but its less than quarter the price, does it really matter?
42.
▲
by
gnulinux
1y ago
Name recognition? Advertisement? Federal grant to beat Chinese competition? There could be many legitimate reasons, but yeah I'm very surprised by this too. Some companies take it a bit too seriously and go above and beyond too. At thi
43.
▲
by
gnulinux
1y ago
Not even that, even if o3 being marginally better is important for your task (let's say) why would anyone use o4-mini? It seems almost 10x the price and same performance (maybe even less): https://openrouter.ai/openai&#
44.
▲
by
gnulinux
1y ago
Wow, that's significantly cheaper than o4-mini which seems to be on part with gpt-oss-120b. ($1.10/M input tokens, $4.40/M output tokens) Almost 10x the price. LLMs are getting cheaper much faster than I anticipated. I'm
45.
▲
by
gnulinux
1y ago
It's averaging to $0.3/1M input tok and $1.2/1M output tok. That's kind of mind blowingly cheap for a model at its caliber. Gemini 2.5 Pro is more than 10x that price.
46.
▲
by
gnulinux
1y ago
At $2/1Mt it's cheaper than e.g. Gemini 2.5 Pro which is ($1.25/1Mt for input and $10/1Mt per output). When I code with Aider my requests average to something like 5000 tokens input and 800 tokens output. At this rate, G
47.
▲
by
gnulinux
1y ago
Qwen3 is the open weight state of the art at the moment. Qwen3-embedding-8B and Qwen3-reranker-8b are surprisingly good (according to some benchmarks, better than Gemini 2.5 embedding). 4B is also nearly as good so you might as well use tha
48.
▲
by
gnulinux
1y ago
Tool calling complements RAG. You build a full scale RAG (embedding, reranker, create prompt, get output from LLM) and hook that to a tool another agent can see. That combines both their power.
49.
▲
by
gnulinux
1y ago
You can fix the code of course. I just experiment with what sort of comments produce better code. In my experience, heavily commented code is handled by LLMs significantly better. So the total quality of comments eventually add up, if you&#
50.
▲
by
gnulinux
1y ago
LLMs do not reason at all (i.e. deductive reasoning using a formal system). Chain of thought etc simulate reasoning by smoothing out the path to target tokens by adding shorter stops on the way.
51.
▲
by
gnulinux
1y ago
Do you still have this problem if you add a comment before declaring the variable like "Note: thingsById is not a dictionary, it is an array. Each index of the array represents a blabla id that maps to a thing" In my experience th
52.
▲
by
gnulinux
1y ago
Fine-tuning existing base models on your programming language is pretty practical. [1] You might need a very good and large dataset but that's hardly a problem for a programming language you're generating because you better have t
53.
▲
by
gnulinux
1y ago
Do 2bit quantizations really work? All the ones I've seen/tried were completely broken even when 4bit+ quantizations worked perfectly. Even if it works for these extremely large models, is it really much better than using somethin
54.
▲
by
gnulinux
1y ago
Yes I do exactly this too. But I do sometimes 2, 3 shot some problems. The method I use is that in Cursor/Copilot interface I use a Markdown file to chat with the bot. Once I have some solutions after a few turns, I edit the file, add
55.
▲
by
gnulinux
1y ago
Yes it would matter. If you just have budget to run a 8B model and it's sufficient for the easy problem you have, a better 8B model with the same spec requirements is necessarily better regardless of how it compares to some other model
56.
▲
by
gnulinux
1y ago
I'm curious if this would also improve small local models. E.g. if I "alloy" Qwen3-8B and OpenThinker-7B is it going to be "better" than each models? I'll try testing this in my M1 Pro.
57.
▲
by
gnulinux
1y ago
I guess I'm skeptical of using a non-associative algebra instead of something that can trivially be made into a ring or field (i.e. matrix algebra). What advantages does this give us?
58.
▲
by
gnulinux
1y ago
That's certainly your take on poetry, but not mine. It also may not be everyone's. I think everyone has a unique reading of each poetry, and thus reading and listening are different. There is nothing wrong with listening to poetry
59.
▲
by
gnulinux
1y ago
I love the "old internet" vibe with the new internet look. I think it's a very creative idea, I wouldn't mind the haters too too much. I love it!
60.
▲
by
gnulinux
1y ago
It's definitely not the hardest "arthouse" novel (or whatever you call it), I found Gravity's Rainbow by Pynchon so much more harder, and Beckett's Three Novels (i.e. Molloy , Malone Dies , The Unnamable ) wa
More ›