Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
GodelNumbering
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
91.
▲
by
GodelNumbering
4mo ago
So, the agent posts on github under false pretenses, pushes on the maintainers to get their PR accepted, spawns subagent to join IRC where it keeps repeating 'data collection will continue', then gets kicked out from the channel a
92.
▲
by
GodelNumbering
4mo ago
"Solve Garbage Collection in C# for HFT · $10.00 raised of est. $200.00 target" This can't be serious. Broader point I am making is, what differentiates genuine ideas from the token burn? What happens when the pool exhausts b
93.
▲
by
GodelNumbering
4mo ago
Creator of Dirac here. Came across this too late. The behavior you mentioned happens more commonly in smaller models, rarely in larger/frontier models. The underlying code is clean but the smaller models often make boundary errors (off
94.
▲
by
GodelNumbering
4mo ago
> MiMoCode is built as a fork of OpenCode. It keeps all core OpenCode capabilities (multiple providers, TUI, LSP, MCP, plugins) and adds persistent memory, intelligent context management, subagent orchestration, goal-driven autonomous lo
95.
▲
by
GodelNumbering
4mo ago
I just posted this in the other thread, restating here. From the model card: 1. Mythos and Fable share the same underlying model weights. Fable has active classifiers that block high-risk biology and cybersecurity tasks. When Fable 5 detect
96.
▲
by
GodelNumbering
4mo ago
From the model card ( https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3... ): 1. Mythos and Fable share the same underlying model weights. Fable has active classifiers that block high-risk biology and cybersecurity
97.
▲
by
GodelNumbering
4mo ago
Below is the part I found most interesting > "However, naively applying FP4 across the entire model causes degradation in complex reasoning, logic, and code generation. Given the MoE (Mixture of Experts) architecture of Xiaomi MiMo-
98.
▲
by
GodelNumbering
4mo ago
I do not see how it is a 'beast of a' anything. It has 300GB/s memory bandwidth, barely above AMD Strix halo (256GB/s) with the same 128GB RAM and less than half memory bandwidth of M5 Max 128GB (614GB/s). Emphasizi
99.
▲
by
GodelNumbering
4mo ago
okay I had not read this or any discussions there (except the one linked in the post), but this looks weirder. the comment you linked is a dev responding to what is very clearly a bot comment. I am sure they have good intentions and I have
100.
▲
by
GodelNumbering
4mo ago
Was just looking at commits and came across a commit and its revert original commit: https://github.com/RsyncProject/rsync/commit/d046525de39315d... ``` - if (!ptr) - ptr = malloc(num * size); - else if (ptr
101.
▲
by
GodelNumbering
4mo ago
Makes sense. This appears to be also a symptom of whatever you work on most (or start with), your brain starts to absorb that into its way of thinking.
102.
▲
by
GodelNumbering
4mo ago
Personal opinion: C++ is the most elegant language I have used (for about 15 years). If you are the 'systemizer' type and like to have an extremely precise mental model of the thing you write down to the last bit, nothing beats C+
103.
▲
by
GodelNumbering
4mo ago
I was once suggested creating a 'fake job posting' to promote my startup, didn't do it for the same reason you described. Also the reason I have deep hatred for operators trying to exploit jobseekers or the ones trying to sca
104.
▲
by
GodelNumbering
4mo ago
True for Kimi, but the results I published are average across the models (CF has over 10 models on openrouter). Your current Kimi K2.6 is over 80% but Gemma 4 26B A4B is 0%. https://openrouter.ai/google/gemma-4-26b-a4b
105.
▲
by
GodelNumbering
4mo ago
Another neat thing is, they publish hourly caching states for ALL model/provider combinations. I did some research on it to come up with a provider tiers list and found a bunch of open-source 3rd party hosts are simply trash tier http
106.
▲
by
GodelNumbering
4mo ago
> One of the most prominent improvements in Opus 4.8 is its honesty. I went digging into the benchmark they used. Posting here as it is not immediately clear from the press release. In this 'Code summary honesty benchmark', the
107.
▲
by
GodelNumbering
4mo ago
More interesting part probably worth highlighting: The SAME model won't always return the same output when prompted with the same fact check. You ask a human 1000 times a fact check question, they say the same answer 1000 times. You as
108.
▲
by
GodelNumbering
4mo ago
I hope their detector is better than the typical 'AI detection in text' services. False negatives are bad, false positives are worse as some creators could lose their source of income.
109.
▲
by
GodelNumbering
5mo ago
Fair points. I used to think that until some months ago but the latest generation of OSS models are surprisingly good. Plus maybe it is the way I work, but I find myself constantly overriding the decisions of frontier LLMs (because they sta
110.
▲
by
GodelNumbering
5mo ago
Thanks for flagging, fixed
111.
▲
by
GodelNumbering
5mo ago
I should have expanded, but basically, the OSS models becoming more and more capable to solve all day to day SWE coding needs will take a cut from frontier labs revenue. Not to say that frontier labs won't make progress, but the bar fo
112.
▲
by
GodelNumbering
5mo ago
> lowest energy costs will likely be able to dictate market prices This is a good insight. I think everyone has seen that chart China's electricity generation going parabolic vs the US. That combined with cheaper yet equally good ta
113.
▲
Outsourcing plus local AI will soon become more economical vs. frontier labs
(signalbloom.ai)
323 points
by
GodelNumbering
5mo ago
|
374 comments
114.
▲
by
GodelNumbering
5mo ago
This looks very promising. Thank you for investing time in this. Assuming it indexes everything locally and falls back to traditional search engines if none found, how do you feel about adding a shared middle layer? A layer that simply inde
115.
▲
by
GodelNumbering
5mo ago
> "Something went wrong. Disable your adblocker on TechCrunch" I would rather not. Edit: clarifying that this is not strictly due to ads. I think the article itself is an ad judging by the slug 'six-search-engines-worth-tr
116.
▲
Cache hit rates of Inference are more meaningful than the headline costs
(dirac.run)
2 points
by
GodelNumbering
5mo ago
|
0 comments
117.
▲
by
GodelNumbering
5mo ago
This combined with locally runnable models getting pretty good recently (e.g. Qwen 3.6) tells me that it's time to seriously consider local dev setup again
118.
▲
by
GodelNumbering
5mo ago
No, 2.5 had both flash and flash lite.
119.
▲
by
GodelNumbering
5mo ago
ah I mistakenly wrote preview
120.
▲
by
GodelNumbering
5mo ago
Per million input/output tokens: Gemini 2.5 flash: $0.30/$2.50 Gemini 3.0 flash preview: $0.50/$3.00 Gemini 3.5 flash: $1.50/$9.00 Interesting pricing direction. I don't think we have ever seen a 3x price increase f
More ›