Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kamranjon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
kamranjon
2mo ago
I know many companies are spending quite a bit of money, I don’t know if it bears out that the increased spend has resulted in increased profits, even if there has been some increase in productivity. I think this is the tough situation many
62.
▲
by
kamranjon
2mo ago
I haven't spun up thinkingcap yet but I'm aware of it and am intending to try it out soon. How did you find it?
63.
▲
by
kamranjon
2mo ago
Would you happen to have a link to that interview? Sounds like an interesting read.
64.
▲
by
kamranjon
2mo ago
Sorry I should have clarified - I meant that a ternary 27b model would outperform a non-quantized or 8 bit quantized 9 or 12b model - which it is generally close to (or much smaller than) in size. So yeah the comparison I was trying to make
65.
▲
by
kamranjon
2mo ago
I’ve actually been really impressed with the 27b model they recently released - amazing performance approaching 40 tok/s on m4 max and I didn’t run into any quality issues in the small set of tasks I tried. Haven’t gone full coding wi
66.
▲
by
kamranjon
2mo ago
Unfortunately Fermion Research appears to entirely AI generate all of their content here, even for the research section: https://www.fermionresearch.com/research/neutrino-8b/ "Neutrino-1 8B was trained native
67.
▲
by
kamranjon
2mo ago
There's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format. PrismML actually targeted the same Qwen 8b model and got it down to 1.75gb here: https:/&
68.
▲
by
kamranjon
2mo ago
There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance characteristics. I'm interested to see if they continue
69.
▲
by
kamranjon
2mo ago
This does seem like a prett niche solution because the electricity is converted into heat and then just used directly as heating - which removes a lot of the utility of electricity. It is a great solution for places where heating is the pri
70.
▲
by
kamranjon
2mo ago
So while SSD streaming is interesting I'm not sure it's exactly the same thing as the per-layer embedding that is being utilized in tandem with streaming here. To utilize per-layer embedding, it would have had to be trained that w
71.
▲
by
kamranjon
2mo ago
Pretty incredible performance for the footprint - really interested to see what could be done on slightly more powerful SBCs like some that have been mentioned in this thread.
72.
▲
by
kamranjon
3mo ago
Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting t
73.
▲
by
kamranjon
3mo ago
You keep saying open weights - but you aren't sharing any information on which models you are using. What benefit does using open weights models provide to the end user if there is zero transparency?
74.
▲
by
kamranjon
3mo ago
Wow I never knew this - what a useful bit of information. Makes me want to try a fountain pen, I do like writing in cursive but I often get hand cramps doing it for extended periods.
75.
▲
by
kamranjon
3mo ago
I’ve recently moved to sourcehut and really enjoy it so far! While I think the creator has strong personal opinions about LLMs - I don’t think they have any intentions of banning LLM generated projects. Anyhow just wanted to share because I
76.
▲
by
kamranjon
3mo ago
Is it a sandbox if the machines hosting your mock package servers have access to the internet? This just seems like a huge failure on the part of OpenAI to secure their environment during testing.
77.
▲
by
kamranjon
3mo ago
No benchmarks, no info on which models are used, ai generated video, just a signup page with nothing else. Anyhow, this kinda reminds me of that quote about architecture: "We replaced our monolith with micro services so that every outa
78.
▲
by
kamranjon
3mo ago
So here is an important question I think. If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model
79.
▲
by
kamranjon
3mo ago
"It went from the start of training to launch in under nine weeks..." This is pretty impressive.
80.
▲
by
kamranjon
3mo ago
What quant are you using?
81.
▲
by
kamranjon
3mo ago
Whoa whoa whoa, 118b params, 8b active MOE, long context reasoning, open weights - music to my ears. Hadn't heard of this lab before but I am very excited, will definitely try this out tomorrow - this is a real sweet spot I think in te
82.
▲
by
kamranjon
3mo ago
What I think is really incredible about this is that they are selling a ton of subscriptions, both personal and business subscriptions. So the fact that they are doing advertising on top of this suggests to me that they are in a deep deep h
83.
▲
by
kamranjon
3mo ago
Pretty bold move to release this at the peak of the Open Models vs Proprietary Models debate. Feels like a really easy decision when it's framed this way.
84.
▲
by
kamranjon
3mo ago
I don't really think the author is misrepresenting the quote in any way, it just seems like many people are reading it as a hard statistic instead of a reference to a properly attributed quote. Just anecdotally though, my company is no
85.
▲
by
kamranjon
3mo ago
"It’s obvious to me that there are ecosystem benefits throughout China, from manufacturing to scientific research; every sector can just plug in these models." I agree with this sentiment and think it's echoed in Fareed Zakar
86.
▲
by
kamranjon
3mo ago
DeepSeek V4 Flash
87.
▲
by
kamranjon
3mo ago
You can run Wordpress locally in your own sandbox to test these vulnerabilities, you don’t have to break any laws.
88.
▲
by
kamranjon
3mo ago
So I am likely coming from a completely different world because I use entirely self hosted models - but is it really common in your workflows to just wait for 12 minutes without any feedback? I am constantly watching my model and following
89.
▲
by
kamranjon
3mo ago
Unfortunately it is becoming more and more common for US alleys to use the pins with cords systems. Many lanes will actually upgrade to this system and I feel it is a mini tragedy in a way. I am very happy to see this because it gives me ho
90.
▲
by
kamranjon
3mo ago
I wouldn’t say I’ve backtracked- I think I’ve been incredibly consistent here. Chinese labs are releasing open weight models, research and analysis. Anthropic is not. They haven’t produced any actual evidence of distillation themselves and
More ›