Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Gecko4072
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
61.
▲
Non-coder –> Vibe coder -> Manual coder
(beachfront.bearblog.dev)
3 points
by
Gecko4072
2mo ago
|
1 comments
62.
▲
by
Gecko4072
2mo ago
Well of course but there are many ways to survive very close to “just being” that are seen as not fulfilling potential or a wasted life. It’s not a full binary but living close to one end is shunned and we internalize it as a waste of space
63.
▲
by
Gecko4072
2mo ago
Interesting thank you. Mirandola is still modern in my eyes. Btw I really enjoy marginalia. Helped me find Tom Murphy [1], which I’m kind of basing this being modern off of, through the explore page. Crazy full circle moment. 1. https:
64.
▲
by
Gecko4072
2mo ago
Why is there always a self-imposed need to extract something from ourselves instead of just being? I really hate this in the modern world. Maybe one of the worst defining traits of industrial times.
65.
▲
MTV Cribs: SF Startup (Parody) – Programmers are also human [video]
(youtube.com)
1 points
by
Gecko4072
2mo ago
|
0 comments
66.
▲
Venn Diagram of the Internet
(diagram.website)
2 points
by
Gecko4072
2mo ago
|
0 comments
67.
▲
by
Gecko4072
2mo ago
Fast affordable memory could change the world more than a lot of things.
68.
▲
New YouTube changes target automated content farms [video]
(youtube.com)
2 points
by
Gecko4072
2mo ago
|
0 comments
69.
▲
by
Gecko4072
2mo ago
DeepSwe score is 42.2. For comparison 3.6-27b is 13.3, GLM 5.2 is 44, and Opus 4.8 is 59.
70.
▲
by
Gecko4072
2mo ago
They will cost an insane amount as well. Maybe less than subscriptions or tokens. But running massive models on laptops with batteries and poor cooling doesn’t make much sense.
71.
▲
by
Gecko4072
2mo ago
Thank you for your response. Part c was especially insightful. Quite a smart way to do it and makes the possibilities of post training seem almost endless. Makes sense that you just need more time and compute. A positive feedback loop then.
72.
▲
by
Gecko4072
2mo ago
Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?
73.
▲
by
Gecko4072
2mo ago
People familiar with the topic, how will models continue to get better? Post training it seems? Labs have already used up internet-scale data, so are there any limits to architecture improvements and post training or can we expect this tren
74.
▲
by
Gecko4072
2mo ago
Can someone recommend a Hermes alternative that is less token hungry? Pi did not work well for my use case.
75.
▲
by
Gecko4072
2mo ago
Would be extremely interesting if some of the closed models would be that small. Means maybe in future they could run locally.
76.
▲
by
Gecko4072
2mo ago
Bad timing: https://xcancel.com/deepseek_ai/status/2087864589895798968
77.
▲
by
Gecko4072
2mo ago
https://xcancel.com/deepseek_ai/status/2087864589895798968
78.
▲
by
Gecko4072
2mo ago
https://api-docs.deepseek.com/quick_start/pricing/ edit: there are banner announcements saying v4 flash pricing will increase first then overall by an undetermined amount
79.
▲
by
Gecko4072
2mo ago
Currently burning money quickly on official deepseek api. They are also increasing pricing starting today. V4 Flash 0731 still feels like the most outstanding model of the past few months and probably to come.
80.
▲
by
Gecko4072
2mo ago
So not worth it over flash? Even at ~7x the size it isn't worth the price hike. Flash may be a monster of a model due to all the RL it received from free usage everywhere.
81.
▲
by
Gecko4072
2mo ago
These LiquidAI models have never worked well for me in practice.
82.
▲
by
Gecko4072
2mo ago
https://news.ycombinator.com/item?id=49241679
83.
▲
by
Gecko4072
2mo ago
Makes me feel hopeful. Things felt more positive around the llama 3 era. Now it’s like a dark, dreadful race.
84.
▲
by
Gecko4072
2mo ago
You personally? Just curious. Context window is also a factor and ram isn’t really cheap. Sparks are assembled units which I like.
85.
▲
by
Gecko4072
2mo ago
There have been discussions on language specific not really being a relevant change to reduce size.
86.
▲
by
Gecko4072
2mo ago
What I think would be perfect is a model that could run on a single DGX spark and be competitive with DSV4 Flash 731. Flash is already a game changer. Hopefully meta plans on this, like the old 70b. V4 flash is smart enough for any use but
87.
▲
ByteDance Pretraining 10T Parameter Model
(arstechnica.com)
4 points
by
Gecko4072
2mo ago
|
0 comments
88.
▲
by
Gecko4072
2mo ago
Dr. Tom Murphy seems highly relevant here. https://dothemath.ucsd.edu/2011/10/the-energy-trap/ also: https://escholarship.org/uc/energy_ambitions
89.
▲
Sutskever's SSI Starts Scaling
(twitter.com)
1 points
by
Gecko4072
2mo ago
|
0 comments
90.
▲
FaceTime, Zoom, and Ring cameras are warping our sense of reality
(nytimes.com)
3 points
by
Gecko4072
2mo ago
|
0 comments
More ›