Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rybosome
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
rybosome
1y ago
That’s cool, I’d love to see the advanced ToolController when it’s available! Great points about not updating priors. I also thought about it a bit more and realized that there’s a way you can largely mitigate the out-of-distribution infere
62.
▲
by
rybosome
1y ago
Thanks for the informative and inspiring post! This is definitely cool, and I can imagine very useful. However I do want to mention that the “recommended” flow these days isn’t to separate out a tool request in the way you have. Eg instead
63.
▲
by
rybosome
1y ago
One could find a 5-minute slice of any highly successful project I’ve worked on where my actions look foolish and my tools look broken. Isolating your analysis of this to a single unflattering interaction is intellectually dishonest; you ha
64.
▲
by
rybosome
1y ago
This appears to me to be intentional and ironic to make a point rather than in earnest. I am interpreting this as a statement about snap judgements in an age where AI will increasingly play the role of a judge or assessor of humans. Perhaps
65.
▲
by
rybosome
1y ago
A very interesting idea. I am curious about this sharing and blending of the various nets; I wonder if something as naive as averaging the weights (assuming the neural nets all have the same dimensions) would actually accomplish that?
66.
▲
by
rybosome
1y ago
I cannot for the life of me think of what you are referring to. If it’s COVID-related mandates like vaccines and lockdowns, then surely it’s obvious that NASA had nothing to do with that? There is no single issue that I can see linking all
67.
▲
by
rybosome
1y ago
Agreed, especially when in this context of training a smaller model on a larger model’s outputs. Distillation is generally accepted as an effective technique. This is exactly what I did in a previous role, fine-tuning Llama and Mistral mode
68.
▲
by
rybosome
1y ago
Echoing this advice in the USA. Don’t fuck with this, the language in the laws is loose enough to prosecute almost anything security-related, even well-meaning efforts like this.
69.
▲
by
rybosome
1y ago
This is quite an impressive list, and many of the things ranking low in difficulty for the author would have been quite high for me. It's definitely inspiring, makes me want to dust off a toy of my own. That said, I feel the conclusion
70.
▲
by
rybosome
1y ago
You clearly have an emotional connection to people that you feel are being harmed by AI, so I’m not going to gaslight you about that. But I will tell you that real humans asking private, real questions of LLMs is also happening, and these t
71.
▲
by
rybosome
1y ago
I was wondering if it is function overloads of the same name, defining it non-variadically: add_function(PythonObject) add_function(PythonObject, PythonObject) add_function(PythonObject, PythonObject, PythonObject) There was a similar patte
72.
▲
by
rybosome
1y ago
The first response definitely wasn’t. It laid out the hazards in great detail, then asserted the likelihood of making it implausibly far. I poked at that conclusion and it backed off until we arrived at shear winds in the outer edges, where
73.
▲
by
rybosome
1y ago
I’ve wondered about this a lot. More so the grim question: if you were in a typical space suit sitting in a ship just outside Jupiter, then propelled yourself towards the planet - what would kill you first? Assume you are close enough that
74.
▲
by
rybosome
1y ago
Global dependencies were disallowed back in 2018 with a tiny handful of exceptions that were difficult or impossible to make fully regional. Chemist, the service that went down, was one of those. Generally GCP wants regionality, but because
75.
▲
by
rybosome
1y ago
Must be hell inside GCP right now. That was a big outage, and they were tired of big outages years ago. It was already extremely difficult to move quickly and get things done due to the reliability red tape, and I have to imagine this will
76.
▲
by
rybosome
1y ago
Exactly the question on my mind. I have an ongoing fantasy of a super powerful, private inference server sitting in my closet that I can throw at stuff like this.
77.
▲
by
rybosome
1y ago
The honest answer is that we are still figuring out where to draw the line between an agent and a human, because that line is constantly shifting. Given your example of the resume screening application from the post and today's capabil
78.
▲
by
rybosome
1y ago
AI code generation tools work like this. Let me reword your phrasing slightly to make an illustrative point: > so for an employee, instead of specifying explicit if conditions, you specify outcomes and you leave the human to figure out w
79.
▲
by
rybosome
1y ago
Adding to that excellent high level explanation of what the attention mechanism is, I’d add (from my reading of the abstract of this paper); This work builds a model that has the ability to “remember” parts of its previous input when genera
80.
▲
by
rybosome
1y ago
I have a very clear memory of offending someone with the use of the word “crap” years ago. As a kid I worked in a restaurant that sold Cincinnati-style chili - noodles with sweet chili and cheese on top. We were encouraged to offer customer
81.
▲
by
rybosome
1y ago
Scale has also built massive amounts of proprietary datasets that they license to the big players in training. Meta, Google, OpenAI, Anthropic, etc. all use Scale data in training. So, the play I’m guessing is to shut that tap off for every
82.
▲
by
rybosome
1y ago
A wonderful approach generally and something we also do to some extent, but not a substitute for fine-tuning in our case. We are working in a domain where there is very limited training data, so what we really want is continued pre-training
83.
▲
by
rybosome
1y ago
Yep! So my use case currently is admittedly very specific. My company uses LLMs to automate hardware design, which is a skill that most LLMs are very poor at due to the dearth of training data. For tasks which involve generation of code or
84.
▲
by
rybosome
1y ago
Haven’t tried it personally, as this was a use case where a classic SFT was effective for what we wanted and none of us had done LoRa before. Really interested in the idea though! The dream is that you have your big, general base model, the
85.
▲
by
rybosome
1y ago
I think the point the author misses is that many applications of fine-tuning are to get a model to do a single task. This is what I have done in my current role at my company. We’ve fine-tuned open weight models for knowledge-injection, amo
86.
▲
by
rybosome
1y ago
I wish that I had the deep biology knowledge to make a joke (not joke?) Large Mushroom Model based on actual signal data. This also makes me wonder about the idea of attempting to expand on this research by artificially stimulating a fungal
87.
▲
by
rybosome
1y ago
This is a great point - communication is a vast array of information channels. “Motion” is communication whether or not it evolved to be, and so is every other aspect of occupying space in a biome. And given that the organism under scrutiny
88.
▲
by
rybosome
1y ago
Yeah, very much agreed that in the spirit of good discussion I should’ve at least asked about their experiences and use case before jumping to finger wagging. But that said, let me reiterate a couple important points from my post: > With
89.
▲
by
rybosome
1y ago
Many, many people are in fact “using the code it generated on its own”. I’ve been putting LLM-assisted PRs into production for months. With no disrespect meant, if you’re unable to find utility in these tools, then you aren’t using them cor
90.
▲
by
rybosome
1y ago
I’ve spent some time optimizing Python performance in a web app and CLI, and yeah it absolutely sucks. Module import cost is enormous, and while you can do lots of cute tricks to defer it from startup time in a long-running app because Pyth
More ›