Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
valine
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
91.
▲
by
valine
2y ago
It has to be a separate model doing the voice output right? I can’t imagine they’ve solved true multimodal output from a single model, they’d be bragging about it. They’re probably predicting tone of voice tokens. Feed that into an audio tr
92.
▲
by
valine
2y ago
OpenAI has been open about their ability to predict model performance prior to training. When Sam talks about GPT-5 he could very easily be talking about the hypothetical performance of a model given their internal projections. I think it’s
93.
▲
by
valine
2y ago
Llama 3 8B is pretty much the king of its model class right now, so yeah. Meta’s instruct fine tune is also a safe choice, really the only thing you have to play with is the quantization level. Llama 8b 4bit isn’t great, but 8bit might be p
94.
▲
by
valine
3y ago
Model merging is usually done with different fine-tunes of the same model. It doesn’t work if the base models are different. One of the more surprising things is that you can actually repeat layers to improve model performance, ie 1-1-2-2 i
95.
▲
by
valine
3y ago
The architecture changes are very straight forward. Model merging has shown that pre-trained transformer layers are very robust. I’ll bet it’s possible to fine tune a pre-trained model like mistral to use this architecture. That would enabl
96.
▲
by
valine
3y ago
I think you're right. I would go a step further and say that all learning is roughly synonymous with reducing the output space, and that humans do the exact same thing. There are more ways to get the wrong answer to a math problem than
97.
▲
by
valine
3y ago
What kind of problems are you seeing that you think can be improved with a fine tune?
98.
▲
by
valine
3y ago
This is cool. Was looking for model weights, but it seems like maybe it will work with a variety of different models. This is like a RAG/agent app built on top of your typical llama. Am I reading that right?
99.
▲
by
valine
3y ago
I’ve been a netlify user since 2017 and I just deleted all my sites. I can’t risk receiving a $100k bill for toy projects. Your “current policy” is not good enough.
100.
▲
by
valine
3y ago
This paper is arguing that it’s impossible for an LLM to know the answer to every question, therefore it’s impossible to eliminate hallucination. It’s easy to imagine an LLM that responds “I don’t know” to all questions. An LLM like that is
101.
▲
by
valine
3y ago
I can read my own writings without overfitting the neurons in my brain. The key I think is contextualization, something LLMs are great at already. The open question is how to utilize that contextualization ability during training. The argum
102.
▲
by
valine
3y ago
The math behind diffusion models is entirely comprehensible. People will always ask why they need to learn calculus and when they will use it in the real world. Here’s another great answer.
103.
▲
by
valine
3y ago
Huh hadn’t thought of that. Thank you
104.
▲
by
valine
3y ago
As part of the training loop I will periodically ask the model a question. At first the model will respond normally, and then get progressively more repetitive until it starts repeating tokens from the training data. The point in time where
105.
▲
by
valine
3y ago
Yeah I apologize, a lot of the information is scattered across threads right now. I should have spent more time compiling everything in one place. This comment chain in particular might have some of what you’re looking for: https:/&#x
106.
▲
by
valine
3y ago
I’ve gone into pretty great detail on the visualization in the README of my repo. The main utility is detecting individual layers being overfit. There are some specifics about OpenPirate that I’m not at liberty to share at the moment, but t
107.
▲
by
valine
3y ago
I’m confident that there are better methods for fine tuning yet to be discovered. I used this visualization method along with some other unpublished research to train a series of models with less data than is typically required for fine tun
108.
▲
by
valine
3y ago
Trial and error. I tried a bunch of different color maps, this one had the best contrast. My process is to periodically prompt the model as I fine-tune. The features that seem to correlate with the model losing coherence are highlighted nic
109.
▲
by
valine
3y ago
An individual snapshot of the model’s output isn’t very useful. What I’ve found enormously helpful is watching how the structure of the output changes over time as I fine-tune Mistral. There will typically be visual artifacts in the heat-ma
110.
▲
Show HN: NeuralFlow – Visualize the intermediate output of Mistral 7B
(github.com)
134 points
by
valine
3y ago
|
20 comments
111.
▲
by
valine
3y ago
My guess is that the threats are perceived as exaggerated and untruthful, which by extension makes the organization running the ads appear unethical. It’s easier to justify stealing from an unethical company.
112.
▲
by
valine
3y ago
This will likely be used as evidence to justify regulating open weight models. It doesn’t matter if the models are actually dangerous, the messaging is a means to an end.
113.
▲
by
valine
3y ago
Transformers are used outside of NLP
114.
▲
by
valine
3y ago
The industry has been consolidating around transformers. I'd say it's getting deeper, not wider.
115.
▲
by
valine
3y ago
It does exist, I’ve been playing around with it all week :). Let me know if you’d like a custom Mistral 7B I’ll train one for you. From the output of bing chat, and the very specific way it goes off the rails, I suspect Microsoft has figure
116.
▲
by
valine
3y ago
Not their fault, hindsight is 20/20.
117.
▲
by
valine
3y ago
If you have rapid generalization you don’t need scale. Large datasets are only necessary to compensate for the lack of good generalization. The model I posted in my earlier comment responds in character for all queries and was trained in 60
118.
▲
by
valine
3y ago
Takes way more than that in my experience. Back prop isn’t sufficient for rapid generalization.
119.
▲
by
valine
3y ago
I would argue the holdup right now is long term memory. GPTs already have the ability to rapidly generalize and incorporate new knowledge within the context window. The trick is to retain what it has learned. It won’t take a very long time
120.
▲
by
valine
3y ago
Why are you letting your AGI run wild? Let’s make the assumption that AGI is just a GPT combined with a new training algorithm that allows it to learn rapidly and generalize from small amounts of information. The moment you turn off the tra
More ›