Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
valine
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
151.
▲
by
valine
3y ago
The coding process. I've been experimenting with fine tuning methods where I freeze various layers or use different loss functions for attention vs feed forward. It's random little things like the names of layers that trip me up.
152.
▲
by
valine
3y ago
This is a really neat idea. Would love to have a similar view for llama like models. I’ve been working with Mistral 7B lately and it’s annoying how many small changes there are between it and llama. Having a view like this would be a good t
153.
▲
by
valine
3y ago
It’s not perfect, if you know of a better alternative I would genuinely love to hear about it.
154.
▲
by
valine
3y ago
You can supply the model with a list of facts already, that’s not the problem. Within the context window the model is able to learn and generalize new information. Fine tuning is very unintelligent in the sense that it doesn’t take the cont
155.
▲
by
valine
3y ago
Training a local LLM on individual facts is a tricky one. Typically it’s not possible to train with a limited quantity of data and expect the model to generalize on that data well. In context learning generalizes well, but it’s a bad fit fo
156.
▲
by
valine
3y ago
I guess the question is how the mix of experts works. Do they predict one token from each model? If so you’re still doing 1.7T worth of computation.
157.
▲
by
valine
3y ago
Not really fair to compare a 4B model and a 1.7T model. Per flop the 4B model here is slightly more expensive than GPT4.
158.
▲
by
valine
3y ago
I suspect OpenAI’s advantage is their ability to synthesize a good fine tuning dataset. My question would be is this leaking data from the fine tuning dataset or from the initial training of the base model? The base model training data is l
159.
▲
by
valine
3y ago
There’s no shortage of 3D mesh data to train on. Who to say scaling up the parameter count won’t allow for increasingly intricate topology the same way scaling language models improved reading comprehension.
160.
▲
by
valine
3y ago
Even if this is “only” mesh autocomplete, it is still massively useful for 3D artists. There’s a disconnect right now between how characters are sculpted and how characters are animated. You’d typically need a time consuming step to retopol
161.
▲
by
valine
3y ago
Sure, the more information you give the model the better. While our current models generate imperfect code, GPT4 code is much more usable out of the box than GPT3.5 code. If that trend continues we will absolutely get to the point where mis
162.
▲
by
valine
3y ago
This 2016 comic failed to predict GPTs. It turns out the same model that writes your code can also write the specification. I don’t blame the author, it’s nearly impossible to predict future technology. Pretending that GPTs don’t exist on t
163.
▲
by
valine
3y ago
GPT-4 can write html from a messy napkin sketch. These models won’t need formal requirements. “I’m having a problem with A, B and C” will likely be all the model needs from you. Your argument was very common and quite persuasive two years a
164.
▲
by
valine
3y ago
Or, more likely, the march towards better abstraction ends with a natural language interface where problem descriptions can be effortlessly translated into functional code. We’re already in spitting distance of this with current language mo
165.
▲
by
valine
3y ago
I walked back nothing. OpenAI was surprised by the mass adoption of ChatGPT, they saw it as an early technical preview. I don’t understand why some people have a such hard time envisioning the potential of new technologies without a polishe
166.
▲
by
valine
3y ago
>> Inversion of burden of proof Nope. OpenAI has already demonstrated the ability to generalize GPT4 to a new modality. Your claim that text models can only generalize to images and not other modalities is utterly unconvincing. Explai
167.
▲
by
valine
3y ago
>> Text and 2D images are a tiny subset of physical reality as perceived by an able-bodied human. Even our best approximation is a poor representation. This is wrong. There’s nothing magical about human perception. You see the world b
168.
▲
by
valine
3y ago
We already have multimodal models that take both images and text as input. The bulk of the training for these models was in text, not images. This shouldn’t be surprising. Text is a great way of abstractly and efficiently representing reali
169.
▲
by
valine
3y ago
Covering an airplane in feathers isn't going to make it fly faster. Biological plausibility is a red haring imho.
170.
▲
by
valine
3y ago
One problem is vanishing gradients. You can only back propagate through the loop so many times before you lose the signal.
171.
▲
by
valine
3y ago
In the grand scheme of human history PCs didn’t come out all that long ago.
172.
▲
by
valine
3y ago
The response gets more reasonable the smaller the model in question. A 1B parameter model passing grade-school math tests would be much more alarming (exciting?) than a GPT-4 sized model doing the same. GPT-4 probably has some version of th
173.
▲
by
valine
3y ago
GPU utilization should be down when using this technique. I’m hoping this could allow for more efficient batch inference on GPUs. If you can predict 10 tokens for the price of 1 it should allow you to do tree of thought much more efficientl
174.
▲
by
valine
3y ago
Hopefully this new model will be a step beyond what you can do with animatediff
175.
▲
by
valine
3y ago
I have seen them, the workflows to create those videos are extremely labor intensive. Control net lets you maintain poses between frames, it doesn’t solve the temporal consistency of small details.
176.
▲
by
valine
3y ago
The rate of progress in ML this past year has been breath taking. I can’t wait to see what people do with this once controlnet is properly adapted to video. Generating videos from scratch is cool, but the real utility of this will be the te
177.
▲
by
valine
3y ago
The largest rocket ever built got to space for the first time, powered by 33 full-flow staged combustion engines, considered to be the holy grail of engine design. First successful use of 33 engines firing in unison. First full-flow engine
178.
▲
by
valine
3y ago
The model page is the only info I’ve found on it. As far as I can tell there’s no paper published on the technique. In the “Merge Process” section they at least give the layer ranges. https://huggingface.co/alpindale/go
179.
▲
by
valine
3y ago
Steve Jobs famously had two iPhone teams working on concepts in parallel. It was click wheel vs multi-touch. Shockingly the click wheel iPhone lost.
180.
▲
by
valine
3y ago
Microsoft has the GPT4 weights, ChatGPT will be available in some form regardless of what happens to OpenAI.
More ›