Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
valine
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
valine
2y ago
The landing legs are heavy. SpaceX would rather have more payload than carry legs to space on every launch. Also landing on the launch pad means you don’t have to transport the rocket. Just have the arms set it down and you’re ready to laun
62.
▲
by
valine
2y ago
Not really the first paper is just fine-tuning on synthetic data. The second paper doesn’t optimize the model weights.
63.
▲
by
valine
2y ago
This period of model scaling at all cost is going to be a major black eye on the industry in a couple years. We already know that language models are few shot learners at inference time, and yet OpenAI seems to be happy throwing petaflops o
64.
▲
by
valine
2y ago
Demonstrating good performance from a non-transformer based architecture is cool. I agree though these particular models aren’t that useful given the current landscape. I think the intent here is probably to justify training a larger 400B m
65.
▲
by
valine
2y ago
That’s a strange perspective. AEB isn’t 100% reliable but we still allow it on the roads because it’s a net increase in safety. When used correctly FSD seems to be a net increase in safety as well. You can make the case they should do a bet
66.
▲
by
valine
2y ago
I did, was about to delete it but then you commented haha.
67.
▲
by
valine
2y ago
Edit: replied to wrong comment
68.
▲
by
valine
2y ago
Reflection 70B is a scam. The creator was just routing requests to Claude. https://old.reddit.com/r/LocalLLaMA/comments/1fc98fu/confirm...
69.
▲
by
valine
2y ago
The model performance is driven by chain of thought, but they will not be providing chain of thought responses to the user for various reasons including competitive advantage. After the release of GPT4 it became very common to fine-tune non
70.
▲
by
valine
2y ago
A single document doesn’t contribute much to the loss during base training. LLMs can absolutely memorize text if it’s duplicated in the training set, not so much if the document in question is obscure.
71.
▲
by
valine
2y ago
NASA wanted to understand if the malfunctioning thrusters could operate within a margin error for safety. Boeing spent a month running a study that produced non-definitive answer as to the root cause of the problem. Had Boeing definitively
72.
▲
by
valine
2y ago
This is such a natural extension to LLMs. I’m shocked it hasn’t been tried before. When I ask a diffusion model to generate a chessboard, I’d expect the pieces to be placed randomly. We are getting closer to image generators that not only k
73.
▲
by
valine
2y ago
Just because you have the dataset doesn't mean you can generate a reference. Let's say I hand you a potato salad recipe and a copy of the entire internet. Say you somehow extract all potato salad recipes from the dataset (non triv
74.
▲
by
valine
2y ago
If you have the model weights you have roughly the same opportunities as the company that trained the model. The code you need to run inference on the Llama weights is very much open source. The only thing you're missing out on is the
75.
▲
by
valine
2y ago
Not sure what this has to do with anything. The paper I was commenting on is using diffusion models to parse raw light hitting the sensor as an alternative to a glass lens. No one wants an image generation model hooked up to weather data— t
76.
▲
by
valine
2y ago
Not sure I understand the question. This paper is about using diffusion models to reconstruct usable images from raw sensor data. The diffusion model in essence replaces the lens.
77.
▲
by
valine
2y ago
I get the feeling that lens free cameras are the future. Obviously the results here are no where near good enough, but given the rapid improvement of diffusion models lately the trajectory seems clear. Would love to lose the camera bump on
78.
▲
by
valine
2y ago
The codebase to do the training is way less valuable than the weights for the vast majority of people. Releasing the training code would be nice, but it doesn't really help anyone but Meta's direct competitors. If you want to trai
79.
▲
by
valine
2y ago
LLMs are bad at counting things just in general. It’s hard to say whether the failures here are vision based or just an inherent weakness of the language model.
80.
▲
by
valine
2y ago
We don’t store color data in full precision usually. People aren’t sensitive to all colors equally, the less sensitive the eye is to a particular color the more efficiently you can store it. You can also typically discard high frequency dat
81.
▲
by
valine
2y ago
Or alternatively a lot of energy is wasted answering simple questions. The whole point of the transformer is to take words and iteratively, layer by layer, use the context to refine their meaning. The vector you get out is a better represen
82.
▲
by
valine
2y ago
It’s a neat approach. VRAM is obviously the major concern here since it sounds like the parameter count grows as you insert more facts. I’m also curious how well the facts communicate with each other. A major problem with RAG is that the mo
83.
▲
by
valine
2y ago
No one knows how it will all shake out. I'm personally skeptical scaling laws will hold beyond GPT4 sized models. GPT4 is likely severely undertrained given how much data facebook is using to train their 8B parameter models. Unless Ope
84.
▲
by
valine
2y ago
Llama 3 8B captures that 'magic' fairly well and runs on a modest gaming PC. You can even run it on an iPhone 15 if you're willing to sacrifice floating point precision. Three years from now I full expect GPT4 quality models
85.
▲
by
valine
2y ago
Happy to see the control center change. Look up for control center is a bad experience.
86.
▲
by
valine
2y ago
Llava1.6, IntenVL, CogVLM2 can all do OCR with nothing but tiled image embeddings and an LLM. Feeding in OCR results from tesseract improves the reliability of the transcript, especially for long strings of random characters, but it’s not s
87.
▲
by
valine
2y ago
Very curious how it performs on OCR tasks compared to InternVL. To be competitive at reading text you need tiling support, and InternVL does tiles exceptionally well.
88.
▲
by
valine
2y ago
The final layer of the transformer prior to the logits pretty much does what you want already. The KV of the final layers are taken into account when generating the final hidden state for your new token. The first layers of the model are re
89.
▲
by
valine
2y ago
I assume so, you can’t really tokenize audio, at least not high fidelity audio. Audio models like Bark don’t output logits from what I understand. For true multimodal output you’d need a model that can output both logits and audio embedding
90.
▲
by
valine
2y ago
It’s definitely multimodal input. Passing Clip embeddings to an LLM is nothing new, and that’s really all you need for document understanding. It’s almost certainly the same thing for audio. They would have trained a dual encoder that maps
More ›