Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ijk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
18 ms
·
211.
▲
by
ijk
2y ago
Answering the question: "what is the impact of code data used in pre-training on a large variety of downstream tasks beyond code generation?"
212.
▲
Automated Design of Agentic Systems
(shengranhu.com)
2 points
by
ijk
2y ago
|
0 comments
213.
▲
Lottery Ticket Adaptation
(github.com)
1 points
by
ijk
2y ago
|
0 comments
214.
▲
by
ijk
2y ago
And revenue has dropped through the floor: https://fortune.com/2024/08/15/elon-musk-tesla-stock-sale-tw... https://mashable.com/article/twitter-x-revenue-falls-x-payme...
215.
▲
by
ijk
2y ago
Where are you getting 2 billion from? The original CLIP paper says: > We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations
216.
▲
by
ijk
2y ago
> Sure, you can check if it's mathematically coherent, but that tells you nothing about whether it describes the physical world correctly. This is a very good point I think a lot of people miss. (Including some who should know bette
217.
▲
by
ijk
2y ago
I'm not sure how "you should never use CrowdStrike" is an argument in CrowdStrike's favor. I guess you're saying they shouldn't have outsourced in the first place? Which does sound like the correct conclusion i
218.
▲
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse
(arxiv.org)
1 points
by
ijk
2y ago
|
0 comments
219.
▲
by
ijk
2y ago
The problem is that some people are running around and saying they are gods. Which I wouldn't care about, but an alarming number of people do believe that they can predict facts.
220.
▲
by
ijk
2y ago
One additional problem with people who write breathless tutorials about doing things with AI is that they are more likely than average to have been written with ChatGPT. Which, given the knowledge cutoff for most models, is not where I'
221.
▲
by
ijk
2y ago
Quite. The big question at the time was "how much data do we need to train GPT-3 equivalent models". Open models had failed to live up to GPT performance, even ones with a massive number of parameters. So getting results that sugg
222.
▲
by
ijk
3y ago
Depends on which branch of RTS design you want to explore. Age of Empires 2 still has a significant player base, for example, so it's both a genre-defining classic and modern ranked multi-player. Other RTS design traditions include Tot
223.
▲
by
ijk
3y ago
I've seen succesful projects that add a constraint solver as a layer in a neural network, so it's potentially something that could be integrated at an even deeper level than our current finetuning for tool use. It's not a pri
224.
▲
by
ijk
3y ago
Speaking as someone who has done VFX professionally, I think Sora is closer about 30% there. It's really not constituted in a way that makes it easy to combine with existing techniques. You either need to come up with a way to make the
225.
▲
by
ijk
3y ago
For translation, you're probably better off with a model that's specifically designed for translation, like MADLAD-400 or DeepL's services.
226.
▲
by
ijk
3y ago
Not to mention that most species evolved to fit particular environments and don't do well when you rapidly alter those environments far outside the parameters they tend to live in.
227.
▲
by
ijk
3y ago
Yeah, there's a research paper idea sitting there for someone who wants to run the numbers on some more ablation tests and see if there are any unwanted side effects. Though if it gets the claimed performance on non-finetuned data, you
228.
▲
by
ijk
3y ago
Soft prompts for language models are still viable, but the most accessible version of open source soft prompt training was a TPU-based implementation that depended on features of Google Colab that were no longer available. You can modify it
229.
▲
by
ijk
3y ago
Turns out that dynamic NTK-Aware scaling further improves on that, with perplexity performance equal to or better than non-scaled at all context lengths: https://www.reddit.com/r/LocalLLaMA/comments/14mrgpr&
230.
▲
by
ijk
3y ago
Funny coincidence, but Fortnite now has Verse, a scripting language that's pushing the envelope and using Haskell-influenced ideas to do distributed computation.
231.
▲
by
ijk
3y ago
I think the most common complaint is that the interviews don't test technical skills, they test things that look like they're adjacent to technical skills but aren't the kind of work that you'd actually be doing.
232.
▲
by
ijk
3y ago
"Okay with paying for it" gives you a wide range of options. Most of the open source stuff people are talking about is things like running a quantized 33B parameter LLaMA model on a 3090. That can be done on consumer hardware, but
233.
▲
Scaling Data-Constrained Language Models
(arxiv.org)
50 points
by
ijk
3y ago
|
5 comments
234.
▲
Improving Factuality and Reasoning in Language Models Through Multiagent Debate
(arxiv.org)
2 points
by
ijk
3y ago
|
1 comments
235.
▲
by
ijk
4y ago
You are probably very aware of it, but just to highlight the importance of this for people who aren't aware: data duplication degrades the training and makes memorization (and therefore plagiarism, in the technical sense) more likely.
236.
▲
by
ijk
5y ago
I agree that the quantum mechanics terms are often more confusing than helpful here. To the point that I wrote a journal article to try to demystify how the algorithm works [1]. If you're familiar with constraint solving, "superpo
237.
▲
by
ijk
10y ago
> You lose copyright if you fail to put a notice on it? Only for pre-1978 things.
238.
▲
by
ijk
10y ago
They're beta-testing .NET 4.6 in Unity 5.5: https://forum.unity3d.com/threads/upgraded-mono-net-in-edito...
239.
▲
by
ijk
10y ago
The Amazon rainforest would die. [1] How much of it actually would die is obviously depends on factors like how big the replacement lake is and how much it affects the phosphorus, but I'm constantly amazed by how seemly separate system
240.
▲
by
ijk
10y ago
Apparently, iTunes no longer depends on QuickTime.
More ›