Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gamegoblin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
91.
▲
by
gamegoblin
3y ago
I follow ~200 extremely high quality people on twitter, never use the algorithmic timeline -- just the "Following" tab, and it is easily 10x more signal to noise than HN for me, I'm almost upset I didn't start using it s
92.
▲
by
gamegoblin
3y ago
AWS has written parts of major services (EC2, Lambda, S3, etc) in Rust.
93.
▲
by
gamegoblin
3y ago
This is due to the tokenization scheme. These LLMs can’t see individual characters, they see chunks of characters that are glued together to form tokens. It’s impressive that they understand individual-character-level questions as well as t
94.
▲
by
gamegoblin
3y ago
In this thread we’re talking about gpt-3.5-turbo-instruct, not GPT4
95.
▲
by
gamegoblin
3y ago
I'm assuming they will price it the same as normal gpt-3.5-turbo. I won't use it if it's more than 2x the price of turbo, because I can usually get turbo to do what I want, it just takes more tokens sometimes. Have you tried
96.
▲
by
gamegoblin
3y ago
One is tuned for chat. It has that annoying ChatGPT personality. Instruct is a little "lower level" but more powerful. It doesn't have the personality. It just obeys. But it is less structured, there are no messages from user
97.
▲
by
gamegoblin
3y ago
Biggest news here from a capabilities POV is actually the gpt-3.5-turbo-instruct model. gpt-3.5-turbo is the model behind ChatGPT. It's chat-fine-tuned which makes it very hard to use for use-cases where you really just want it to obey
98.
▲
by
gamegoblin
3y ago
It's about as useful as having an obedient 2nd year CS student with access to a juypter notebook. Which is to say, pretty useful for certain stuff! You can do flows like: - Upload CSV of stock data - "Graph the closing price each
99.
▲
by
gamegoblin
3y ago
Misremembered, the main thrust of the comment still stands, the 100K context window isn't "real", it would be absurdly expensive to do it for real. They are using a lot of approximation tricks to get there.
100.
▲
by
gamegoblin
3y ago
I agree with you that learning certain things is wasteful. For instance, one could imagine an RNN that learned to do some approximation of tree search for game playing Chess and Go. But we have very good reason to think that tree search i
101.
▲
by
gamegoblin
3y ago
I think the jury is still out if these will actually scale to ultra-long language understanding sequences. KWKV, for example, is still trained like GPT, but is architected so it can be run as an RNN during inference time. This is awesom
102.
▲
by
gamegoblin
3y ago
Better phrasing would have been "the important tokens are roughly evenly distributed throughout the text", that was the intended reading.
103.
▲
by
gamegoblin
3y ago
Yes, Claude 1M is using all sorts of approximation tricks to get that 1M context window. IMO this is actually quite deceptive marketing.
104.
▲
by
gamegoblin
3y ago
The benefit of "traditional" O(N^2) transformer attention is you correlate every token to every other token. So, in the limit, your network won't "miss" much. When you abandon O(N^2) attention, you are forced to sta
105.
▲
by
gamegoblin
3y ago
The biggest plant in the US, near San Diego (Carlsbad), cost $1B to build and produces 200,000 m^3 per day. Israel's biggest desalination plant, Sorek, cost $400MM and produces over 600,000 m^3 per day, so 40% the cost and 300% the wat
106.
▲
by
gamegoblin
3y ago
I was surprised how cheap it is. Desalinated water costs ~50 cents per 1000 liters [1]. That's about the same amount of water as a typical American household uses per day. 50 cents per day for a fully desalinated water supply is... i
107.
▲
by
gamegoblin
3y ago
It still feels very clear to me, I feel like the people debating this have probably never written and trained an LLM. Consider a simpler case: a small neural network that takes 2 numbers and adds them together, producing 1 number as output.
108.
▲
by
gamegoblin
3y ago
This is what we did for our KV store at S3: https://www.amazon.science/publications/using-lightweight-fo... Using https://github.com/awslabs/shuttle which works on our real Rust code.
109.
▲
by
gamegoblin
3y ago
The point of my comment is that even the distribution represents intelligence. If you give it a tricky Yes/No question that results in a distribution that's 99.97% "Yes" and negligible values for every other token, that
110.
▲
by
gamegoblin
3y ago
GPT can give a single Yes/No answer that indicates a fair amount of intelligence for the right question. No iteration there. Just a single pass through the network. Hofstadter is surprised by this.
111.
▲
by
gamegoblin
3y ago
I think it has to do with the training regime and fixed-computation time nature of feedforward neural networks. Recurrent neural networks have the recursion as part of the training regime . GPT only has auto-regressive "recursion"
112.
▲
by
gamegoblin
3y ago
Just being a part of any auto-regressive system does not contradict his statement. Go look at the GPT training code, here is the exact line: https://github.com/karpathy/nanoGPT/blob/master/train.py#L12...
113.
▲
by
gamegoblin
3y ago
All the other responses to you at the time of writing this comment are confidently wrong. Definition of Feedforward (from wiki): ``` A feedforward neural network (FNN) is an artificial neural network wherein connections between the nodes do
114.
▲
by
gamegoblin
3y ago
Email in bio if anyone reads this, been on waitlist basically since launch
115.
▲
by
gamegoblin
3y ago
If you don't have a preference between seeing Vermont's beautiful hilly forests, and seeing the same forests with a tacky billboard drawing your gaze, I don't really know what to tell you other than that most people feel diff
116.
▲
by
gamegoblin
3y ago
I agree with you that connecting consumers to products is valuable. I've bought several things from instagram ads (has the best targeting of any platform, IMO), but OP specifically mentioned public places. Advertising in public places
117.
▲
by
gamegoblin
3y ago
One could bolt such a system on top of an LLM. An LLM is "just" a document completer. Given text, it predicts the following text. So there is nothing stopping you from bolting on a system which works like: - Given the current stre
118.
▲
by
gamegoblin
3y ago
Beam search is like a breadth-limited breath-first search. This is more akin to a depth-first search where you give the model the ability to backtrack.
119.
▲
by
gamegoblin
3y ago
Just listening to it, it's subjectively not better, but if it's > 10x faster/cheaper, I would use it anyway -- it's good enough to be listenable. Eleven Labs is the first voice synthesis that is good enough that I
120.
▲
by
gamegoblin
3y ago
I've used a mac for 10 years now, and TIL The Windows snapping UI is much more intuitive, though.
More ›