Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
eldenring
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
91.
▲
by
eldenring
3y ago
GPT-3.5 is much, much smarter than Llama2. Its not nearly as close as the benchmarks make it seem.
92.
▲
by
eldenring
3y ago
Isn't this just ignoring all the lessons we learned from pretraining large models? Seems like a different flavor of the bitter lesson.
93.
▲
by
eldenring
3y ago
Yes kafka is definitely in an awkward latency spot.
94.
▲
by
eldenring
3y ago
I'd guess these model's understand works more closely to people so encoding in text is more token efficient and things like comments help. Also syntax seems a lot easier to understand for them than semantics/logic. If you
95.
▲
by
eldenring
3y ago
Since the cost per token of a forward/backward pass is so high on the models, the relative cost of a training set is small, and bandwidth (not talking about interconnect bandwidth) is even cheaper relatively.
96.
▲
by
eldenring
3y ago
Multiple servers adds a pretty hefty layer of networking, orchestration, security, failure handling (even in the datacenter), and serialization - even if the business logic stays mostly the same.
97.
▲
by
eldenring
3y ago
I'd be surprised if 99.9% of the targeted userbase of this app cares about this. I don't.
98.
▲
by
eldenring
3y ago
This is absolutely not true, having shared memory or even just being able to communicate over local os pipes is massivley simpler than introducing a network.
99.
▲
by
eldenring
3y ago
I mean you'd get the same number of page faults with both examples, so I think the discrepency is explainable by caches/cpu prefetching.
100.
▲
by
eldenring
3y ago
This is one of the most common arguments for non reference counted GCs
101.
▲
by
eldenring
3y ago
Lockless is not the same as lock-free, which is not the same as wait-free which seems to be what you are describing
102.
▲
by
eldenring
3y ago
It would be cool to see this compared with a more simple bin packing technique, or even just stuffing the most commonly used instructions at the top and those ordered by size.
103.
▲
by
eldenring
3y ago
Atomics scale very well if you are reading often and writing rarely.
104.
▲
by
eldenring
3y ago
Good article, but if i had to guess the subsequent L3 cache access after an increment is likely far overshadowed by the overhead of coherence messages, or the cores communicating between each other.
105.
▲
by
eldenring
3y ago
Doom is the same. It uses evil mode by default and SPC as the default leader key. Also in my experience it is a lot faster.
106.
▲
by
eldenring
4y ago
It's hard to blame interest rates for all the wealth in tech when you look at all the new products and innovations that have happened over the last 10 years.
107.
▲
by
eldenring
4y ago
Huh? pretty much every one of the recently successful AI startup employees (OpenAI, Anthropic, etc.) have had stints at Google Brain.
108.
▲
by
eldenring
4y ago
There might not be a "Z" token for some of these names. A made-up example is "Lukasz" might tokenize to ["Luk", "asz"], so the model doesn't have any notion of how words are actually spelled. I s
109.
▲
by
eldenring
4y ago
Could I get a source for that? Not that I don't believe you, but my napkin math puts the cost of training the 65b parameter model alone at a lot higher than 100k.
110.
▲
by
eldenring
4y ago
I see a bullet point saying Varies is for "random or non-constant state"
111.
▲
by
eldenring
4y ago
> Trained using the Chinchilla formula, these models provide the highest accuracy for a given compute budget. I'm confused as to why 111 million parameter models are trained with the Chinchilla formula. Why not scale up the training
112.
▲
by
eldenring
4y ago
Yes! I have been spending the last couple months pulling out completely unnecessary redis caching from some of our internal web servers. The only loss here is network latency which negligible when you're colocated in AWS. Postgres'
113.
▲
by
eldenring
4y ago
Would an ASIC really be better than GPUs? I'm not 100% sure but aren't low power GPUs essentially ASICs for matrix multiplication.
114.
▲
by
eldenring
4y ago
Maybe the output was long enough that the context window would miss your previous prompts?
115.
▲
by
eldenring
4y ago
Sounds like OCaml.
116.
▲
by
eldenring
4y ago
The ChatGPT API is using a smaller model than the original "GPT 3.5". The original, presumably larger, model is used by Legacy ChatGPT.
117.
▲
by
eldenring
4y ago
oh god sorry I didn't even realize. Its a bit sad that nowadays a comment like that can even be mistaken for a serious one.
118.
▲
by
eldenring
4y ago
what an evil and disgusting comment.
119.
▲
by
eldenring
4y ago
> Writing basic data structures isn't a niche, esoteric edge case Maybe it isn't an edge case (although it should be) it also isn't `easy` in a non GC'd language, and a huge source of memory bugs. I wouldn't say
120.
▲
by
eldenring
4y ago
> you can't (or don't want to) allocate memory dynamically. this crate is intended to allow all metrics storage to be declared in statics, for use in embedded systems and other no-std use-cases. Basically, the library just tr
More ›