Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kir-gadjello
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
kir-gadjello
7mo ago
I don't think it's strictly better than GLM 5, more like they are peers (but in math competitions StepFun is stronger than most), and in my experience have similar coding/bugfix ceiling where world knowledge is not the decidi
2.
▲
by
kir-gadjello
7mo ago
Much worse - from my experience minimax is not suitable for high autonomy on hard projects. The real distant second in my experience is mimo flash v2 (but I did not try the latest version, might be closer to parity). I would not use minimax
3.
▲
by
kir-gadjello
7mo ago
Yeah, my github is in the profile. Soon (tm). Feel free to follow.
4.
▲
by
kir-gadjello
7mo ago
I use pi, but I'm almost done writing a better alternative that doesn't have pi's stability issues. 80K Rust SLOC and a few hundred tests btw.
5.
▲
by
kir-gadjello
7mo ago
They are not equivalent 1:1, esp. in knowledge coverage (given OOM param size difference) and in taste (Sonnet wins, but for taste one can also use Kimi K2.5), but in my hardcore use (high-performance realtime simulations of various kinds)
6.
▲
by
kir-gadjello
7mo ago
I think we are at this point where the hard ceiling of a strong model is pretty hard to delineate reliably (at least in coding, in research work it's clearer ofc) - and in a good sense, meaning with suitable task decomposition or a tes
7.
▲
by
kir-gadjello
7mo ago
I just use openrouter, it's free for now. But I would pay 30-100$ to use it 24/7.
8.
▲
by
kir-gadjello
7mo ago
Respectfully, from my experience and a few billions of tokens consumed, some opensource models really are strong and useful. Specifically StepFun-3.5-flash https://github.com/stepfun-ai/Step-3.5-Flash I'm workin
9.
▲
by
kir-gadjello
2y ago
While llama3-8b might be slightly more brittle under quantization, llama3-70b really surprised myself and others[1] in how well it performs even in the 2..3 bits per parameter regime. It requires one of the most advanced quantization method
10.
▲
by
kir-gadjello
2y ago
Synthetic data researchers, connoisseurs and artisans, obviously. Feedback loop architects. Mindscaping artists.
11.
▲
by
kir-gadjello
3y ago
It is quite likely GPT-4 uses one or even two sparsity approaches on top of each other (namely, coarse grained switch transformer-like and fine grained intra-tensor block sparsity), if you look at the openly available contributors' res
12.
▲
by
kir-gadjello
3y ago
It is a guess informed by some familiarity with the literature and by going over the papers authored by researchers credited in the OpenAI's "GPT-4 contributors" web page. I have an expanded list of foundational research that
13.
▲
by
kir-gadjello
3y ago
It allows for modifications and commercial use: https://creativecommons.org/licenses/by-sa/4.0/ >You are free to: >Share — copy and redistribute the material in any medium or format >Adapt — remix, t
14.
▲
by
kir-gadjello
3y ago
Impressive model, thank you for releasing it under a business-friendly license! Have you considered using Google's sparse "scaling transformer" architecture as the base? Even at 3B scale it can generate 3-4x more tokens per F
15.
▲
by
kir-gadjello
3y ago
Thank you for developing the pipeline and amassing considerable compute for gathering and preprocessing this dataset! I'm not sure if this is the right place to ask about this, but could you consider training an LLM using a more advanc
16.
▲
by
kir-gadjello
4y ago
That's cool, thanks for noting, Alan! Would you mind adding a reference link to the source, so that other people could visit my blog? I'm just starting out with blogging, it would help me to get more readers and feedback on this d
17.
▲
by
kir-gadjello
4y ago
This is cool, but SSD read bandwidth is still the bottleneck. On my non-mac machine it still takes several seconds to load the model.
18.
▲
by
kir-gadjello
4y ago
I think envying closed source closed weights proto-AGI systems is counterproductive. We have opensource models with available weights that are almost as powerful: https://huggingface.co/maderix/llama-65b-4bit (nonfree
19.
▲
by
kir-gadjello
4y ago
This document doesn't contain the architecture and training details of GPT-4. As an engineer, these details would be the most interesting part of it! Driven by interest in GPT-4 and cutting edge LLMs I studied the research literature a
20.
▲
by
kir-gadjello
4y ago
No, a typical LLM is a pure function of its input, if you (and not the LLM hosting company) control all of input context , and if your sampler uses pseudorandom number generator. But you could create an LLM for which it wouldn't be th
21.
▲
by
kir-gadjello
4y ago
Given a list of contributors it's not that hard to reverse-engineer the specific engineering choices made by looking up their publication history. My analysis https://kir-gadjello.github.io/posts/gpt4-some-technica
22.
▲
by
kir-gadjello
4y ago
It could be done in a dozen ways. One beautiful method is just using the xPos positional embedding pioneered by Microsoft and scale the context window size at runtime (even better if your attention is subquadratic - again there is a dozen o
23.
▲
by
kir-gadjello
4y ago
If you have questions about my rationale for this or that technique included in the list, please, ask! For example, I think Google's paper "Sparse is enough for scaling transformers" was very underrated, as it provided more t
24.
▲
by
kir-gadjello
4y ago
It's no problem to put model's architecture and even some python code into the generous 32k context window, the real problem seems to be as you say "awareness" - at least the facet of it that'd allow to answer compl
25.
▲
GPT-4 architecture: what we can deduce from research literature
(kir-gadjello.github.io)
10 points
by
kir-gadjello
4y ago
|
6 comments
26.
▲
by
kir-gadjello
4y ago
As the discussion of GPT-4 heats up, the absence of details on its technical implementation becomes only more glaring. As an engineer, I have not learned anything applicable I haven't known yesterday from the newest OpenAI publication!
27.
▲
by
kir-gadjello
4y ago
Charitably speaking the researchers had little time to execute this, so they just ended up using the well known OpenAI API. Still, it would be very useful if someone used LLaMA-65B instead of text-davinci-003 here. Someone should ask the re
28.
▲
by
kir-gadjello
4y ago
If you are interested in the infrastructure-level details of how similar models are trained by lesser known groups, take a look at this paper: https://arxiv.org/abs/2204.06745 Quotes from the paper: Our model is train
29.
▲
by
kir-gadjello
4y ago
"Accomodate" is the word to scrutinize here. Yes, it will cost a lot to outright buy physical HPC infrastructure to train and infer a series of large models deployed for customers all over the globe. No, it won't cost nearl
30.
▲
by
kir-gadjello
4y ago
Please do it, people shouldn't put up with the apathetic siloed status quo. I'm sure people will find all sorts of beneficial uses for these models they are going to run on their own hardware!
More ›