Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pico_creator
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
pico_creator
2y ago
That hurts - i used those GPUs before at their peak Now any random GPU in the computer store murders it
32.
▲
by
pico_creator
2y ago
Exactly, it covered in the article that there is a segmentation happening via GPU cluster size. Is it big enough for foundation model training from scratch = ~$3+ Otherwise it drops hard Problem is "big enough" is a moving goal po
33.
▲
by
pico_creator
2y ago
When I did the screenshot a month ago, it wasn't public info yet. Now its public: SFCompute list it on their main page - https://sfcompute.com/ And they are *not* the only one
34.
▲
by
pico_creator
2y ago
Problem is - u can find some at $1
35.
▲
by
pico_creator
2y ago
If we $2 H100 this year or next. Either AI is super dead, or a new alien GPU rained from the sky
36.
▲
by
pico_creator
2y ago
Covered in the article. They are below MSRP essentially
37.
▲
by
pico_creator
2y ago
I’m hopping more for an open weights AI boom With cheap compute for everyone to finetune :)
38.
▲
by
pico_creator
2y ago
Hey there, Twitter author / guy from the RWKV team. Some updates: Its also arriving to windows 10 (so that's 1.5B deploys) Our main guess would be the copilot beta features being tested - local copilot - local memory recall And it
39.
▲
by
pico_creator
2y ago
Thanks! We had to rewrite lots of code to optimize for the hot-swap load speeds. So we prioritized llama as it was the most popular group on Hugging Face. And RWKV (which is what we work on in open source space) But other architectures are
40.
▲
Show HN: Run any Llama model finetune and more, instantly
(featherless.ai)
7 points
by
pico_creator
2y ago
|
2 comments
41.
▲
by
pico_creator
3y ago
nice! - so once sora is open for you folks, it would be huge jump
42.
▲
by
pico_creator
3y ago
Are you all using stability based diffusion models? What are your teams thoughts on SORA ?
43.
▲
by
pico_creator
3y ago
Thanks for flagging that out, swapped it out with another map I found, that shows different colors for different thresholds So its not "as strict"
44.
▲
by
pico_creator
3y ago
Yea, thats why I focused only on the top 25 languages, despite the model being trained for 100+ languages. Was not confident, on the languages beyond the 25th cut-off, until we build better datasets (which we are in works on with various re
45.
▲
by
pico_creator
3y ago
Yea, the map is poorly generated IMO as well - sadly there is only a few online tools online i could fine that "help me highlight, places where this list of languages is supported" So if we want a better map, we might need to redo
46.
▲
by
pico_creator
3y ago
Currently the main policy is only around copyright - and nothing about AI safety: https://www.linuxfoundation.org/legal/generative-ai Also in the full power of opensource, if LF really force something the group disagre
47.
▲
by
pico_creator
3y ago
makes sense, considering there lacks an open std to port back and forth. people stick to datadog / elastic / etc passionately, to avoid "redoing their charts" also low hanging fruit to evals prompt and check (does it JSO
48.
▲
by
pico_creator
3y ago
LLMOps / PromptOps, seems to be a new monitoring tool segment which is heating up. And pretty much essential as AI models are constantly being updated "in production" these days. Seeing that you interviewed LangChain, which
49.
▲
RWKV – worlds first OSS AI model to join Linux Foundation
(twitter.com)
9 points
by
pico_creator
3y ago
|
2 comments
50.
▲
by
pico_creator
3y ago
plain simple apache 2 licensing, no strings attached - do whatever you want with it like any other Linux OSS software
51.
▲
by
pico_creator
3y ago
I feel that langchain is the early jQuery of the AI world. It is useful if you are abstracting multiple models, less so if you are only using openAI with no vectorDB (where it may become an overhead) It is opinionated, bloated, but it works
52.
▲
by
pico_creator
3y ago
Thanks for clarifying, and pointing out where docs can be improved. For most parts I am trying to ensure everything do get properly reorganised under the official wiki page here. I have just added a guideline of sort to help better navigate
53.
▲
by
pico_creator
3y ago
Btw, I also realise on rereading, both statements can be true at the same time. In the podcast as well, I was being very transparent that I am happily using transformers (Salesforce codegen) in production for my primary work use cases, and
54.
▲
by
pico_creator
3y ago
Ideas and trends move much slower, than we think, especially outside of silicon valley, or the current bubbles we as individual are in. We as humans have a tendency of not wanting things to change, nor accept new ideas that challenge our ex
55.
▲
by
pico_creator
3y ago
As a user from the early internet forum era, I do find it a step back as well. But unfortunately, the people as large has kinda spoken, and voted by their usage. And ultimately for community engagement, their decision is something I have to
56.
▲
by
pico_creator
3y ago
Agreed that as a whole, we do need more public benchmarks (both transformers and RWKV), with much longer sequence length. I view this as a case of the "benchmarks" and "datasets" not keeping up with the current progress,
57.
▲
by
pico_creator
3y ago
Do let me know your use case where it perform poorly? And I would look into getting it better supported in upcoming datasets we train with =) ( I mean this genuinely, we have an entire discord channel for #failed-task to help us keep track
58.
▲
by
pico_creator
3y ago
What other architectures should we be taking note of (that stability is supporting)? Especially in the text space that you feel will change the current transformer LLM paradigm?
59.
▲
by
pico_creator
3y ago
Quantisation thankfully is applicable to RWKV as much as transformers. Most notably in our RWKV.cpp community project: https://github.com/saharNooby/rwkv.cpp Tooling/Ecosystem is something that I am actively worki
60.
▲
by
pico_creator
3y ago
It is important to note the claims are made for similar dataset & params. If you like some independent verification outside eleuther circle. There is microsoft paper on retnet: https://arxiv.org/abs/2307.08621 And
More ›