Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
icyfox
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
Should you fine tune for JSON output?
(freeman.vc)
2 points
by
icyfox
3y ago
|
0 comments
62.
▲
by
icyfox
3y ago
I know this is what most of the SpaceX mission control systems do. They're set up with ~3x/5x redundancy for effectively consumer hardware. If they detect memory misalignment between 1/3 of the nodes, they'll power cycle
63.
▲
by
icyfox
3y ago
A friend of mine created a bot to do basically this, except it also looks at the current page rank associated with researchers recommending that paper. I've seen a lot of good looking papers (decent school/group/conference su
64.
▲
by
icyfox
3y ago
Just for connecting to an existing service: https://github.com/ollama/ollama-python/blob/main/ollama/_cl...
65.
▲
by
icyfox
3y ago
Well, Lora works just because it's a low rank approximation of full updates - much in the same way that SVD works, and regular gradient updating works. It delivers good results by both acting as a regularizer and by allowing larger mod
66.
▲
by
icyfox
3y ago
Like most things in ML, the answer of which layers to use come down to empirical evidence more than theory. In a typical Lora training pipeline, you freeze the contents of the base model and just adjust the Lora layers. The more layers you
67.
▲
by
icyfox
3y ago
I've done some research involving parameterized summarization recently. GPT-4 does a really good (like, pretty outstanding) job in following the tone and density of a few one-shot example inputs. It's a lot more difficult to encou
68.
▲
by
icyfox
3y ago
Favorite part of this piece: > We were initially skeptical whether we would get to 10,000 results. But with nightly leaderboard gamification, we managed to break 15,000 results within a week. Out of fear of eating into our productivity,
69.
▲
by
icyfox
3y ago
Right now it seems that they've trained a true foundational model. It can recreate its training inputs, but that only matters as far as downstream performance goes. Based on a brief scanning of the repo & website: - Their training
70.
▲
by
icyfox
3y ago
Well thankfully my day job is actually ML research. :) Lines of code is not strictly speaking a bottleneck for the next generation of models, but it ties with other objectives that are: researcher productivity, hardware efficiency, and mode
71.
▲
by
icyfox
3y ago
I'm a big believer in the approach you're laying out too. Bugs are much easier to diagnose, crash reporting is more straightforward, and you don't need augmented services to consolidate everything at the output layer. That sa
72.
▲
by
icyfox
3y ago
For those hearing about Cerebras for the first time, they make a chipset that's similar to a GPU in matrix multiplication speed but way bigger (a whole wafer) so it can fit more transistors and memory onto one chip. They achieve this s
73.
▲
by
icyfox
3y ago
Common wisdom in most industries is to release bad PR announcements on a Friday and good ones towards the start of the week. It's interesting how the advent of twitter communication has shifted the ML ecosystem to publishing work whene
74.
▲
by
icyfox
3y ago
The only production case I've seen of someone switching to a vector database totally in-lieu of a RDBMS is when their raw data lives in parquet / arrow files and they need fast read-only searches in production. Then a periodic ind
75.
▲
by
icyfox
3y ago
This is the best I've found as well. They also have a new offering (different internal model, similar quality evidently) that I've used back when it was in a free unlimited trial. Not sure the status right now: https://
76.
▲
by
icyfox
3y ago
Sometimes gradients are small but meaningful, if you constrain them to too few bits / degrees of freedom they'll be unable to backprop successfully. This can hamper training and therefore results quality. You can also think about
77.
▲
by
icyfox
3y ago
Practically speaking, most models today infer at 8bit or 16bit (sometimes, rarely 32). You don't see an empirical lift at more bits of precision. Size of the memory is far more important.
78.
▲
by
icyfox
3y ago
This will be an interesting test to see how fast you can bootstrap GPT-4 level performance with unlimited funds and talent that already has deep knowledge of the internals. With the initial adoption of ChatGPT alongside Copilot, OpenAI'
79.
▲
by
icyfox
3y ago
It fits a pretty nice niche between small data (32GB in-memory, pandas level) and truly huge data where the disk files can't even fit on one machine (2TB+). Most data that people want to process on a daily cadence fit somewhere in betw
80.
▲
Lorax: Serve 100s of Fine-Tuned LLMs in Production for the Cost of 1
(github.com)
5 points
by
icyfox
3y ago
|
1 comments
81.
▲
by
icyfox
3y ago
I've actually seen this work quite well, in terms of voice recognition. The commands are specific enough and a lavalier / directed microphone can filter out most of the surrounding noise. Whether other people will be as happy with
82.
▲
by
icyfox
3y ago
I'm really happy that Montreal made this work. We need more alternative examples of public transit build-outs in the west that actually work and come in on-budget. Copenhagen is another good exemplar here with their public/private
83.
▲
by
icyfox
3y ago
My general rule of thumb at the moment is: - Tasks requiring knowledge synthesis between multiple different datapoints (in your case wiki pages) should be fine-tuned so the model's able to do some basic chain of reasoning to reach a ne
84.
▲
RFDs: A Tool for Discussion
(oxide.computer)
1 points
by
icyfox
3y ago
|
0 comments
85.
▲
by
icyfox
3y ago
I'm not sure I'd advocate for writing a static site generator, although I'm certainly guilty of writing a few myself. Instead I always encourage people who are trying to start blogging to do the writing first. Figure out a wo
86.
▲
by
icyfox
3y ago
Fast can also measure some more connection metadata if you click "show more info". A bit annoying it requires the extra step to start measuring ping and upload speed, but it is there.
87.
▲
by
icyfox
3y ago
I have a similar harness going for my recent experiments, except instead of hosting with huggingface I have a dataframe with pointers to the files on S3 and then just download them during local preprocessing. Every time I see DVC mentioned
88.
▲
by
icyfox
3y ago
Last time I looked into this their training code is kept internal and they're just releasing the model weights.
89.
▲
by
icyfox
3y ago
There was an Information article a few weeks ago on Apple increasing its daily spend to millions of dollars in ML training costs per day. With their historic desire for control over the whole stack, I'd put money on them rolling their
90.
▲
by
icyfox
3y ago
That's the interesting part of this story; Congress didn't think this requirement existed, neither did the lobbyists. But the language that congress adopted (with the consultation of this lobbying group) made the agencies _think_
More ›