Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
desideratum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
desideratum
6mo ago
Yes my findings and thoughts were pretty much identical. I actually think you can get something reasonable at 1.3B params with the correct training recipe, but definitely not at this compute/token budget. One thing I found was that the
2.
▲
by
desideratum
6mo ago
This is a gross simplification of the process - you would typically use order(s) of magnitude more data and compute, and a substantial amount of online reinforcement learning to elicit emergent tool use capabilities. Many recent OSS models
3.
▲
by
desideratum
6mo ago
I appreciate the kind words very much : )
4.
▲
by
desideratum
6mo ago
I see what you mean, but I disagree. I expect that Claude Code is backed by a separate post-train of Claude base which has been trained using the Claude Code harness and toolset.
5.
▲
by
desideratum
6mo ago
Oh I wouldn't be surprised. This is a sample from one of the OSS code datasets I'd used, which are all generated synthetically using LLMs. Data is indeed the moat.
6.
▲
by
desideratum
6mo ago
This is a great question. You definitely aren't training this to use it, you're training it to understand how things work. It's an educational project, if you're interested in experimenting with things like distributed t
7.
▲
Nanocode: The best Claude Code that $200 can buy in pure JAX on TPUs
(github.com)
219 points
by
desideratum
6mo ago
|
26 comments
8.
▲
Batrachochytrium Dendrobatidis
(en.wikipedia.org)
4 points
by
desideratum
7mo ago
|
0 comments
9.
▲
by
desideratum
7mo ago
Oh, and if you want to utilize 120Hz on the XDR display, you're going to have to replace your perfectly functioning Mac. > Mac models with M1, M1 Pro, M1 Max, M1 Ultra, M2, and M3 support Studio Display XDR at up to 60Hz. All other
10.
▲
by
desideratum
7mo ago
It's mind-boggling that Apple is considering the base 27 inch Studio Display with the same 4 year old panel, but with some new accessories slapped on an "upgrade".
11.
▲
by
desideratum
10mo ago
Thanks for sharing this. I agree w.r.t. XLA. I've been moving to JAX after many years of using torch and XLA is kind of magic. I think torch.compile has quite a lot of catching up to do. > XLA isn't at present particularly use
12.
▲
by
desideratum
10mo ago
The Scaling ML textbook also has an excellent section on TPUs. https://jax-ml.github.io/scaling-book/tpus/
13.
▲
by
desideratum
10mo ago
Aside: this guy regularly posts on the Discord server for an open-source post-training framework I maintain, demanding repayment for bugs in nightly builds and generally abusing the maintainers.
14.
▲
Finetuning GPT-OSS with Axolotl
(github.com)
3 points
by
desideratum
1y ago
|
0 comments
15.
▲
Accelerate ND-Parallel: A Guide to Efficient Multi-GPU Training
(huggingface.co)
3 points
by
desideratum
1y ago
|
0 comments
16.
▲
Training LLMs with GRPO and Interpreter Feedback Using WebAssembly
(huggingface.co)
3 points
by
desideratum
1y ago
|
0 comments
17.
▲
Training Large Language Models with Interpreter Feedback Using WebAssembly
(huggingface.co)
1 points
by
desideratum
1y ago
|
0 comments
18.
▲
DeepSeek-V3-0324
(huggingface.co)
5 points
by
desideratum
2y ago
|
1 comments
19.
▲
Training Process Reward Models in Axolotl
(axolotlai.substack.com)
2 points
by
desideratum
2y ago
|
0 comments
20.
▲
by
desideratum
2y ago
This is an exceptional salary for the UK.
21.
▲
by
desideratum
2y ago
I'd reccomend checking out the CUDA mode Discord server! They also have a channel for Metal https://discord.gg/ZqckTYcv
22.
▲
Torchtune – a native PyTorch library for fine-tuning LLMs
(github.com)
2 points
by
desideratum
2y ago
|
0 comments
23.
▲
by
desideratum
2y ago
torchtune ( https://github.com/pytorch/torchtune ) - a PyTorch library for fine-tuning LLMs, particularly for memory-constrained setups. Try it out and fine-tune Llama3.1 8B on a single RTX 4090!
24.
▲
(Deep Learning Based) Opportunistic Screening to Improve Statin Rates
(ahajournals.org)
1 points
by
desideratum
2y ago
|
0 comments
25.
▲
The theory of Proximal Policy Optimisation implementations
(salmanmohammadi.github.io)
1 points
by
desideratum
2y ago
|
0 comments
26.
▲
by
desideratum
6y ago
Hi Stefano. I'm an ML Engineer/Researcher looking for a new role, and very interested in learning more about Epistemic. Am I correctly visualizing something similar to https://www.connectedpapers.com/ as part of y
27.
▲
by
desideratum
6y ago
Nick Bostrom's "Superintelligence" is a sober perspective on this issue and a very worthwhile read.
28.
▲
by
desideratum
6y ago
Hi David, wonder if you'd be open to remote within the UK? I'd be interested in the ML Engineer role primarily, but would also happily be considered for other roles.
29.
▲
by
desideratum
6y ago
Some truly impressive results. I'll pick my usual point here when a fancy new (generative) model comes out, and I'm sure some of the other commenters have alluded to this. The examples shown are likely from a set of well-defined (
30.
▲
by
desideratum
6y ago
Thanks for this. I had the same thought about this being a lesson they'll need to learn. The other two engineers put up very little resistance and seem to be perfectly happy with their TC.
More ›