Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
MasterScrat
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
MasterScrat
2y ago
His intro to RL (not for LLM) blog post is a great read FYI https://karpathy.github.io/2016/05/31/rl/
62.
▲
by
MasterScrat
2y ago
I'd be curious to hear how you'd express these notions in more scientific terms?
63.
▲
by
MasterScrat
2y ago
I think both of these projects miss what makes TikTok (and reels in general) so effective. A good TikTok video gets "injected into your brain". You have zero effort to provide and suddenly this stuff is in your mind. I'm not
64.
▲
by
MasterScrat
2y ago
You mean, like the fastapi documentation used to be? https://github.com/fastapi/fastapi/issues/3273
65.
▲
by
MasterScrat
2y ago
How do bandwidth costs work now? do you pay the ISPs a flat fee, or is it still usage-based? how much cheaper is it compared to cloud providers?
66.
▲
by
MasterScrat
2y ago
Can you talk about the main non-LLM NLP tools you use? e.g. BERT models? > One prompt? Fair. 10? Still ok. 100? You're pushing it. 10M - get help. Assuming you could do 10M+ LLM calls for this task at trivial cost and time, would
67.
▲
by
MasterScrat
2y ago
This hasn't been my experience, using TikTok from Switzerland, I almost exclusively see English language, with a focus on my interests
68.
▲
by
MasterScrat
2y ago
Thanks, amazing story! I found this nice coverage of the events: https://www.mentalfloss.com/article/92007/why-us-federal-cou...
69.
▲
by
MasterScrat
2y ago
Multi-touch support (at least on iOS) is a nice touch
70.
▲
by
MasterScrat
2y ago
Very insightful to have a number from him here: > LLMs are trained on much more than the whole Internet -- they also consume handcrafted answers produced by armies of highly qualified data annotators (often domain experts). Today approxi
71.
▲
by
MasterScrat
2y ago
Here's how it looks like apparently: https://youtu.be/hXUbTKd6pmo?t=29
72.
▲
by
MasterScrat
2y ago
Interesting point. In a way this is a "simulated game engine", trained from actual game engine data. But I would argue a working simulated game engine becomes a game engine of its own, as it is then able to "propell the game&
73.
▲
The Problem with Startup "Experts" [video]
(youtube.com)
1 points
by
MasterScrat
2y ago
|
0 comments
74.
▲
by
MasterScrat
2y ago
We sell text-to-image model finetuning (aka "Dreambooth") as a service and yes, this is one of the use cases. Recently a travel agency used our platform to generate images of people in the destinations they were advertising.
75.
▲
by
MasterScrat
2y ago
Ok I now understand better what happened: The price for using images as part of your prompt has indeed not changed between GPT-4o-mini and GPT-4o Yet overall, captioning 500 images now costs me 5x less. This is because when I'm caption
76.
▲
by
MasterScrat
2y ago
It almost sounds shady... "it's 30x cheaper per token but you now need 30x more tokens per image"? Has anyone already validated this based on billed cost? running a batch myself to check EDIT: Ok so I captioned 500 images in
77.
▲
by
MasterScrat
2y ago
It also provides decent spam prevention, and high probability people will check the associated email address
78.
▲
by
MasterScrat
2y ago
It's a tangent, but reading this linked article from 2014: https://blog.cloudflare.com/the-relative-cost-of-bandwidth-a... Am I reading correctly that egress in Europe costs $8 Mbps/month which is $0.0004/GB,
79.
▲
by
MasterScrat
2y ago
Mostly from the tight JAX-TPU integration yeah
80.
▲
by
MasterScrat
2y ago
We've built our startup from scratch on JAX, selling text-to-image model finetuning, and it's given us a consistent edge not only in terms of pure performance but also in terms of "dollars per unit of work"
81.
▲
by
MasterScrat
2y ago
I’m not saying you’re a GPT4-based bot, but I want to point out this comment really triggered my LLM radar
82.
▲
by
MasterScrat
3y ago
I used Windows 10 for the first time in years this weekend to play games While playing , Windows would switch from the game to a stupid "Restart your computer now or in 1 hour?" modal - with no option to get it to permanently go
83.
▲
by
MasterScrat
3y ago
We started on V3s, now fully moved to V4s with some V5Es, investigating a full move towards V5E & V5P
84.
▲
by
MasterScrat
3y ago
Interesting! this was already the case with TPUs easily beating A100s. We sell Stable Diffusion finetuning on TPUs (dreamlook.ai), people are amazed how fast and cheap we can offer it - but there's no big secret, we just use hardware t
85.
▲
by
MasterScrat
3y ago
Scaling up compute can improve throughput, but can't easily improve latency between tokens. Generation is usually bottlenecked by the time it takes to go through the network for each token. To speed that up, you need to perform these c
86.
▲
by
MasterScrat
3y ago
And how does Mistral do "accurate long content retrieval"?
87.
▲
There's a supercomputer lodged inside this Barcelona church
(lonelyplanet.com)
1 points
by
MasterScrat
3y ago
|
0 comments
88.
▲
by
MasterScrat
3y ago
Heartbreaking to hear that a company that's really helping moving things forward is struggling
89.
▲
by
MasterScrat
3y ago
Interesting - so Azure OpenAI was not affected? Do they get system updates at the same time as the OpenAI API? Is the pricing the same?
90.
▲
by
MasterScrat
3y ago
TPUs are amazing for Stable Diffusion. We've been doing training (Dreambooth) and inference on TPUs since the beginning of the year at https://dreamlook.ai . We basically get 2.5x the training speed for Stable Diffusion 1.5
More ›