Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mirker
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
mirker
4y ago
Even so, you can run models on a private cloud. No need to do anything but sense on an embedded device.
32.
▲
by
mirker
4y ago
The alignment portion of training requires you to have upvote/downvote data on many LLM responses. Google’s attempt at that (at least according to the news so far) was asking all employees to volunteer time ranking the responses. Combi
33.
▲
by
mirker
4y ago
I’d argue that tokenization is more analogous to having a losslessly encoded image where you don’t have access to the decoder. Humans don’t perform well on that one.
34.
▲
by
mirker
4y ago
You didn’t mention how to gather high quality data. OpenAI has never and will never release that.
35.
▲
by
mirker
4y ago
“This parser is getting complicated” -> time for GPU offload to 5000 GPU nodes running GPT-5.
36.
▲
by
mirker
4y ago
Backtracking is easily solved with a shortest path algorithm. I don’t see any need for masking if you are simply maximizing likelihood of the entire sequence.
37.
▲
by
mirker
4y ago
You may be shocked to hear this but dijkstra’s short path algorithm is the technical answer to this question. We just don’t use it because it’s expensive.
38.
▲
by
mirker
4y ago
I’d say that any technical limitation that doesn’t apply to humans is not AGI. Context windows are the most apparent ones; humans don’t have a stroke after reading N characters.
39.
▲
by
mirker
4y ago
There are domain specific languages that use these primitives/abstractions e.g., Halide https://halide-lang.org/tutorials/tutorial_lesson_05_schedul...
40.
▲
by
mirker
4y ago
Without accounting for data and model architecture, it’s not a very useful number. For all we know, they may have sparse approximations which would throw this off by a lot. For example, if you measure a fully connected model over images of
41.
▲
by
mirker
4y ago
It seems more difficult to do with a target moving so fast. It’s possible costs drop by orders of magnitude every year.
42.
▲
by
mirker
4y ago
Is it failures or is this some backfill/budget scheduling while everyone is sleeping?
43.
▲
by
mirker
4y ago
The model has less technical debt from being newer. The quality per cost is supposed to be good.
44.
▲
by
mirker
4y ago
The common free to play guard is you need to play X number of unranked games before you can play ranked. The account is “paid” for with some proof of work.
45.
▲
by
mirker
4y ago
Isn’t this very similar to Karpathy’s nanoGPT?
46.
▲
by
mirker
4y ago
It’s very efficient in processing strings since it’s linear time with respect to input string length if implemented properly.
47.
▲
by
mirker
4y ago
Easiest way would be to classify the query to go to either Bing proper or ChatGPT. Example: “What is today’s date?” -> Bing “Write a rap song about hippos” -> ChatGPT
48.
▲
by
mirker
4y ago
Agreed. GitHub and OpenAI are the current branding. Though they did have some twitter bots go bad years ago and maybe they learned from that?
49.
▲
by
mirker
4y ago
My interpretation of his claim is it’s no longer a good career move to focus on big-picture research. It’s better to attack problems with narrow scope or superficial impact. Some of those papers are useful, but they’re not the most importa
50.
▲
by
mirker
4y ago
Inference is latency sensitive so the front end is still relevant.
51.
▲
by
mirker
4y ago
I mean, why not? I have alerts on cost in my cloud usage corresponding to orders of magnitude increases in expected cost. If they trigger a few hours faster (more accurately) that’s good. But I don’t see what I would do bar shutting down th
52.
▲
by
mirker
4y ago
It’s plausible to do what you’re saying. I still think it’s more trouble than it’s worth. If you somehow managed to engineer everything perfectly and tune it as you say, you’d still have customers who wanted their service to stay online pa
53.
▲
by
mirker
4y ago
I agree you can get something working. However, if 1% of the time, a customer is overcharged $100+ because they assumed the limit was guaranteed, you’d probably reconsider if you want to offer that service at all.
54.
▲
by
mirker
4y ago
See my other reply. The reason I believe it’s architected this way is to have a good tradeoff between timeliness and overhead cost. Full correctness at global scale would involve serializing all billable events through a series of global tr
55.
▲
by
mirker
4y ago
There actually are some fundamental laws related to this problem. It’s a distributed system, so you can lose availability of the metrics data and be bounded by the CAP theorem. You’d also need to keep the metrics synchronized at some granu
56.
▲
by
mirker
4y ago
The latency to query cost is on the order of hours on user side so the service would be off the limit by the cost rate * cost latency. That’s my guess for why it’s not a feature.
57.
▲
by
mirker
4y ago
The system approves requests by how trustworthy you are. It does this so: 1. You can’t denial of service the cloud or something else 2. Launder money through mining cryptocurrency 3. Take popular instances away from high profit customers
58.
▲
by
mirker
4y ago
Do you feel different? There is a theory that vitamin D is merely correlated with a healthy body. Sunlight may be required.
59.
▲
by
mirker
4y ago
Thought: If your beef is horse meat from a bad supply chain, your vegetarian burger may also be horse meat.
60.
▲
by
mirker
4y ago
Usually grad students get projects to investigate or prove a particular idea. It’s especially straightforward to do this when the advisor is iterating on a line of ideas, and you resume where the last grad student left off. There isn’t a po
More ›