Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Majromax
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
Majromax
4mo ago
From the prompt timings above, it seems like 'prompt eval time' is the equivalent to 'processing time for input tokens'. Hyperscalers can perform this evaluation very quickly because evaluation can be significantly paral
32.
▲
by
Majromax
5mo ago
> They are all just variations of "insert a canned prompt", varying only along the dimensions of (a) how and where the prompt is installed and from where it is sourced, and (b) which context or contexts the prompt runs in. Ther
33.
▲
by
Majromax
5mo ago
> You think tax incentives are what makes VC work in California but not other places in the US let alone Canada? I believe that my comment above was aligned with your premise here. I say that the tax difference is not sufficient by it
34.
▲
by
Majromax
5mo ago
> I'm sure a lot of Canadian tech workers would repatriate and foreign workers would immigrate to Canada if they could lower taxes across the board and make life easier for tech companies and workers I'm not sure that Canadian
35.
▲
by
Majromax
5mo ago
> Even if it is common (i don't think this is required any more anyways), just why? As far as Canadian law goes, there are two factors at play in the parent's events; * NAFTA work permits are applied for at the border, on entry
36.
▲
by
Majromax
5mo ago
> It’s territorial waters belonging to Iran and Oman. The trick is that it's still an 'international strait', or a segment of water that forms the only connection between two areas of high seas -- in this case the Persian
37.
▲
by
Majromax
5mo ago
> I don't know enough about the current state of naval warfare but I've assumed this is related to the asymmetry that's emerged around protecting capital warships, especially in the scenario of a very narrow strait and a l
38.
▲
by
Majromax
5mo ago
> If the person using a tool is an attorney, then that communication should be protected whether it's by pen or keyboard. But the tool is not your attorney, so it can't be the originator of attorney-client privilege. The situ
39.
▲
by
Majromax
5mo ago
> if your question is "what is the capital of france" the LLM could presumably extract out "paris" from the value vector during attention computation instead of needing the FFN for that. But how do you get 'Paris
40.
▲
by
Majromax
5mo ago
> That said, I was sympathetic to the recent bug reports —- to trigger one, you’d need to have a session that waited an hour doing nothing and then very specifically tested for in-context retrieval. I don’t want to run that test, do you
41.
▲
by
Majromax
6mo ago
> Since the devs on HN (& the whole world) is buying what looks like nonsense to me - what am I missing? Input tokens are expensive, since the whole model has to be run for each token. They're cheaper than output tokens because
42.
▲
by
Majromax
6mo ago
Context length 1e6, vector length 1e3, and 1e2 model layers for 100e9 context size. Costs will go up even more with a richer latent space and more model layers, and the western frontier outfits are reasonably likely to be maximizing both.
43.
▲
by
Majromax
6mo ago
> I believe what they're saying is they attempted to fine tune both Qwen and Pythia using Karoline Leavitt's "corpus" (I guess transcripts of press conferences) where she is presumably using the word "deportation
44.
▲
by
Majromax
6mo ago
My reading of the article is that the first audience for this test is the vendors themselves. The test is long and comprehensive to give the vendor confidence in its own hosting.
45.
▲
by
Majromax
6mo ago
> That nudge is the flinch. It is the gap between the probability a word deserves on pure fluency grounds and the probability the model actually assigns it. Hold up, what is the 'probably a word deserves on pure fluency grounds'
46.
▲
by
Majromax
6mo ago
Only for the range of tasks where 4.7 performs well but 4.6 performed suboptimally. If both models can one-shot the task without retries, then the number of iterations is already at the lower bound. This also applies at the sub-task level.
47.
▲
by
Majromax
6mo ago
> Similarly, it's unlikely you can measure a significant performance difference between models like GPT 5.4-xhigh and GPT 5.2 unless you have a task where one of them almost always fails or one almost always succeeds That feels like
48.
▲
by
Majromax
6mo ago
Decoupling capacitors don't remove the ground reference, they just allow high-frequency signals a faster path to ground. Typically, you need dedicated circuitry (and usually inductive coupling) to provide full isolation, but if the c
49.
▲
by
Majromax
6mo ago
Caveman doesn't and cannot change the tokenizer, so the relative token count differences by input category will remain unchanged.
50.
▲
by
Majromax
6mo ago
While that's a nice effort, the inter-run variability is too high to diagnose anything short of catastrophic model degradation. The typical 95% confidence interval runs from 35% to 65% pass rates, a full factor of two performance diff
51.
▲
by
Majromax
6mo ago
Don't confuse the many voices of a crowd with a single person's fickle view. If you can track an individual person or organization who changes their mind 'every other week' then more power to you, but unless you're
52.
▲
by
Majromax
6mo ago
Are Macs/etc compute bound with their 'it fits in unified memory' language models? Certainly by the time you're streaming weights from SSD you must be back in a bandwidth-bound regime.
53.
▲
by
Majromax
6mo ago
You meant “reasonable,” but you did not apply reason. Situations such as this can be handled with a quota set at something like 150% of median use, but then extended upon a justified request. It can work in a lab where there’s a human touch
54.
▲
by
Majromax
6mo ago
That's just the binary nature of the bet. You address that in a real trading strategy with (fractional) Kelly position sizes. Anyone doing this for actual money would also be well served by implementing continuous monitoring and acti
55.
▲
by
Majromax
6mo ago
> picking up pennies in front of a steamroller. This or any other statistical play is only 'in front of a steamroller' if you do it with leverage, especially if notionally uncorollated bets suddenly move together. Bets on Poly
56.
▲
by
Majromax
6mo ago
> Bookmakers price purely based on facts and statistics. Their pricing isn't affected by excitement nor by how many people are betting a certain way. A bookmaker is a market maker, and they ideally want to end up with no net intere
57.
▲
by
Majromax
6mo ago
> Edit: conversely, if the average no costs _more_ than 73 cents, but the 73% of all polymarkets resolve to No, that would imply that an everything-always-happens strategy is profitable (neglecting slippage) Or just the bid-ask spread; p
58.
▲
by
Majromax
6mo ago
'Enough' [capital] is doing a lot of work in that sentence. In the limit of a one-sided irrational market, the 'rational actor' would need to take the other side of every open transaction.
59.
▲
by
Majromax
6mo ago
Wet streets cause rain? You don’t show that the Polymarket signal lead Nasdaq.
60.
▲
by
Majromax
6mo ago
> Not a lawyer, but as I understand it the license is a matter of copyright, and the copyright only applies to the design files. So as long as you're making that keyboard for yourself then you should be good to do anything you want
More ›