Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
spindump8930
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
spindump8930
6mo ago
Remember that models on different inference platforms might not necessarily give exactly the same results, adding another axis of non-determinism to development. Things like quantization, custom model serving silicon, batching, or other inf
32.
▲
by
spindump8930
6mo ago
Any more context on the copilot training note? More pointers would be very interesting, but we'd need to keep in mind how many different underlying models were (are?) branded as copilot. I thought at some points the "copilot"
33.
▲
by
spindump8930
6mo ago
> The researchers tested five LLMs: OpenAI’s GPT-4o (before the highly sycophantic and since-sunset GPT-5) Interesting, I always thought the sycophancy peaked with 4o and the associated personality (such as when myboyfriendisai users beg
34.
▲
by
spindump8930
6mo ago
Hopefully this money means more compute infrastructure to help Anthropic counter the efficiency changes that have created this perceived downtrend in claude quality.
35.
▲
by
spindump8930
6mo ago
Having known some folks who did recurse, I think places like this want to select for those who consider coding a type of craft or art or self-expression. You can use LLMs, but stand by what you do and have pride in construction.
36.
▲
by
spindump8930
6mo ago
Not clear that they even have any GPUs yet: > Allbirds, which will be renamed “NewBird AI,” said it executed a $50 million deal with an unnamed institutional investor to acquire “high-performance GPU assets” to begin transitioning into a
37.
▲
by
spindump8930
6mo ago
Yes, the paper itself tells a different story than the bullet points in this article.
38.
▲
by
spindump8930
6mo ago
The article seems quite editorialized, shifting between describing "large-scale AI models" and "neural network-based approaches". The underlying paper itself is more precise, comparing against LUAR, a 2021 method based o
39.
▲
by
spindump8930
6mo ago
Yes, it's far more certain that meta released this, which is less convincing on evals, as a result of the mythos previews.
40.
▲
by
spindump8930
6mo ago
Re: changes, there's been enormous turnover in AI organizations, and in theory this one was developed by a "new" org. Whether that means less or more benchmaxxing is anyone's guess.
41.
▲
by
spindump8930
6mo ago
Spending tons of money on Claude and the recent token benchmarks came WELL after Meta's huge investments in compute infrastructure for AI as well as the long history of language model development inside science divisions at the company
42.
▲
by
spindump8930
6mo ago
Only for poor quality systems. Unfortunately there are many systems that tried to make easy hype, but are the equivalent of an ML 101 classifier class project. If one measures for perplexity (how likely text is under a certain language mode
43.
▲
by
spindump8930
6mo ago
Pangram has time after time been shown as the only detector that mostly works. And that paper is pretty old now! There are recent papers from academics independently bench-marking and studying detectors e.g. https://arxiv.org
44.
▲
by
spindump8930
1y ago
Is the proxy here linkedin messaging/mail instead of direct email?
45.
▲
by
spindump8930
1y ago
It was mentioned elsewhere in the thread but this article is relevant: https://www.cbssports.com/mlb/news/guardians-reliever-emmanu... The ability to bet on short term individual events (such as a single pitch) me
46.
▲
by
spindump8930
1y ago
For many of us a better Turing test is contextual to a topic we CARE about. Lots of LLMs sound better than a randomly sampled human on a topic I don't know too much about (e.g. opinions on new movies). They're decent on engineerin
47.
▲
by
spindump8930
1y ago
Fairly certain that AI (meaning an expensive llm type model) isn't needed to detect spam a large amount of the time. Classical classification methods could work while also being more privacy friendly (e.g. running on device). > Acco
48.
▲
by
spindump8930
1y ago
I recall that there were similarly motivated lawsuits for the earlier answer boxes that used to appear (prior to direct genai injection into the SERP page). What ever happened with those? Finding it difficult to search for.
49.
▲
by
spindump8930
1y ago
That's not what this is about. "I had no problem getting deterministic LLM outputs when I experimented with this 6 months ago" looks like you're using llama-cpp in that repo. This is about vllm serving many requests at o
50.
▲
by
spindump8930
1y ago
This topic is interesting, but the repo and paper have a lot of inconsistencies that make me think this work is hiding behind lots of dense notation and language. For one, the repo states: > This implementation follows the framework from
51.
▲
by
spindump8930
1y ago
Don't forget that theoretical peak performance is (probably) half the performance listed on the nvidia datasheet because they used the "with sparsity" numbers! I've seen this bite folks who miss the * on the figure or ar
52.
▲
by
spindump8930
1y ago
It's good to have this support in APIs but grammar constrained decoding has been around for quite a while, even before the contemporary LLM era (e.g. [1] is similar in spirit). Local vs global planning is a huge issue here though - if
53.
▲
by
spindump8930
1y ago
Great article, thanks for sharing. The tension between "SQL is declarative" and "Write the query like this or it will OOM" has always made me uncomfortable. I've used closed, mature, systems with custom cost based o
54.
▲
by
spindump8930
1y ago
Data is at least meant to be AGI given e.g. the measure of a man storyline, as a non-human but generally intelligent. Less clear for HAL, though the story has HAL undergoing a human-like learning process and the later storylines show that e
55.
▲
by
spindump8930
1y ago
The title makes it sound nice but the reported results are worse than random baselines on several benchmarks, including ones to claim superiority over BERT. At a glance, Hellaswag, boolq, winogrande are all at or below random guessing. At b
56.
▲
by
spindump8930
1y ago
The many sources of stochastic/non-deterministic behavior have been mentioned in other replies but I wanted to point out this paper: https://arxiv.org/abs/2506.09501 which analyzes the issues around GPU non determ
57.
▲
Give Me FP32 or Give Me Death?
(arxiv.org)
3 points
by
spindump8930
1y ago
|
2 comments
58.
▲
by
spindump8930
1y ago
People often don't understand why LLMs can be non deterministic even with deterministic seeding, temperature, sampling. This paper shows how bad it can be with different hardware and gpu hosts.
59.
▲
by
spindump8930
1y ago
First off, this interface is very nice and a pleasure to use, congrats! Are you using word2vec for these, or embeddings from another model? I also wanted to add some flavor since it looks like many folks in this thread haven't seen som
60.
▲
by
spindump8930
1y ago
If you consider most of the dominate architectures in deeplearning type approaches, transformers are remarkably generic. If you reduce transformer like architectures to "position independent iterated self attention with intermediate tr
More ›