Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thorum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
151.
▲
by
thorum
2y ago
I agree in principle, but most people do not have the necessary education to make accurate judgments across all fields from first principle reasoning. Eg someone with a PhD in physics probably doesn’t have the knowledge to debunk every argu
152.
▲
by
thorum
2y ago
If public health experts are caught being deceptive in one case, or the media publishes as truth something that ends up being political propaganda- even once - they lose trust and people will correctly choose to treat future information fro
153.
▲
by
thorum
2y ago
As someone who cares deeply about science education, here are some hard truths that many people flinch away from: “Anti-science” Americans are not less intelligent on average than “pro-science” Americans. It’s simply a difference of who the
154.
▲
by
thorum
2y ago
> The historian Yuval Noah Harari has argued that humans should figure out how to work together and establish trust before developing advanced AI. In theory, I agree. If I had a magic button that could slow this whole thing down for 30 o
155.
▲
by
thorum
2y ago
Law enforcement is treated with more skepticism, and judges are more worried about being held to account for their decisions. It’s hard to see either of these developments as anything but positive. The rise of amateur sleuthing is less clea
156.
▲
by
thorum
2y ago
Fair enough.
157.
▲
by
thorum
2y ago
This isn’t a counterargument. Yes, mainstream US sites are not saints and have many serious issues. Also yes, they are least subject to US jurisdiction and can be held to some account. If these restrictions drive users to websites that are
158.
▲
by
thorum
2y ago
You’re right, and this is unfortunately becoming very common on GitHub, even for otherwise great projects. Good technical writing is concise and tells you what you need to know as directly as possible. ChatGPT written documentation always r
159.
▲
by
thorum
2y ago
The goal is for the LLMs we have to be so good at long context tasks and reasoning about code that you can simply paste the entire documentation for your new language and they’ll figure it out. Once you have that working, you can generate s
160.
▲
Towards Benchmarking LLM Diversity and Creativity
(gwern.net)
1 points
by
thorum
2y ago
|
0 comments
161.
▲
by
thorum
2y ago
Most of the comments here are responding to the title by discussing whether current AI represents intelligence at all, but worth noting that the author’s concerns all apply to human brains too. He even hints at this when he dismisses “human
162.
▲
by
thorum
2y ago
Reasoning models are supposed to be able to work around that kind of limitation (o1 was code named strawberry after all) so it’s not a bad test IMO.
163.
▲
by
thorum
2y ago
The solution is user interfaces that are stable, but infinitely customizable by the user for their personal needs and preferences, rather than being fixed until a developer updates it.
164.
▲
by
thorum
2y ago
Reminds me of the article “Language Models Model Us”: > “On a dataset of human-written essays, we find that gpt-3.5-turbo can accurately infer demographic information about the authors from just the essay text, and suspect it's infe
165.
▲
by
thorum
2y ago
Thanks! The archived snapshot doesn’t have that, it must have been added later.
166.
▲
by
thorum
2y ago
Am I missing something or does this NYT article not actually say what the judge based his decision on? That would be an important piece of information to have when deciding whether the decision was good or not.
167.
▲
by
thorum
2y ago
Thanks for this, I was wondering why people would care so much about these files that they go and complain. Reading through these forum questions, looks like most people are trying to free up disk space by clearing temp files, but the SQLit
168.
▲
by
thorum
2y ago
Clever way to get more training data.
169.
▲
by
thorum
2y ago
> I’m pretty sure it’s a cool feature but, what a mouthful. Imagine you decided to start developing websites today, how do you even start? Easy: By not using that optional feature. If you’re like 99% of developers you really don’t need s
170.
▲
by
thorum
2y ago
This is a 2 1/2 hour video essay from Angela Collier, a YouTube physics educator, talking about the difference between the 'cult of personality' perception of Feynman vs the more complicated reality of who he actually was (th
171.
▲
The sham legacy of Richard Feynman [video]
(youtube.com)
8 points
by
thorum
2y ago
|
1 comments
172.
▲
Chinese lab DeepSeek releases a 'reasoning' AI model that rivals OpenAI's o1
(techcrunch.com)
8 points
by
thorum
2y ago
|
1 comments
173.
▲
by
thorum
2y ago
It’s pretty amazing that we’re almost at the point where a single laptop computer with no internet connection can contain (1) most of humanity’s recorded knowledge and (2) an intelligence that can explain it to you.
174.
▲
I used vision models to help me win at Age Of Empires 2
(old.reddit.com)
4 points
by
thorum
2y ago
|
0 comments
175.
▲
by
thorum
2y ago
> We don't have a clue how to make a computer capable of System 1 thinking. I think you’re overthinking this. System 1 thinking as the term is being used by AI researchers means making a fast decision based on reasoning processes
176.
▲
Claude Artifact Runner
(github.com)
1 points
by
thorum
2y ago
|
0 comments
177.
▲
by
thorum
2y ago
There seems to be an ongoing mass exodus of their best talent to Anthropic and other startups. Whatever their moat is, that has to catch up with them at some point.
178.
▲
The deconstructed Standard Model equation
(symmetrymagazine.org)
43 points
by
thorum
2y ago
|
28 comments
179.
▲
by
thorum
2y ago
CoT moderately improves model performance, but all non-o1 models suck at actually thinking step by step effectively. If the task is not straightforward, they make obvious mistakes or default to guessing. OpenAI trained o1 to pick better ste
180.
▲
by
thorum
2y ago
> This along with a few tokenizer related tests people ran, made people suspect that we are just serving Claude with post-processing where we filter out words like Claude. Didn't these "few tokenizer related tests" prove t
More ›