Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nsagent
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
151.
▲
by
nsagent
2y ago
> And crucially, we made sure to tell the model not to guess if it wasn’t sure. (AI models are known to hallucinate, and we wanted to guard against that.) Prompting an LLM not to confabulate won't actually prevent it from doing so.
152.
▲
by
nsagent
2y ago
As a postdoc with a background in NLP, I've come to the same conclusion. I feel like I'm stagnating in academia and have decided to quit my postdoc and self fund research that I can hopefully monetize in the future. It certainly d
153.
▲
by
nsagent
2y ago
Well, my recruiter callback ratio is likely 1 out of 5, despite having a very VERY niche profile: a PhD focused on NLP for creative text generation, especially in video games, and a prior career as a game developer. Needless to say, I'
154.
▲
by
nsagent
2y ago
> Since answers will take longer to generate, ChatGPT will display a progress bar and send an in-app notification if you switch away to another conversation. I wonder if they are optimizing for batching, which might introduce even more d
155.
▲
by
nsagent
2y ago
Please change the title to be less sensationalist. The title from the article is perfectly acceptable.
156.
▲
by
nsagent
2y ago
When the videos are completely unrelated to your search, but just happen to be new/popular videos, then yes it's useless. Surfacing relevant videos would make sense. For example, searching for climbing comp videos and getting a
157.
▲
by
nsagent
2y ago
Interesting perspective, considering a paper ByteDance just released yesterday [1] has much worse video quality. If your comparison is to real videos, then for sure the quality isn't great. If instead you compare to other released rese
158.
▲
by
nsagent
2y ago
> The moment you post to reddit, they can do anything they want with the content. It is now theirs. This is patently false. The content is not theirs, users grant Reddit a license to their content. See the TOS [1] you spoke of: > You
159.
▲
by
nsagent
2y ago
It's been tried and seemingly works well! https://arxiv.org/abs/2409.06131
160.
▲
by
nsagent
2y ago
I actually did invest in projects rather than papers and I definitely feel I paid the price. It's the main reason I opted for a postdoc rather than going straight on the academic market: the quality of my research was high, but the num
161.
▲
by
nsagent
2y ago
Both models generate an answer after multiple turns, where each turn has access to the outputs from a previous turn. Both refer to the chain of outputs as a trace. Since OpenAI did not specify what exactly is in their reasoning trace, it&#x
162.
▲
by
nsagent
2y ago
Sorry, but that does not seem to be the case. A friend of mine who runs a long context benchmark on understanding novels [1] just ran an eval and o1 seemed to improve by 2.9% over GPT-4o (the result isn't on the website yet). It's
163.
▲
by
nsagent
2y ago
This seems like an even more general version of GameNGen [1] which only looked at Doom. [1]: https://news.ycombinator.com/item?id=41375548
164.
▲
by
nsagent
2y ago
Not looking good. Apparently the model was broken when they released it yesterday. The version they uploaded 8hrs ago only has an 8k context length, so we can't test it on the novels. Here's the updates to the model config on hugg
165.
▲
by
nsagent
2y ago
If this does indeed beat all the closed source models, then I'm flabbergasted. The amount of time and resources Google, OpenAI, and Anthropic have put into improving the models to only be beaten in a couple weeks by two people (who as
166.
▲
by
nsagent
2y ago
This statement from Wyden's press release seems to be in contrast to Chris Cox's reasoning in his journal article [1] (linked in the amicus). It is now firmly established in the case law that Section 230 cannot act as a shield
167.
▲
by
nsagent
2y ago
If anything it sounds like "related" is not what they are actually doing. Rather they are looking at ways to uniquely fingerprint users through optimizing how they split "related" sites. Reminds me of the research that s
168.
▲
by
nsagent
2y ago
The current comments seem to say this is rings the death knell of social media and that this just leads to government censorship. I'm not so sure. I think the ultimate problem is that social media is not unbiased — it curates what peop
169.
▲
by
nsagent
2y ago
That's definitely around the corner! You should check out the paper [1] I presented at the Wordplay workshop [2] a few weeks ago. The paper discusses automatic evals using LLM agents, but the real fun was actually playing with the mode
170.
▲
by
nsagent
2y ago
This rings so true to me. In fifth grade after being given an IQ test, I took Algebra I early, only to retake it when I moved to junior high the next year. Then I had to retake trig because I did no homework, but aced the tests. Luckily the
171.
▲
by
nsagent
2y ago
Was just talking with a microbiology researcher yesterday about ChatGPT. They use it for helping to write code for analyzing their lab data and they were surprised I said it's likely to lead them astray at times, especially if they don
172.
▲
by
nsagent
2y ago
> the user competition is meant to be a fun side-project that we threw together today, I think it's cool that people hack things like that so quickly :) The website comes off as a marketing strategy rather than a fun one-day hackath
173.
▲
by
nsagent
2y ago
If you're looking for serialization/deserialization, you might consider confactory [1]. I created to be a factory for objects defined in configs. It actually builds the Python objects without much effort on the user. It simply mak
174.
▲
by
nsagent
2y ago
I had the same underlying reason to stay informed and also quit reading the news. It's certainly jarring to hear major events like the assassination attempt from friends and family first, but it's very reminiscent of my experience
175.
▲
by
nsagent
2y ago
Yeah, I switched from Homebrew to MacPorts a few years ago and couldn't be happier.
176.
▲
by
nsagent
2y ago
Using Occam's razor, that is less probable than the model picking up on statistical regularities in human language, especially since that's what they are trained to do.
177.
▲
by
nsagent
2y ago
You're quite right that LLMs can seemingly do some abstract reasoning problems, but I would not say they aren't in the training data. Sure, the exact form using the made up word gronk might not be in the training data, but the gen
178.
▲
by
nsagent
2y ago
Monocultures are known to be points of failure, but people keep going down that path because they optimize for efficiency (heck, most modern economics is premised on the market being efficient). This problem is pervasive and effects everyth
179.
▲
by
nsagent
2y ago
Reminds me of a recent podcast with Staša Gejo [1], a top competition climber. She basically says the same thing. At times she hated being told to do drills growing up, but really valued that later because as a kid she sometimes didn't
180.
▲
by
nsagent
2y ago
Having worked with Girls Who Code and been in many DEI discussions with faculty, industry, and current/prospective students, a lot of the disparities in supposed interest comes down to women feeling like they don't fit into the CS
More ›