Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Cynddl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
Cynddl
6mo ago
> "Unavailable Due to the UK Online Safety Act" Anyone outside the UK can share what this is about?
32.
▲
by
Cynddl
6mo ago
> “These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure.” As a researcher in the same field
33.
▲
ChatGPT Edu feature reveals researchers' project metadata across universities
(fastcompany.com)
2 points
by
Cynddl
7mo ago
|
0 comments
34.
▲
AI no better than other methods for patients seeking medical advice, study shows
(reuters.com)
3 points
by
Cynddl
8mo ago
|
0 comments
35.
▲
by
Cynddl
8mo ago
Link to the study: https://www.nature.com/articles/s41591-025-04074-y Co-author here and happy to answer questions!
36.
▲
AI chatbots pose 'dangerous' risk when giving medical advice, study suggests
(bbc.co.uk)
4 points
by
Cynddl
8mo ago
|
2 comments
37.
▲
by
Cynddl
8mo ago
Have you tried https://huetone.ardov.me/ ? Multiple color scales, P3, export to CSS and figma, as well as APCA & WCAG for accessibility.
38.
▲
Show HN: Small, anonymous app for teams to do retrospective sessions
(retrospective.rocher.lc)
1 points
by
Cynddl
8mo ago
|
0 comments
39.
▲
by
Cynddl
11mo ago
Looks like a new model trained to be warmer and friendlier to users. Time to reshare our work: https://arxiv.org/html/2507.21919 > Artificial intelligence (AI) developers are increasingly building language models wi
40.
▲
Measuring What Matters: Construct Validity in Large Language Model Benchmarks
(arxiv.org)
1 points
by
Cynddl
11mo ago
|
0 comments
41.
▲
AI Capabilities May Be Overhyped on Bogus Benchmarks, Study Finds
(gizmodo.com)
43 points
by
Cynddl
11mo ago
|
17 comments
42.
▲
AI's capabilities may be exaggerated by flawed tests, according to new study
(nbcnews.com)
3 points
by
Cynddl
11mo ago
|
0 comments
43.
▲
Experts find flaws in tests that check AI safety and effectiveness
(theguardian.com)
3 points
by
Cynddl
11mo ago
|
0 comments
44.
▲
Measuring What Matters: Construct Validity in Large Language Model Benchmarks
(oxrml.com)
3 points
by
Cynddl
11mo ago
|
2 comments
45.
▲
by
Cynddl
1y ago
Anyone knows what Mojo is doing that Julia cannot do? I appreciate that Julia is currently limited by its ecosystem (although it does interface nicely with Python), but I don't see how Mojo is any better then.
46.
▲
The quiet software tooling Renaissance
(pdx.su)
3 points
by
Cynddl
1y ago
|
0 comments
47.
▲
Facial recognition works better in the lab than on the street, researchers show
(theregister.com)
4 points
by
Cynddl
1y ago
|
1 comments
48.
▲
We Shouldn't Trust Facial Recognition's Glowing Test Scores
(techpolicy.press)
2 points
by
Cynddl
1y ago
|
0 comments
49.
▲
by
Cynddl
1y ago
Hi, author here, this is exactly what we tested in our article: > Third, we show that fine-tuning for warmth specifically, rather than fine-tuning in general, is the key source of reliability drops. We fine-tuned a subset of two models (
50.
▲
by
Cynddl
1y ago
Hi, author here! We used a dataset of conversations between a human and a warm AI chatbot. We then fed all these snippets of conversations to a series of LLMs, using a technique called fine-tuning that trains each LLM a second time to maxim
51.
▲
Training language models to be warm and empathetic makes them less reliable
(arxiv.org)
358 points
by
Cynddl
1y ago
|
375 comments
52.
▲
by
Cynddl
1y ago
That's not how GDPR works and in this case the data is clearly anonymised despite the authors' claims. Amongst others, there needs to be mechanisms for users to delete their data, whether it was at some point public or not.
53.
▲
AI's limited understanding of gender puts health equity at risk
(oii.ox.ac.uk)
4 points
by
Cynddl
1y ago
|
0 comments
54.
▲
by
Cynddl
1y ago
I see on the landing page a screenshot with "Test for GDPR PII compliance", suggesting that this tool is probably not ready for any serious usage. Anyone in the regulation landscape would know that GDPR is a EU data protection law
55.
▲
by
Cynddl
2y ago
Previous discussions: https://news.ycombinator.com/item?id=40599419 (9 months ago) https://news.ycombinator.com/item?id=38412162 (2 years ago) https://news.ycombinator.com/item?id=37186039
56.
▲
Establishing meaningful data access for algorithm audits
(syntheticsociety.oii.ox.ac.uk)
1 points
by
Cynddl
2y ago
|
0 comments
57.
▲
by
Cynddl
2y ago
I'm going to repeat myself as I do everytime I encounter such tools. These tools DO NOT provide anonymization, and especially not at the level required by the EU's GDPR (where the notion of PII does not exist). As a computer scien
58.
▲
by
Cynddl
2y ago
Thanks for sharing it! I'm the author of this research article, happy to answer any question about our work. :)
59.
▲
Alpha Lyrae: This font 'randomly' pixelates characters in a block of text
(vegaprotocol.github.io)
1 points
by
Cynddl
2y ago
|
0 comments
60.
▲
by
Cynddl
2y ago
Just today, every French newspaper and hundreds around the world. Two examples: https://www.thetimes.com/world/europe/article/pavel-durov-te... “Chief executive of the encrypted messaging app reportedly det
More ›