Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
echen
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
RL Environments and the Hierarchy of Agentic Capabilities
(surgehq.ai)
4 points
by
echen
11mo ago
|
0 comments
2.
▲
Explaining Reinforcement Learning with Human Feedback (RLHF)
(surgehq.ai)
11 points
by
echen
4y ago
|
0 comments
3.
▲
Users report Google Calendar bug creating random, fake events
(theverge.com)
2 points
by
echen
4y ago
|
0 comments
4.
▲
McDonald’s is testing its first robot restaurant with no human contact
(twistedfood.co.uk)
2 points
by
echen
4y ago
|
0 comments
5.
▲
The Chatbots Are Coming for Google
(bloomberg.com)
7 points
by
echen
4y ago
|
1 comments
6.
▲
ChatGPT Crushes Google on Coding Queries, and Matches It at General Information
(surgehq.ai)
11 points
by
echen
4y ago
|
1 comments
7.
▲
by
echen
4y ago
We're a human/AI data company (Surge AI) and work with many of the LLM companies to red team their systems. We actually just wrote up a blog post about it: https://www.surgehq.ai/blog/ai-red-teams-for-adversar
8.
▲
AI Red Teams for Adversarial Training: Making ChatGPT and LLMs More Robust
(surgehq.ai)
9 points
by
echen
4y ago
|
0 comments
9.
▲
HellaSwag: 36% of this popular large language model benchmark contains errors
(surgehq.ai)
49 points
by
echen
4y ago
|
8 comments
10.
▲
The Violence, Racism, & Sexism Uncaught by Twitter's Content Moderation Systems
(surgehq.ai)
3 points
by
echen
4y ago
|
0 comments
11.
▲
Move Over, Google: The TikTokification of Next-Gen Search
(surgehq.ai)
13 points
by
echen
4y ago
|
4 comments
12.
▲
Sci-Fi Reddit Community Bans AI-Art for Being 'Low Effort' Posting
(vice.com)
2 points
by
echen
4y ago
|
0 comments
13.
▲
Stability AI, the startup behind Stable Diffusion, raises $101M
(techcrunch.com)
13 points
by
echen
4y ago
|
2 comments
14.
▲
DALL·E vs. Imagen, and Evaluating Astral Codex Ten's Bet on AI Progress
(surgehq.ai)
13 points
by
echen
4y ago
|
0 comments
15.
▲
The $250K Inverse Scaling Prize and Human-AI Alignment
(surgehq.ai)
11 points
by
echen
4y ago
|
0 comments
16.
▲
The company that makes GIFs says they're 'cringe' and 'out of fashion'
(businessinsider.com)
11 points
by
echen
4y ago
|
0 comments
17.
▲
TikTok-addicted students delete app during exams
(bbc.com)
1 points
by
echen
4y ago
|
0 comments
18.
▲
Google bars Truth Social download over violent content
(cbsnews.com)
3 points
by
echen
4y ago
|
0 comments
19.
▲
Evaluation of TikTok vs. Instagram Reels
(surgehq.ai)
222 points
by
echen
4y ago
|
263 comments
20.
▲
A bartending robot that can engage in personalized interactions with humans
(techxplore.com)
1 points
by
echen
4y ago
|
1 comments
21.
▲
Scientists revived the cells of pigs an hour after death
(livescience.com)
1 points
by
echen
4y ago
|
0 comments
22.
▲
Optimizing Facebook's Algorithms for Human Values Instead of Clicks
(surgehq.ai)
7 points
by
echen
4y ago
|
1 comments
23.
▲
Egregious Failures in Gmail Spam Detection
(surgehq.ai)
9 points
by
echen
4y ago
|
0 comments
24.
▲
by
echen
4y ago
I've been experiencing the same. Here are a bunch of egregious spam mistakes we collected from different people to illustrate the problem: https://www.surgehq.ai/blog/are-the-spammers-winning-failure...
25.
▲
Smarter, Better, Faster: Using Machine Learning to Review Emotes
(blog.twitch.tv)
1 points
by
echen
4y ago
|
0 comments
26.
▲
How Good is Hugging Face's BLOOM? Human Evaluation of Large Language Models
(surgehq.ai)
10 points
by
echen
4y ago
|
0 comments
27.
▲
by
echen
4y ago
Agree that context is often needed! (Which is why it was strange to us that raters weren't presented with any context besides the comment text itself -- not even the subreddit, much less the original Reddit post.) One interesting quest
28.
▲
by
echen
4y ago
Gaming is a really fun and interesting labeling domain, given the community jargon (I'm actually a big Twitch user, but still couldn't tell you what many common emotes mean... took me years to understand "poggers") and c
29.
▲
by
echen
4y ago
Great question! I'd love to measure that more rigorously too. Although from what we've seen, the amount context sensitivity matters really depends on the labeling task / application. For example, when you're trying to la
30.
▲
by
echen
4y ago
I'd love to chat. Want to reach out to the email in my profile? I'm the founder of a startup solving this exact problem ( https://www.surgehq.ai ), and previously built the human computation platforms at a couple FAANGs
More ›