Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
paradite
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
paradite
1y ago
I recently did a livestream on trying to understand attention mechanism (K, Q, V) in LLM. I think it went pretty well (was able to understand most of the logic and maths), and I touched on some of these terms. https://youtube.com
32.
▲
by
paradite
1y ago
I've came to the pain conclusion on Next.js: It's a bad framework for building full stack apps, but it's better than anything else. https://x.com/paradite_/status/1941016421934551477
33.
▲
by
paradite
1y ago
Claude Sonnet 4 is the best coding model. Period. Nothing else comes close. Anthropic probably has 80% of AI coding model market share. That's a trillion dollar market.
34.
▲
by
paradite
1y ago
There is automatic code indexing from Cursor. Autocomple is also automatically triggered when you place your cursor inside the code.
35.
▲
by
paradite
1y ago
Claude models have been weaker in vision tasks compared to models from OpenAI and Google. https://eval.16x.engineer/evals/image-analysis For them to roll out a browser extension must mean that they have found a walkaro
36.
▲
LLM Context Management: How to Improve Performance and Lower Costs
(eval.16x.engineer)
2 points
by
paradite
1y ago
|
0 comments
37.
▲
by
paradite
1y ago
It won't work because of too many false positives. People are already trained to ignore warnings, like how they blindly accept T&C without reading.
38.
▲
by
paradite
1y ago
Hey. I like your roast on benchmarks. I also publish my own evals on new models (using coding tasks that I curated myself, without tools, rated by human with rubrics). Would love you to check out and give your thoughts: Example recent one o
39.
▲
by
paradite
1y ago
It's pay wall for me.
40.
▲
by
paradite
1y ago
JSR does that? Now that might be a good reason to move my packages over to get rid of tsup.
41.
▲
GPT-5 Coding Evaluation: Underwhelming Performance Given the Hype
(eval.16x.engineer)
6 points
by
paradite
1y ago
|
0 comments
42.
▲
The Identity Crisis: Why LLMs Don't Know Who They Are
(eval.16x.engineer)
1 points
by
paradite
1y ago
|
0 comments
43.
▲
by
paradite
1y ago
Is this from Moonshot AI (company behind the Kimi K2), or a 3rd party? Judging from the design, I assume it's not officially related to the model.
44.
▲
The Pink Elephant Problem: Why "Don't Do That" Fails with LLMs
(eval.16x.engineer)
3 points
by
paradite
1y ago
|
0 comments
45.
▲
by
paradite
1y ago
The performance not only depends on the tool, it also depends on the model, and the codebase you are working on (context), and the task given (prompt). And all these factors are not independent. Some combinations work better than others. Fo
46.
▲
by
paradite
1y ago
What kind of questions / domains were you encountering false information on?
47.
▲
by
paradite
1y ago
It's actually more complex than just input and output tokens, there are more pricing rules by various providers: - Off-peak pricing by DeepSeek - Batch pricing by OpenAI and Anthropic - Context window differentiated pricing by Google a
48.
▲
by
paradite
1y ago
I believe everyone should run their own evals on their own tasks or use cases. Shameless plug, but I made a simple app for anyone to create their own evals locally: https://eval.16x.engineer/
49.
▲
by
paradite
1y ago
I am not pro or against AI-generated posts. I was just making an observation and testing my AI classifier.
50.
▲
by
paradite
1y ago
"Hard truth" and "reality check" in the same post is dead giveaway. I read and generate hundreds of posts every month. I have to read books on writing to keep myself sane and not sound like an AI.
51.
▲
by
paradite
1y ago
This is obviously AI generated, if that matters. And I have an AI workflow that generates much better posts than this.
52.
▲
by
paradite
1y ago
Agents have been a field in AI long since 1990s. MDP, Q learning, TD, RL, PPO are basically all about agent. What we have today is still very much the same field as it was.
53.
▲
by
paradite
1y ago
This is an incredibly fascinating read into how OpenAI works. Some of the details seem rather sensitive to me. I'm not sure if the essay is going to stay up for long, given how "secretive" OpenAI is claimed to be.
54.
▲
by
paradite
1y ago
I think you are confusing "I don't like it" with "It's not going to happen". Just because you don't like it, it doesn't mean it's not going to happen. Observe the world without prejudice. Think r
55.
▲
by
paradite
1y ago
As dang said, presume good faith. It's part of the HN guideline. Also, "Never attribute to malice that which is adequately explained by stupidity"
56.
▲
My Claude Code Workflow and Personal Tips
(thegroundtruth.substack.com)
1 points
by
paradite
1y ago
|
0 comments
57.
▲
Claude 4, Gemini 2.5 Pro, and GPT-4.1: Understanding Their Unique Quirks
(eval.16x.engineer)
2 points
by
paradite
1y ago
|
0 comments
58.
▲
by
paradite
1y ago
I chanced up some endpoint that mentioned "makersuite" that seems to be what's behind genai. It's interesting to see how one team / product gets rebranded to another in Google.
59.
▲
by
paradite
1y ago
"To continue, please install the Microsoft 365 Copilot app" I got this on mobile. Seems to be pretty apt.
60.
▲
AI Coding Landscape (June 2025)
(paradite.github.io)
2 points
by
paradite
1y ago
|
0 comments
More ›