Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
paradite
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
New Claude Models Default to Full Code Output, Stronger Prompt Required
(eval.16x.engineer)
1 points
by
paradite
1y ago
|
0 comments
62.
▲
by
paradite
1y ago
The results are not surprising, but it's good to have these findings formalized as publications, so that we (or LLMs) can refer to them as ground truth in the future.
63.
▲
by
paradite
1y ago
I hate things that are translucent. I find them very distracting, and hurt my eyes. I hope Apple gives the option to turn this whole thing off. I notice the borders now also have shadows / gradients due to reflection, that's also
64.
▲
by
paradite
1y ago
It can open PR via GitHub actions integration. I just did: https://x.com/paradite_/status/1931644656762429503 Docs: https://docs.anthropic.com/en/docs/claude-code/github-action...
65.
▲
by
paradite
1y ago
Okay maybe I need to make myself more clear, and start from your claim: > It has to do with mental capacities to "bring objects under a concept", partition experience into its conceptual structure ("conceptualise"), s
66.
▲
by
paradite
1y ago
You have biased view on the definition of "concept" based on English language and logic. In Chinese language, concept is 概念. In Chinese language, happy dog is 快乐的狗. Notice it has an extra "的" that is missing in English l
67.
▲
by
paradite
1y ago
You seem to believe "concept" is a concept in English language. It is not. Concept is an abstraction layer above human languages. Here's a good article that touched on this topic: https://www.neelnanda.io/mech
68.
▲
by
paradite
1y ago
> LLMs as they currently exist cannot master a theory, design, or mental construct because they don't remember beyond their context window. Only humans can can gain and retain program theory. False. > An LLM is a token predictor.
69.
▲
by
paradite
1y ago
Opus 4 beats all other models in my personal eval set for coding and writing. Sonnet 4 also beats most models. A great day for progress. https://x.com/paradite_/status/1925638145195876511
70.
▲
by
paradite
1y ago
I don't think it's a black and white distinction between agentic and non-agentic tools. Not to mention tools are constantly evolving and changing. For example, Cursor a year ago was not agentic at all. GitHub Copilot only recently
71.
▲
by
paradite
1y ago
Missing OpenAI Codex cli Also missing a class of non-IDE desktop apps like 16x Prompt and Repo Prompt.
72.
▲
by
paradite
1y ago
There are actually a lot of tools that do this: https://prompt.16x.engineer/cli-tools
73.
▲
by
paradite
1y ago
I built a tool that helps with copy pasting code into chat UI called 16x Prompt. https://prompt.16x.engineer/ It is used by quite a lot of people. So the problem is definitely there. The app supports API integration as well
74.
▲
by
paradite
1y ago
I'm getting so much sci-fi vibes from this post. I've read so many sci-fi stories where big tech corporations have similar control over people as countries. Now we are actually heading there. I'm both excited and a bit worrie
75.
▲
by
paradite
1y ago
It's kind of interesting if you view this as part of RLHF: By processing the system prompt in the model and collecting model responses as well as user signals, Anthropic can then use the collected data to perform RLHF to actually "
76.
▲
by
paradite
1y ago
I wonder what if this is just a decoy to get the more sophisticated candidate in.
77.
▲
by
paradite
1y ago
I can't belive Ollama haven't fix the context window limits yet. I wrote a step-by-step guide on how to setup Ollama with larger context length a while ago: https://prompt.16x.engineer/guide/ollama TLDR o
78.
▲
by
paradite
1y ago
Thinking takes way too long for it to be useful in practice. It takes 5 minutes to generate first non-thinking token in my testing for a slightly complex task via Parasail and Deepinfra on OpenRouter. https://x.com/paradite_
79.
▲
by
paradite
1y ago
https://eval.16x.engineer/ - 16x Eval: A desktop GUI app to evaluate prompts and models With 16x Eval, you can manage your prompts, contexts, and models in one place, locally on your machine, and test out different combinat
80.
▲
by
paradite
1y ago
Seems like the story of Stackoverflow.
81.
▲
by
paradite
1y ago
If you want to evaluate your personal prompts against different models quickly on your local machine, check out the simple desktop app I built for this purpose: https://eval.16x.engineer/
82.
▲
by
paradite
1y ago
I think you should try more tools and use cases. Yes some of the current AI coding tools will fail at some use cases and tasks, but some others tools might give you good results. For example, Devin is pretty bad at some trivial frontend tas
83.
▲
by
paradite
1y ago
All I care about as a certbot user is what do I need to do. Do I need to update certbot in all my servers? Or would they continue to work without the need to update?
84.
▲
How We Got Here – AI Timeline from 2015 to 2024
(thegroundtruth.substack.com)
2 points
by
paradite
1y ago
|
0 comments
85.
▲
by
paradite
1y ago
Ah I see what you mean. I was trying to convey that this is a limitation, hence not a tick symbol. But I guess it could be interpreted differently like you said.
86.
▲
by
paradite
1y ago
Have you tried using a tool like 16x Prompt to send only relevant code to the model? This helps the model to focus on a subset of codebase thst is relevant to the current task. https://prompt.16x.engineer/ (I built it)
87.
▲
by
paradite
1y ago
I forgot to mention that OpenAI also invented PPO, which is the default algorithm that everyone uses for RL since 2017: https://en.wikipedia.org/wiki/Proximal_policy_optimization DeepSeek's GRPO is also just a min
88.
▲
by
paradite
1y ago
You can bypass this problem by embedding relevant source code files directly in the prompt itself. I built a desktop GUI tool called 16x Prompt that help you do it: https://prompt.16x.engineer/
89.
▲
by
paradite
1y ago
The author mentioned AlphaGo and Alpha Zero without mentioning OpenAI gym and OpenAI Five. Those products show OpenAI was innovating and leading in RL at that stage around 2017 to 2019. https://github.com/openai/gym h
90.
▲
by
paradite
1y ago
Here's a longer blog post I wrote on the same topic, with new updates daily: https://prompt.16x.engineer/blog/quasar-alpha-openai-stealth...
More ›