Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dimitri-vs
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
dimitri-vs
1y ago
Yes, second this, I would go as far to say I use it more than o3.
92.
▲
by
dimitri-vs
1y ago
Maybe the individual tokens, but from experience of using LLMs something upstream encouraged the model to think it was okay to take the action of deleting the DB, something that would override safety RL, Replit system prompts and supposed u
93.
▲
by
dimitri-vs
1y ago
> I panicked and ran database commands without permission The AI responses are very suspicious. LLMs are extremely eager to please and I'm sure Replit system prompts them to err on the side of caution. I can't see what sequenc
94.
▲
by
dimitri-vs
1y ago
More work, without a doubt - any productivity gain immediately becomes the new normal. But now with an additional "2%" error rate compounded on all the tasks you're expected to do in parallel.
95.
▲
by
dimitri-vs
1y ago
An overly eager intern with short term memory loss, sure.
96.
▲
by
dimitri-vs
1y ago
I feel like one coding benchmark should be just telling it to double check or fix something that's actually perfectly fine repeatedly and watch how bad it deep fries your code base.
97.
▲
by
dimitri-vs
1y ago
Given the tidal wave of vibecoded startups were about to see - the US may want to take a look at what the EU is doing.
98.
▲
by
dimitri-vs
1y ago
source? this would defy a lot of convention and would cause a lot of instability
99.
▲
by
dimitri-vs
1y ago
...but also 1000x harder to setup than just copy pasting into ChatGPT
100.
▲
by
dimitri-vs
1y ago
I would think you can get pretty accurate results by including the top 10 subreddits they are active in and their last 20 comments (and their score). Comments alone may not be enough, the reaction to them is more telling.
101.
▲
by
dimitri-vs
1y ago
Realistically, how many people do you think have the time, skills and hardware required to do this?
102.
▲
by
dimitri-vs
1y ago
Agreed. To add to this I often buy warehouse deals which is an order of magnitude more risky and the percentage of returns I've had to do is in the low single digits. Almost always it's an obviously brand new never opened product.
103.
▲
by
dimitri-vs
1y ago
Same. I have a very efficient workflow with Cursor Edit/Agent mode where it pretty much one-shots every change or feature I ask it to make. Working inside a CLI is painful, are people just letting Claude Code churn for 10-15 minutes an
104.
▲
by
dimitri-vs
1y ago
This might be an obvious questions but why is Claude Code not included?
105.
▲
by
dimitri-vs
1y ago
IMO it's pretty clear what vibe coding is: you don't look at the code, only the results. If you're making judgement on the code, it's not vibe coding.
106.
▲
by
dimitri-vs
1y ago
It's really just technical writing. Majority of tricks from the GTP-4 era are obsolete with reasoning models.
107.
▲
by
dimitri-vs
1y ago
It's become more subtle but still there. You can bias the model towards more "expert" responses with the right terminology. For example, a doctor asking a question will get a vastly different response than a normal person. A
108.
▲
by
dimitri-vs
1y ago
Google Workspace is always like 6 months behind on features. Even basic things like being able to delete chats: https://support.google.com/a/thread/321841996/why-can-t-i-de...
109.
▲
by
dimitri-vs
1y ago
Agreed and if you're on Windows I would highly recommend this one: https://chris.dziemborowicz.com/apps/hourglass/
110.
▲
by
dimitri-vs
1y ago
Interesting approach, I'm definitely going to steal your wording for "generate an implementation plan that...". I do something similar but entirely within Cursor: 1. create a `docs/feature_name_spec.md`, use voice-to-tex
111.
▲
by
dimitri-vs
1y ago
Cursor with gemini-2.5 MAX and agentic mode. I really like the idea of Claude Code but its rare that I fully spec out a feature on my first request and I can't see how it can be used for frontend features that require a lot of browser-
112.
▲
by
dimitri-vs
1y ago
It's actually very easy to see for yourself. When the agent "looks" at a file it will say the number of lines it looks at, almost always its the top 0-250 or 0-500 but might depend on model selected and if MAX mode is utilize
113.
▲
by
dimitri-vs
1y ago
You could say the same for code benchmarks, no? For general medical Q&A I can't see how a specialized system would be better than base o3 with web search and a good prompt. If anything RAG and guardrail prompts would degrade perfo
114.
▲
by
dimitri-vs
1y ago
You never forgot your reusable grocery bag, umbrella, or sun glasses? You've never reassembled something and found a few "extra" screws?
115.
▲
by
dimitri-vs
1y ago
Easy, any research task that will take you 5 minutes to complete it's worth firing off a Deep Research request while you work on something else in parallel. I use it a lot when documentation is vague or outdated. When Gemini/o3 ca
116.
▲
by
dimitri-vs
1y ago
This to me is a strong indicator that they are seeing the limits of current architectures and don't have a good solution to scaling to AGI.
117.
▲
by
dimitri-vs
1y ago
*Currently iOS only: > The Pendant syncs with the native phone app. We currently only have a Limitless iOS app, and the Limitless Android app will be built in [~Q3 2025].
118.
▲
by
dimitri-vs
1y ago
Not really, the point was contrasting sentimental labels with professionally defined titles, which seems precisely the distinction needed here. It's easy enough to look up on the agreed upon term for software engineer / developer
119.
▲
by
dimitri-vs
1y ago
If my 5 yo daughter draws a square with a triangle on top is she an architect?
120.
▲
by
dimitri-vs
1y ago
Have you tried building agents? They will go from PhD level smart to making mistakes a middle schooler would find obvious, even on models like gemini-2.5 and o1-pro. It's almost like building a sandcastle where once you get a prompt wo
More ›