Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ankit219
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
91.
▲
by
ankit219
1y ago
From Incognito window's note > Others who use this device won’t see your activity, so you can browse more privately. This won't change how data is collected by websites you visit and the services they use, including Google. Dow
92.
▲
by
ankit219
1y ago
Yes, we are getting there. I think compiler is a bigger problem than unit tests given most verticals don't even have that. With unit tests, there would be some reward hacking but would be controlled at the model level + tests. (this is
93.
▲
by
ankit219
1y ago
The bottleneck for automation is verification. With human work, verification was fast(er) because you know where to look with certain assumptions that your upstream tasker would not have made trivial mistakes. For automation, AI needs to ve
94.
▲
by
ankit219
1y ago
This is super cool. Usually you dont see effective models at 270M out in the wild. The architectural choices are new and interesting as well. Would it be okay for you to divulge some more training information here? With 170M embedding param
95.
▲
Show HN: I created a random vibe coding cost estimation site via lovable
(preview--ai-lever-finance.lovable.app)
1 points
by
ankit219
1y ago
|
0 comments
96.
▲
by
ankit219
1y ago
No justification for a $100k number. For $100k a year or about $8k a month, you will end up using 1B tokens a month (that too a generous blended $8 per million input/output tokens including caching while the number is lower than that).
97.
▲
by
ankit219
1y ago
This is speculated as the reason as blockbuster acquihires have risen: https://www.bloomberg.com/opinion/articles/2025-07-17/meta-g... https://natlawreview.com/article/rise-acquihiring-po
98.
▲
by
ankit219
1y ago
Meta's case might be different, and hence i wrote it as scale would have been acquired in past times instead of an investment like they did today. For other companies, the hiring aspect avoids M&A filings, balance sheet consolidati
99.
▲
by
ankit219
1y ago
They arent public and pitch in front of sophisticated guys who can probably tell this apart.
100.
▲
by
ankit219
1y ago
This is a highly speculative post, with conjectures presented as facts. Some things that irked me: - Cursor did not hire Anthropic's "researchers". It hired the guys who built Claude code (PM and dev). Who then promptly went
101.
▲
by
ankit219
1y ago
Anthropic in the latest fundraise announce their gross margins are around 60%. Typically in the industry now, the gross margins are reported as revenue - cost of only serving paid users (free users etc. go in either cac or r&d). The thi
102.
▲
by
ankit219
1y ago
I think CLI is a good idea for now. Next abstraction seems to be Github PRs where someone (likely me) files an issue/feature, then I click a button, and the agent fixes the issue/feature. Github has talked about something similar,
103.
▲
by
ankit219
1y ago
This was a version where they wanted to collect data. The next version where gemini cli/jules harness is likely part of the RL training environment and the model would work a lot better. Thats the trajectory of improvement.
104.
▲
by
ankit219
1y ago
Likely op does not mean ai slop, but more a signal of human carelessness that they could not write it in a proper manner.
105.
▲
by
ankit219
1y ago
Interesting article, full of speculation and some logical follows, but feels like it feels short of admitting what the true conclusion is. Model building companies can build thinner wrapper / harness and can offer better prices than th
106.
▲
by
ankit219
1y ago
The article does not say anything substantial, but just some opposite viewpoints. 1/ Openai's technical staff were using Claude Code (API and not the max plans). 2/ Anthropic's spokesperson says API access for benchmarki
107.
▲
by
ankit219
1y ago
Somewhat related to the article, but mostly anecdotal. In SF, i have had chats with (~15) engineers who after some prodding admitted that they feel the whole AGI thing is passing them by (not that it is close). In a sense that they want to
108.
▲
by
ankit219
1y ago
The study where developers lost productivity due to using AI tools should always be mentioned with caveats. This comment by simonw for reference: https://news.ycombinator.com/item?id=44522772 1/ They tested on 16 exper
109.
▲
by
ankit219
1y ago
That is the model everyone gravitate towards. Openai's Fidji too started with the note about how superintelligence is for everyone. I think it would be back to income based tiers though. You want more assistance, pay $200 per month. Ev
110.
▲
by
ankit219
1y ago
API has fewer limits, and practically limitless. Claude is also on Aws and gcp, where you get more quotas (probably credits as well) and different rate limits.
111.
▲
by
ankit219
1y ago
The current paradigm is driven by two factors: one is the reliability of the models and that constraints how much autonomy you can give to an agent. Second is about chat as a medium which everyone went to because ChatGPT became a thing. I s
112.
▲
by
ankit219
1y ago
In my day job, working with biotech and life science research companies to automate FDA compliance. That is automatically generate sections of their submissions based on their results/protocols and FDA rules[1]. Here is a short demo:
113.
▲
by
ankit219
1y ago
i meant in a way that programming is always synthesis. (to my brain) combining stuff from multiple places is good, but not great synthesis you can do on your own. Many developer feel there is not much of a difference between the two. No com
114.
▲
by
ankit219
1y ago
One unpopular opinion I hold is that in recent times programming became a lot more about integrating libraries and frameworks vs writing your own thing. This was fine because if someone open sourced it, why repeat the same work. That ended
115.
▲
Elsa, FDA's AI tool has been hallucinating research and fabricating studies
(cnn.com)
5 points
by
ankit219
1y ago
|
0 comments
116.
▲
by
ankit219
1y ago
The very premise of the article is that tasks are needed for humans to learn and maintain skills. Learning should happen independently, it is a tautological argument that since human wont learn with agents which can do more, we should not h
117.
▲
by
ankit219
1y ago
This is not an NxM verifier hell. I explicitly talked about one way which is parallel generation + classifier. You can also use majority voting here. Both would give you the right answer at each step without having to write code or test cas
118.
▲
by
ankit219
1y ago
I talk about it from experience. How else do you think people are training RL agents if not based on verifiers? You don't have to verify every output at every step, you just need enough to course correct the agent and catch early when
119.
▲
by
ankit219
1y ago
They are not completely independent. It's a good assumption though. If a model encounters something out of distribution then all five of the generations will fail. If the model knows and went in a wrong direction (due to lack of reliab
120.
▲
by
ankit219
1y ago
These are all solvable problems. The issue is given the race to get to a certain ARR quickly, many startups end up not focusing on these. There is some truth to AI agents being not as useful as their promise, but the problems mentioned are
More ›