Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ankit219
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
13 ms
·
61.
▲
Using coding LLM agents to hack Catan's browser game
(ankitmaloo.com)
1 points
by
ankit219
11mo ago
|
0 comments
62.
▲
Using GLM-4.6 to reverse engineer Catan Universe browser game (WebGL and unity)
(ankitmaloo.com)
2 points
by
ankit219
11mo ago
|
0 comments
63.
▲
by
ankit219
1y ago
I agree it's a workaround. Ideally the model should follow instructions directly, or check before running another server to see if it's starting. Though training cannot cover every usecase and different devs work differently, so i
64.
▲
by
ankit219
1y ago
create a instruction.md file with yaml like structure on top. put all the instructions you are giving repeatedly there. (eg: "a dev server is always running, just test your thing", "use uv", "never install anything
65.
▲
by
ankit219
1y ago
The headline is provocative. The issue is simple: Meta has a recommended feed and another feed which is reverse chronological. The option is hidden and not default, and all the NL judge is asking is for that option to be preserved and not b
66.
▲
by
ankit219
1y ago
Location: San Francisco Remote: okay Willing to relocate: Not outside bay area Technologies: More of a generalist. can code. Currently working on RL posttraining. In the past worked on agentic flows, and some foundational level research. Ré
67.
▲
RL environments need to account for changing priors to work
(ankitmaloo.com)
2 points
by
ankit219
1y ago
|
0 comments
68.
▲
Notes on RL Environments
(ankitmaloo.com)
1 points
by
ankit219
1y ago
|
0 comments
69.
▲
by
ankit219
1y ago
I see your point. I thought the first one was already known when deepseek came out. Perplexity team showed how they removed this kind of bias via finetuning and their finetune could answer sensitive questions. I mistakenly thought you went
70.
▲
by
ankit219
1y ago
Your claim and the original claim are vastly different. Refusing to assist is not the same as "writing less secure code". This is clearly a filter before the request goes to the model. In the article's case, the claim seems t
71.
▲
by
ankit219
1y ago
Hi! Thank you for the clarification. I was just saying it might be possible in the future (in a way you can determine how much compute - which model - a specific query needs today as well). And the experience has definitely improved w route
72.
▲
by
ankit219
1y ago
The router introduced in gpt-5 is probably the biggest signal. A router, while determining which model to route query, can determine how much $$ a query is worth. (Query here is conversation). This helps decide the amount of compute openai
73.
▲
by
ankit219
1y ago
My conjecture is that their memory implementation is not aimed at building a user profile. I don't know if they would or would not serve ads in the future, but it's hard to see how the current implementation helps them in that reg
74.
▲
by
ankit219
1y ago
The difference is implementation comes down to business goals more than anything. There is a clear directionality for ChatGPT. At some point they will monetize by ads and affiliate links. Their memory implementation is aimed at creating a u
75.
▲
by
ankit219
1y ago
there is clearly overhype. Given the influx of people building here (and i am not aware if it happened previously too), in a bid to differentiate from other startups building the same thing, many just stripped the nuance out of any technica
76.
▲
by
ankit219
1y ago
> your particular persuasive triggers through chatbot memory features, where they train and fine-tune based on your past conversations Represents a fundamental misunderstanding of how training works or can work. Memory is more to do with
77.
▲
by
ankit219
1y ago
> We use a single agent architecture (as we found this reduces hallucinations) Do you have a benchmark for this? in my experience, hallucinations have nothing to do with what framework you use.
78.
▲
by
ankit219
1y ago
I like the thesis, (and probably not a fully informed opinion here), it's easier for openai like router to go the affiliate model than ads model. Router can determine how much a query is worth (eg: help plan a vacation is worth $50 vs
79.
▲
by
ankit219
1y ago
The problem for the judge seems to be that there is no alternative at this point. No other company can bid for or credibly pay Apple/Mozilla as much as Google did. Apple testified they would spend less on innovation if the payment goes
80.
▲
by
ankit219
1y ago
In Jan, when deepseek launched, Dario Amodei had to disclose they spent about $10M to train the last generation of models (his arguments was deepseek was on the curve, not breaking it). They earned $250M in May based on ARR, and about $400M
81.
▲
by
ankit219
1y ago
Re Chrome divesture: > The remedy also extends beyond the conduct Plaintiffs seek to redress. It was Google’s control of the Chrome default, not its ownership of Chrome as a whole, that the court highlighted in its liability finding. See
82.
▲
by
ankit219
1y ago
Their projections for ARR at the end of this year at a high of $9B[1] at the end of this year. And reported gross margins of 60% (-30% with cloud providers partnerships). All things considered, if this pans out, it's a 20x multiple. Hi
83.
▲
by
ankit219
1y ago
Calling SEO content high quality is overestimating the nature and level of an SEO article. The other aspect that it implies is businesses are the reason we get so much high quality content. That is provably wrong. The moment people saw that
84.
▲
by
ankit219
1y ago
Oh nothing official. There are people who estimate the sizes based on tok/s, cost, benchmarks etc. The one that most go on is https://lifearchitect.substack.com/p/the-memo-special-editio... . This guy estimated Cla
85.
▲
by
ankit219
1y ago
A discounted Azure H100 will still be more than $2 per hour. Same goes for AWS. Trainium chips are new and not as effective (not saying they are bad) but still cost in the same range. For inference, gross margins are exactly: (what companie
86.
▲
by
ankit219
1y ago
This seems very very far off. From the latest reports, anthropic has a gross margin of 60%. It came out in their latest fundraising story. From that one The Information report, it estimated OpenAI's GM to be 50% including free users. T
87.
▲
by
ankit219
1y ago
They are still gating it by usecase (I presume). But this way, they are not limited to the creativity of what their self selected group of beta testers could come up with, and perhaps look at security against a more diverse set of usecases.
88.
▲
by
ankit219
1y ago
This is what they need for the next generation of models. The key line is: > We view browser-using AI as inevitable: so much work happens in browsers that giving Claude the ability to see what you're looking at, click buttons, and f
89.
▲
by
ankit219
1y ago
I like the idea. I also think you may need a mechanism to detect adverse actions before they are executed. This becomes important because an email cannot be unsent, and if I dont review the text, the 1 in 100 chance of the email sounding we
90.
▲
by
ankit219
1y ago
Building with non deterministic systems isnt new. It does not take a scientist. Though people who have experience with these systems are fewer in number today. You saw the same thing with TCP/IP development where we ended up developing
More ›