Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alsima
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
alsima
2mo ago
You can also check out another demo here which compares Bullet speed to Codex/CC: https://youtu.be/rWVmG5fRKgE
32.
▲
by
alsima
2mo ago
Thanks sensho :)
33.
▲
Show HN: A faster coding agent than Codex and Claude Code
(codewithbullet.com)
9 points
by
alsima
2mo ago
|
6 comments
34.
▲
Anyone else think Claude Code and Codex are too slow? We hate it
(bullet.davidhf.com)
4 points
by
alsima
2mo ago
|
0 comments
35.
▲
by
alsima
2y ago
Potentially true as well haha
36.
▲
by
alsima
2y ago
Definitely not saying multi-agents is all you need for SWE-bench haha. I touch on this at the end of the blog post, where I mention jumps in progress require better base models or tooling.
37.
▲
by
alsima
2y ago
A lot...as you might imagine the costs of running the whole organization scale immensely.
38.
▲
by
alsima
2y ago
It's the new cerebral valley slang dude
39.
▲
by
alsima
2y ago
I see. Agree with the point about marginal improvements at a hefty increase in computational cost (I touch on this a bit at the end of the blog post where I mention that better performance requires better tooling/base models). Though I
40.
▲
by
alsima
2y ago
Looking at Table 3: "Our [sampling and voting] method outperforms other methods used standalone in most cases and always enhances other methods across various tasks and LLMs", which benchmark did the majority vote algorithm perfor
41.
▲
by
alsima
2y ago
Hmmm I get what you mean...I think it's hard to sell a solution around this idea, but I think it will become something more like a common practice/performance improvement method. James Huckle on Linkedin ( https://www.li
42.
▲
by
alsima
2y ago
honestly agree. When I first started working with agents I didnt fully understand what it really was either but I eventually fell on a definition of an LLM call that performs a unique function proactively ¯\_(ツ)_/¯.
43.
▲
by
alsima
2y ago
Well, you have Cognition AI and Devin that became a recent unicorn startup (partnerships with Microsoft and stuff) but true, I can't think of an agent that actually lives up to the hype (heard Devin wasn't great).
44.
▲
by
alsima
2y ago
I would check out this company, Swarms ( https://github.com/kyegomez/swarms ) who's working with enterprises to integrate multi-agents. But definitely a great point to focus on, the research paper mentions that the
45.
▲
AI agents but they're working in big tech
(alexsima.substack.com)
66 points
by
alsima
2y ago
|
55 comments
46.
▲
by
alsima
2y ago
If we structured AI agents like big tech org charts, which company structures would perform better? Inspired by James Huckle's thoughts on how organizational structures impact software design, I decided to put this to the test: https:
47.
▲
by
alsima
2y ago
While I was working on building SIMA, a multi-agent software-engineer (which recently achieved 27.67% on SWE-bench-lite: https://www.swebench.com/ ), I realized that evaluations on the actual benchmark were quite inconsisten
48.
▲
by
alsima
3y ago
Thank you for bringing this to out attention! We are currently working on a fix.
49.
▲
by
alsima
3y ago
Awesome to hear that it was useful for you!!!
50.
▲
by
alsima
3y ago
Most likely, the model would be less inclined to answer questions/hallucinate for prompts not related to AWS—this is definitely be a future path for improvement
51.
▲
by
alsima
3y ago
We're in the process of doing just that and adding chat context/basically remembering your past questions.