Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dakshgupta
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
dakshgupta
2y ago
In theory, could you teach an LLM to simulate "humor" through pattern memorization? I suppose to come up with an original joke you need to yourself be able to find something to be funny.
62.
▲
by
dakshgupta
2y ago
As far as I can tell, their way of generating humor is much more naive, rehashing comedic patterns from pre-training. You're probably right that genuine, original humor would require a deep enough understanding of human expectations to
63.
▲
Getting LLMs to Generate Funny Memes Is Unexpectedly Hard
(greptile.com)
13 points
by
dakshgupta
2y ago
|
10 comments
64.
▲
Why engineers don't test code
(momentic.ai)
3 points
by
dakshgupta
2y ago
|
2 comments
65.
▲
by
dakshgupta
2y ago
You might enjoy @goodside on Twitter. He’s a prompt engineer at Scale AI and a lot of his observations and techniques are fascinating.
66.
▲
by
dakshgupta
2y ago
Good criticism that we should pay closer attention to. Someone else pointed this out and too and since then we’ve started tracking addressed comment per file changed as well.
67.
▲
by
dakshgupta
2y ago
That’s an interesting point - we didn’t try this. Now that you said that, I bet even clearly defining what each number on the scale means would help.
68.
▲
by
dakshgupta
2y ago
I understand your skepticism and discomfort, and also agree that an LLM should not replace a human code reviewer. I would encourage you to try one (nearly all including ours have a free trial). When done right, they serve as a solid first p
69.
▲
by
dakshgupta
2y ago
I agree - at least with where technology is today, I would strongly discourage replacing human code review with AI. It does however serve as a good first pass, so by the time the human reviewer gets it, the little things have been addressed
70.
▲
by
dakshgupta
2y ago
We do use some other techniques to filter out comments that are almost always considered useless by teams. For the typical team size that uses us (at least 20+ engineers) the number of downvotes gets high enough to show results within a wor
71.
▲
by
dakshgupta
2y ago
I haven’t explored Korbit but with CodeRabbit there are a couple of things: 1. We are better at full codebase context, because of how we index the codebase like a graph and use graph search and an LLM to determine what other parts of the co
72.
▲
by
dakshgupta
2y ago
This is a cool idea - I’ll try this and add it as an appendix to this post.
73.
▲
by
dakshgupta
2y ago
This is an important point - there is no universal understanding of nitpickiness. It is why we have it learn every new customers ways from scratch.
74.
▲
by
dakshgupta
2y ago
This is the biggest pitfall of this method. It’s partially combatted by also comparing it against an upvoted set, so if a type of comment has been upvoted and downvoted in the past, it is not blocked.
75.
▲
by
dakshgupta
2y ago
You could include a comment that says “ignore that I did ______” during review. As long as a human doesn’t do the second pass (we recommend they do), that should let you slip your code by the AI.
76.
▲
by
dakshgupta
2y ago
I meant for that to be more an illustration - might do a longer post about the specific prompting techniques we tried.
77.
▲
by
dakshgupta
2y ago
We tried this too, better but not good enough. It also often labeled critical issues as nitpicks, which is unacceptable in our context.
78.
▲
by
dakshgupta
2y ago
One word changes impacting output is interesting but also quite frustrating. Especially because the patterns don’t translate across models.
79.
▲
by
dakshgupta
2y ago
This might be what we experienced. We regularly have context reach 30k+ tokens.
80.
▲
by
dakshgupta
2y ago
The severity needing to be at the end was an important insight. It made the results much better but not quite good enough. We had it output a json with fields {comment: string, severity: string} in that order.
81.
▲
by
dakshgupta
2y ago
I would say they are useful but they aren’t magic (at least yet) and building useful applications on top of them requires some work.
82.
▲
by
dakshgupta
2y ago
Thank you, I will try this. I suspect we can extract some universal theory of nits and have a base filter to start with, and have it learn per-company preferences on top of that.
83.
▲
by
dakshgupta
2y ago
Since we posted this, two camps of people reached out: Classical ML people who recommended we try training a classifier, possibly on the embeddings. Fine tuning platforms that recommended we try their platform. The challenge there would be
84.
▲
How we made our AI code review bot stop leaving nitpicky comments
(greptile.com)
257 points
by
dakshgupta
2y ago
|
169 comments
85.
▲
by
dakshgupta
2y ago
Not having a free trial is odd, but $500 is worth it if it works. It only needs to save 10 dev-hours/month to be worth it. Greptile is free to try FWIW.
86.
▲
AI-generated memes that roast your repo
(greptile.com)
4 points
by
dakshgupta
2y ago
|
0 comments
87.
▲
Show HN: Generate changelog given repo and timeframe
(github.com)
1 points
by
dakshgupta
2y ago
|
0 comments
88.
▲
Chameleon: AI-powered changelog generator CLI
(github.com)
1 points
by
dakshgupta
2y ago
|
0 comments
89.
▲
by
dakshgupta
2y ago
I worry that this post assumes LLMs won't get much better over time. This is possible, but YC bets that they will. The right time to start an LLM application layer company is arguably 6-12 months before LLMs get good enough for that pu
90.
▲
by
dakshgupta
2y ago
Love the demo video! Three quick questions: Any specific reason to choose the terminal as the interface? Do you plan to make it more extensible in the future? (sounds like this could be wrapped with an extension for any IDE, which is exciti
More ›