Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jerpint
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
jerpint
2y ago
Chefs hiss (sorry couldn’t resist)
62.
▲
Python's 'shelve' is useful for LLM debugging
(jerpint.io)
1 points
by
jerpint
2y ago
|
0 comments
63.
▲
by
jerpint
2y ago
It doesn’t solve the biggest problem with RAG, which is retrieving the correct sources in the first place. It sounds like they just use a secondary LLM to check if everything that was generated can be grounded in the provided sources. It mi
64.
▲
by
jerpint
2y ago
I’ve been noticing this too. The other day, cursor’s version of Claude sonnet (3.7) added a with open(file) as f: pd.read_csv(f) This was a mistake not worthy of even gpt3… I’ve also noticed I get overall better suggestions from Clau
65.
▲
by
jerpint
2y ago
Interesting, that would be a pretty blunder
66.
▲
by
jerpint
2y ago
VSCode in cloud would be great, GitHub tried something similar with GitHub.dev , I haven’t tried it in a while but it didn’t feel quite ready at the time, maybe things have changed
67.
▲
by
jerpint
2y ago
If they’re launching a cloud-based service, doesn’t this effectively remove the risk of running it locally ?
68.
▲
by
jerpint
2y ago
I wonder if this has to do with catastrophic forgetting to some extent; fine tuning on a large enough dataset to make the RLHF go away. Genius to add the “negative” code intent to the mix
69.
▲
by
jerpint
2y ago
Just tried but it returned an error
70.
▲
by
jerpint
2y ago
I think the big thing overlooked is how much the human steering the models matters. If you know what you’re doing and what changes you need, cursor and other tools make you so productive. If you don’t know what you’re doing, these things ca
71.
▲
by
jerpint
2y ago
you're right - I hadn't noticed! I fixed it now, thanks for pointing it out
72.
▲
by
jerpint
2y ago
will probably focus on getting the text out of the papers first, figures might be a good next step after that
73.
▲
by
jerpint
2y ago
whoa - i haven't yet played with MCP - might be a good first project!
74.
▲
by
jerpint
2y ago
For now , yes - abstracts and other metadata
75.
▲
Show HN: ArXiv-txt, LLM-friendly ArXiv papers
(arxiv-txt.org)
22 points
by
jerpint
2y ago
|
11 comments
76.
▲
Show HN: Realtime HTML Rendering with LLMs
(jerpint.io)
1 points
by
jerpint
2y ago
|
1 comments
77.
▲
by
jerpint
2y ago
The other day I tried an open source deep research implementation, and a ton of links returned 403s because I was using an agent. But it is for legitimate purposes. I think we need better ways of identifying legitimate agents working on my
78.
▲
by
jerpint
2y ago
The ability to add watermarks to text is really interesting. Obviously it could be worked around , but could be a good way to subtly watermark e.g. LLM outputs
79.
▲
by
jerpint
2y ago
I think a good analogy is being able to drive a car vs understanding the engine of a car. Both are useful, but you wouldn’t hire a mechanic drive you around
80.
▲
by
jerpint
2y ago
> In 50% and 90% experimental trials, they succeed in creating a live and separate copy of itself respectively I mean any half decent coding LLM can literally do ``` from transformers import SomeLLM model = SomeLLM.from_pretrained(“…”) `
81.
▲
by
jerpint
2y ago
Similar feeling with Gemini g suite integration
82.
▲
by
jerpint
2y ago
I wonder if this is a good alternative for running potentially malicious tools like random fine tunes of LLMs or coding agent outputs with little to no risk
83.
▲
by
jerpint
2y ago
I really wonder how much the self-supervised data flywheel will be enough to reproduce this model or not. There are also so many different tweaks that could be made to try to get these models to perform better, kudos to the team reproducing
84.
▲
I had different agents play 'The Password Game' – they didn't do so well
(jerpint.io)
2 points
by
jerpint
2y ago
|
0 comments
85.
▲
by
jerpint
2y ago
Browser use supports gemini via langchain
86.
▲
by
jerpint
2y ago
Thank you for all of your incredible contributions!
87.
▲
by
jerpint
2y ago
It’s funny because I actually use vim mostly when I don’t want LLM assisted code. Sometimes it just gets in the way. If I do, I load up cursor with vim bindings.
88.
▲
"The Password Game" Is a Solid Benchmark for Multimodal Agents
(jerpint.io)
1 points
by
jerpint
2y ago
|
0 comments
89.
▲
by
jerpint
2y ago
> This code repository and the model weights are licensed under the MIT License. DeepSeek-R1 series support commercial use, allow for any modifications and derivative works, including, but not limited to, distillation for training other
90.
▲
by
jerpint
2y ago
You can still game a test set without training on it, that’s why you usually have a validation set and a test set that you ideally seldom use. Routinely running an evaluation on the test set can get the humans in the loop to overfit the dat
More ›