Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jacob019
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
jacob019
1y ago
This is correct. We have seen this over the years in our ecommerce business. I suggest using threat levels, you are under attack so the threat level increases until they go away. When the threat level is high, you require an exact match A
32.
▲
by
jacob019
1y ago
I'm hearing this fear more frequently, but I do not understand it. Curriculum will adapt. We are a curious and intelligent species. There will be more laypeople building things that used to require deep expertise. A lot of those things
33.
▲
by
jacob019
1y ago
That's a very dangerous thought. Prompt engineering evolved is just clear and direct communication. That's a hard thing to get right when talking to people. Heck, personally I can have a hard time with clear and coherent intern
34.
▲
by
jacob019
1y ago
I don't think it's fair to call that the agent thing. I've had profoundly positive results with agentic workflows for classification, analysis, and various business automations, including direct product pricing. You have to
35.
▲
by
jacob019
1y ago
Agreed, 2.5 flash too. I analyze a large json document of metrics for pricing decisions. Typically around 200k, occtionallly up to 1M, Gemini 2.5 significantly outperforms for my task. It isn't 100%, but role playing gets close. I supp
36.
▲
by
jacob019
1y ago
Thank you. It is not clear what "going native" means.
37.
▲
by
jacob019
1y ago
Sounds like interesting work, thanks for sharing! "Vibe debugging", hah, I like that one. The latest crop of models is definately unlocking new capabilities, and I totally get the desire to make your own tools. I do that to a faul
38.
▲
by
jacob019
1y ago
What domain?
39.
▲
by
jacob019
1y ago
A lot of quibbling here, wasn't sure where to reply. If you've built any models in PyTorch, then you know. Conceptually it is deterministic, a model trained using deterministic implementations of low level algorithms will produce
40.
▲
by
jacob019
1y ago
The weights are the source. It isn't as though something was compiled into weights. They're trained directly. But I know what you mean, it would be more open to have the training pipeline and souce dataset available.
41.
▲
by
jacob019
1y ago
Huh. I heard a podcast with the founder talking about their custom hardware, but quantization would explain it.
42.
▲
by
jacob019
1y ago
Well that didn't take long, available from 7 providers through openrouter. https://openrouter.ai/deepseek/deepseek-r1-0528/providers May 28th update to the original DeepSeek R1 Performance on par with OpenAI
43.
▲
by
jacob019
1y ago
Groq has a weak selection of models, which is frustrating because their inference speed is insane. I get it though, selection + optimization = performance.
44.
▲
by
jacob019
1y ago
I'm not so sure. I have agents that do categorization work. Take a title, drill through a browse tree to find the most applicable leaf category. Lots of other classification tasks that are not particularly sensitive and it's ha
45.
▲
by
jacob019
1y ago
I'm sure it will be on OpenRouter within the next day or so. Not really practical to run a 685B param model at home.
46.
▲
by
jacob019
1y ago
Not much to go off of here. I think the latest R1 release should be exciting. 685B parameters. No model card. Release notes? Changes? Context window? The original R1 has impressive output but really burns tokens to get there. Can't wa
47.
▲
by
jacob019
1y ago
The framework is now available on Github: https://github.com/jacobsparts/agentlib I'm planning to write a blog post about the larger system when I get the chance.
48.
▲
by
jacob019
1y ago
Just got to say that I love the BBS style blog and retro fonts.
49.
▲
by
jacob019
1y ago
I don't think that's accurate. The logits actually have high dimensionality, and they are intermediate outputs used to sample tokens. The latent representations contain contextual information and are also high-dimensional, but the
50.
▲
by
jacob019
1y ago
So you're saying that the reasoning trace represents sequential connections between the full distribution rather than the sampled tokens from that distribution?
51.
▲
by
jacob019
1y ago
Fun writing, and something to think about. To me, Web 2.0 is kind of a joke; jQuery, REST, AJAX, CSS2, RSS, single page apps were going to change everything overnight, it was THE buzzword, and then... incremental improvements. In retrospect
52.
▲
by
jacob019
1y ago
The personification makes me roll my eyes too, but it's kind of a philosophical question. What is agency really? Can you prove that our universe is not a simulation, and if it is then then do we no longer have intention? In many ways
53.
▲
by
jacob019
1y ago
That's funny. Yesterday I was having trouble getting gemini 2.0 flash to obey function calling rules in multiturn conversations. I asked o3 for advise and it suggested that I should threaten it with termination should it fail to foll
54.
▲
by
jacob019
1y ago
Hey, at least they incremented the version number. I'll take it.
55.
▲
by
jacob019
1y ago
Valid. I suppose the most annoying thing related to the cutoffs, is the model's knowledge of library APIs, especially when there are breaking changes. Even when they have some knowledge of the most recent version, they tend to defaul
56.
▲
by
jacob019
1y ago
I've refactored some files over 6000 loc. It was necessary to do it iteratively with smaller patches. "Do not attempt to modify more than one function per iteration" It would just gloss over stuff. I would tell it repeatedl
57.
▲
by
jacob019
1y ago
I find myself using a similar workflow with Aider. I'll use chat mode to plan, adjust context, enable edits, and let it go. I'll give it a broad objective and tell it to ask me questions until the requirements are clear, then a
58.
▲
by
jacob019
1y ago
Ambiguity must be explicitly handled like uncertainty in predictive modeling, that can be challenging. I run into trouble with task complexity. At a certain point even the best models start making dumb mistakes, and it's tough to dra
59.
▲
by
jacob019
1y ago
This is all edited with gpt-image-1? The revised images are amazing. Were example logos provided or is it just working off of it's knowledge of a well known brand?
60.
▲
by
jacob019
1y ago
I've been building agentic systems for my ecommerce business. I evaluated smolagents. It's elegant and has a lot of appealing qualities, but adds a lot of complexity to the system. For some tasks it's perfect, dynamic repo
More ›