Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zurfer
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
zurfer
8mo ago
Yes, but honestly what's the best source when reporting about a person? Their personal website no? I think it's a hard problem and I feel there are a lot of trade-offs here. It's not as simple as saying chatgpt is stupid or t
62.
▲
by
zurfer
8mo ago
to be fair, i think usage has increased a lot because of coding agents and some things that worked well for now can't scale to the next 10x level.
63.
▲
by
zurfer
8mo ago
6x more expensive
64.
▲
by
zurfer
9mo ago
we use codex review. it's working really well for us. but i don't agree that it's straightforward. moving the number of bugs catched and signal to noise ratio a few percentage points is a compounding advantage. it's a va
65.
▲
by
zurfer
9mo ago
I certainly don't think people on HN are dumb, I'm surprised that the sentiment towards this is just talking so much about the downside and not the upside. And look I do agree that humans should be the one responsible for the thin
66.
▲
by
zurfer
9mo ago
I'm a bit shocked to see so many negative comments here on HN. Yes, there are security risks and all but honestly this is the future. It's a great amplifier for hackers and people who want to get stuff done. It took some training
67.
▲
by
zurfer
9mo ago
"Nobody" is against open source. People are against bad products, high maintenance, high complexity... .. or against working for free without much reward. It would be great if there was a way for humanity to have and develop free
68.
▲
by
zurfer
9mo ago
It's an interesting industry that needs billions to bring a new drug to market. At the same time it creates a lot of value to a patient. But the manufacturing of a single dose is usually tiny. Now how do you price that? The profits her
69.
▲
by
zurfer
9mo ago
Yes it's a bit disappointing but probably captures the current American and Chinese opinion quite well. Europe as a whole has a lot of good things going for it but I do agree that it's less ambitious on average than these 2 power
70.
▲
by
zurfer
10mo ago
You're too kind. Even the CEO of Google retweeted how well Gemini 2.5 did on Pokemon. There is a high chance that now it's explicitly part of the training regime. We kind of need a different kind of game to know how well it genera
71.
▲
by
zurfer
10mo ago
Yes I tried it with minimal and it's roughly 3 seconds for prompts that take flash 2.5 1 second. On that note it would be nice to get these benchmark numbers based on the different reasoning settings.
72.
▲
by
zurfer
10mo ago
It's a cool release, but if someone on the google team reads that: flash 2.5 is awesome in terms of latency and total response time without reasoning. In quick tests this model seems to be 2x slower. So for certain use cases like quick
73.
▲
by
zurfer
10mo ago
How do you understand meritocracy? It seems natural that those that do valuable things get rewarded a lot. Ideally everyone would get the same chances to do valuable things but that's not how the world is setup. Unfortunately. However
74.
▲
by
zurfer
10mo ago
It drives me a bit crazy when people say OpenAI has no moat. Yes, companies like Google can catch up and overtake them, but a moat is merely making it hard and expensive. 99.999.. perc of companies can't dream of competing with OpenAI.
75.
▲
by
zurfer
11mo ago
I think it's a great article that should discourage a lot of people to waste resources. To really do it you have to treat this article as a to-do list of challenges to overcome. If you have no ideas on how to address those challenges y
76.
▲
by
zurfer
11mo ago
well, maybe no one felt informed enough to write this, so it was outsourced to the llm (imposter syndrom) or it was pure laziness.
77.
▲
by
zurfer
11mo ago
There are shareholders/owners and CEOs. You can certainly have an AI CEO if the board of directors wants that. Although depending on the jurisdiction CEOs might need be humans, but surely not everywhere. And you could even imagine AI o
78.
▲
by
zurfer
11mo ago
It also tops LMSYS leaderboard across all categories. However knowledge cutoff is Jan 2025. I do wonder how long they have been pre-training this thing :D.
79.
▲
by
zurfer
11mo ago
I heard the story once on how you migrate these old systems: 1 you get a super extensive test suite of input - output pairs 2 you do a "line by line" reimplementation in Java (well banks like it). 3 you run the test suite and trac
80.
▲
by
zurfer
11mo ago
You're right, I didn't produce evidence in my comment. Smoking and lung cancer are more clear cut than social media problems in general. The effects of social media are more complex and nuanced than smoking. There are a lot of stu
81.
▲
by
zurfer
11mo ago
I don't buy it. The most relevant critique is see is that it's hard to control the age of your users without removing anonymous accounts thus limiting privacy. Well, it's a hard problem but it doesn't feel impossible to
82.
▲
by
zurfer
11mo ago
The issue is rather the algorithmic feed optimized for grabbing our attention. It's definitely addictive and should be regulated like other drugs. Give people technology, but let's have an honest conversation about it finally. As
83.
▲
by
zurfer
11mo ago
The current go to solution for the kinds of problems that TabPFN is solving would be something like XGBoost. In general it's a good baseline, but the challenge is always that you need to spend a lot of time feature engineering and twea
84.
▲
by
zurfer
11mo ago
this was the article I had in mind, when writing this: https://dynomight.substack.com/p/chess
85.
▲
by
zurfer
11mo ago
Yeah I was not precise; it was `gpt-3.5-turbo-instruct`, other variants weren't trained on it apparently. https://dynomight.substack.com/p/chess
86.
▲
by
zurfer
11mo ago
It makes me wonder if we'll see an explosion of purpose trained LLMs because we hit diminishing returns on invest with pre training or if it takes a couple of months to fold these advantages back into the frontier models. Given the siz
87.
▲
by
zurfer
1y ago
Well, I can't comment much on Genie, but the core question is always how you scale the complexity. In Dot, it's divide and conquer. If you have several different teams each of them has to maintain their knowledge base. A bunch of
88.
▲
by
zurfer
1y ago
Reliability is the important dimension to focus on, but what is your baseline? I've worked in data and me and my colleagues regularly had to confront bugs (on many levels) that were communicated to end users
89.
▲
by
zurfer
1y ago
Germany is the 5th biggest economy in the world, then there is Austria and Switzerland. The claim that nobody would pay for German to SQL seems a bit pessimistic ;) I'd also love to understand better why you think that there is no &quo
90.
▲
by
zurfer
1y ago
Cofounder of one of those analytics agents here ( https://getdot.ai ). The promise of the technology is not that it can deal with any arbitrarily complex Enterprise setup, but rather that you expose it with enough guidance on a co
More ›