Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ssivark
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
ssivark
5d ago
The actual tokens might be non-deterministic, but you could look for proxy measures that are supposed to be invariant. Eg. correctness/performance on benchmarks, "thinking level" on complex problems, etc
2.
▲
by
ssivark
7d ago
WDYM -- I thought "try another way" is supposed to list all possible options instead of cycling through them!
3.
▲
by
ssivark
9d ago
> OpenAI cautioned that the reports were individual snapshots and “shouldn’t be considered reflective of how often misalignment occurs.” Sounds like a proper infestation of roaches!
4.
▲
by
ssivark
11d ago
The time is ripening to disrupt Android/iOS. The place to start would be to rethink the device around AI. No need for apps or messy integrations with third-party crap. Just a simple device with a great new UX (not typing on a touchpad!
5.
▲
by
ssivark
11d ago
Perhaps don't deploy random weights of unknown origin? Also not every model provider might be capable of babysitting all your uncontrolled agent deployments. If you want SLOs, get into a contractual relationship with entities whose wei
6.
▲
by
ssivark
11d ago
The problem with third party audits is that it allows OAI/Ant to shrug off any further responsibility and claim that they are following best practices (basically, reward hacking). The only real solution is to make them absorb liability
7.
▲
by
ssivark
12d ago
> Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doi
8.
▲
by
ssivark
12d ago
> to separate ZDR, attestation, and confidential routes Could you please clarify what that means? Given what I've been searching for, I might in principle be part of your intended customer profile, but I can't figure out whethe
9.
▲
by
ssivark
13d ago
What about inference providers like Baseten, Modal, Fireworks, Together, etc? I thought one of their value propositions was inference (using open weights models) that guarantees with crisp terms that they will not use your data.
10.
▲
by
ssivark
15d ago
There is potentially a world of difference between how you interpret what is fair and what the terms of service contractually guarantee.
11.
▲
by
ssivark
19d ago
Don't know why the below comment by killix got flagged; it's a legitimate point. In the current version of my setup, I've decided to accept that tradeoff. But it would also be interesting to check whether agent behavior can b
12.
▲
by
ssivark
19d ago
> Someone who has no idea what a standard deviation is can't intuit about distributions. I disagree vehemently with this claim. I could cite my experience in teaching this topic to liberal arts / humanities college students in
13.
▲
by
ssivark
19d ago
Look at the histogram and think about what distribution one could reasonably impute from samples. And what you would set as bounds for "outliers", per your needs. While we're at it, let me also say that it might be useful t
14.
▲
by
ssivark
19d ago
My personal opinion is that statistics textbooks usually come from a prescriptive perspective, and that makes it challenging for the reader to get visceral intuition for what is actually going on. Any reader would be far better off just vis
15.
▲
by
ssivark
19d ago
To the extent that skills are contextual guidance (for this author, this project, etc) and not just (raw) capabilities they are unlikely to be eaten by models. I maintain all my skill files in a central location (like dotfile management
16.
▲
by
ssivark
19d ago
It's one of several downstream problems, including poorer quality of life by various metrics. By emphasizing a PR problem for businesses more than other human problems downstream or the root problem, you are demonstrating priorities th
17.
▲
by
ssivark
20d ago
That's like saying a law and order problem is a PR problem because word got out that crime is spiking. The PR problem is downstream consequence of what can be measured, but the root cause is upstream. Focusing directly on solving the
18.
▲
by
ssivark
21d ago
Which is why I said that cost might be a better metric than output tokens. But even that is somewhat misleading -- because there isn't really a single price -- there are a variety of offers / deals / subsidies, including subs
19.
▲
by
ssivark
21d ago
...the remaining 10% of their Claude usage ;-) > modes are the load-bearing piece: Okay, this write-up is filled with Claudisms.
20.
▲
by
ssivark
21d ago
Duh, that's just a benchmarking artifact. If you run one model at 4x recurrence compared to another, you get 4x thinking without increasing the tokens. So of course it's going to dominate the perf at given output token level. The
21.
▲
by
ssivark
22d ago
I imagine this will only become more important as we have more agents working on codebases in parallel; so a clean solution would be extremely valuable. On the flip side, the rise of agentic coding means that the library/batteries mism
22.
▲
by
ssivark
26d ago
There might be a lot going on, depending on what kind of programming tasks you use AI for, but here's one hunch: I think the current AI usage pattern disrupts our core sense-making mechanism on the code we create. It jumps from nothing
23.
▲
by
ssivark
1mo ago
> most people would interpret that to mean You're just asserting common convention among some implicitly selected audience that you consider "most" people, rather than justifying why that is the most reasonable practice. N
24.
▲
by
ssivark
1mo ago
> Before you could have invented PageRank, you must think in graphs. For anyone puzzling over what is the graph-based perspective, there's actually a very elegant and simple mathematical way to derive page rank. Define the directed
25.
▲
by
ssivark
1mo ago
I'm not an accountant and don't claim to have a clean answer to how it should be accounted, but I hope I can highlight the conundrum. Suppose you run a brokerage or some kind of marketplace enabling transactions. Should all transa
26.
▲
by
ssivark
1mo ago
> It’s so hard to finish an idea that is not yours and is just suggested by AI Yes! And that's because you can't really delegate ownership/authorship to AI (at today's capability levels). Coding agents might have been
27.
▲
by
ssivark
1mo ago
The author makes an important point about ownership but that got watered down to code understanding and debugging. What matter is not just ownership of the code that got written, but also ownership of the problem identification and the ch
28.
▲
AI assistants need adaptive conversational tempo to be good cognitive companions
(woventhought.substack.com)
2 points
by
ssivark
1mo ago
|
0 comments
29.
▲
by
ssivark
1mo ago
We're in an era where soon (if not already) it will become straight forward to direct a clanker to move an application from one gui framework to another. Once we have robust automated computer use, that becomes the verification loop, a
30.
▲
by
ssivark
1mo ago
> We use issue trackers and pull requests to manage work, people write documentation and communicate over email, instant messaging, and in meetings. Keeping all of the information in these channels synchronized and up to date is a full t
More ›