Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lmeyerov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
lmeyerov
5mo ago
We have been getting increasingly hit by this. We do defense, not offense, and AI refusals to run defense prompts has been going noticeably up. Historically, tasks used to only get randomly rejected when we were doing disaster management AI
62.
▲
by
lmeyerov
6mo ago
It's been fun benchmarking AI investigations at botsbench.com . Part of it is checking for these kinds of issues - we recently started seeing contamination in our first generation challenge, and less obvious, agent sandbox escapes for
63.
▲
by
lmeyerov
6mo ago
arxiv: https://arxiv.org/html/2311.02206v5 I've been a fan of this series! The work by this team, as well as kuzudb (acquired by apple) and relational.ai, have similar vibes. One area that has been especially inte
64.
▲
by
lmeyerov
6mo ago
At RSAC, there were a ton of agentic security startups converging on ebpf monitors for this reason. Eg, sondera gave a fun talk at graph the planet where they did that + exposed with a policy layer over agent traces via Cedar (used in AWS I
65.
▲
by
lmeyerov
6mo ago
Maybe it's useful to split out B1) KG pipelines from the choice of B2) simple property graph ontologies & queries vs advanced rdf ontologies and sparql queries It sounds like you are thinking about KG pipelines, but I'm unclea
66.
▲
by
lmeyerov
6mo ago
It's interesting to think of where the value comes from. Afaict 2 interesting areas: A: One of the main lessons of the RAG era of LLMs was reranked multiretrieval is a great balance of test time, test compute, and quality at the expens
67.
▲
by
lmeyerov
6mo ago
This is great We reached a similar conclusion for GFQL (oss graph dataframe query language), where we needed an LLM-friendly interface to our visualization & analytics stack, especially without requiring a code sandbox. We realized we c
68.
▲
by
lmeyerov
6mo ago
When Anthropic's CPO left Figma's board this week, that was my first question . Oof.
69.
▲
by
lmeyerov
6mo ago
Most companies and their vendor ecosystems run on OSS Worse, "attackers no longer break in, they log in", so the supply chain attacks harvesting credentials have been frightening
70.
▲
by
lmeyerov
6mo ago
We have this issue in GFQL right now. We wrote the first OSS GPU cypher query language impl, where we make a query plan of gpu-friendly collective operations... But today their steps are coordinated via the python, which has high constant o
71.
▲
by
lmeyerov
6mo ago
This is great work by Dawn Song 's team. A huge part of botsbench.com for comparing agents & models for investigation has been in protecting against this kind of thing. As AI & agents keep getting more effective & tenacious
72.
▲
by
lmeyerov
6mo ago
Instead of scanning more code, afaict what you seem to want is instead, scan on the same small area, and compare on how many FPs are found there. A common measure here is what % of the reported issues got labeled as security issues and fixe
73.
▲
by
lmeyerov
6mo ago
Maybe there's a fundamental miscommunication here of what evals are? Evals apply not just to LLMs but to skills, prompts, tools, and most things changing the behavior of compound AI systems, and especially like the productivity claims
74.
▲
by
lmeyerov
6mo ago
We find it true in Louie.ai evals (ai for investigations), about a 10-20% lift which meaningful. It'd measured here: botsbench.com . Unfortunately, undesirable in practice due to people being token-constrained even before. One case is
75.
▲
by
lmeyerov
6mo ago
I've found value in architectural research before r&d tier projects like big changes to gfql, our oss gpu cypher implementation. It ends up multistage: - deep research for papers, projects etc. I prefer ChatGPT Pro Deep Research he
76.
▲
by
lmeyerov
6mo ago
It sounds like the answer is "No, there is no repeatable eval of the core AI coding productivity claim, definitely not on one of the many AI coding benchmarks in the community used for understanding & comparison, and there will not
77.
▲
by
lmeyerov
6mo ago
I'm not too familiar with etsy, but presumably most etsy sellers are closer to being lemonade stands than they are to being ikea And yes, sometimes it's nice to support a local lemonade stand. For my family's income, I know w
78.
▲
by
lmeyerov
6mo ago
My question was on claims like "5x productivity boost in merged PRs (lots of open PR & merge rate goes down, but net positive)", eg, does this change anything on swe-bench or any other standard coding eval?
79.
▲
by
lmeyerov
6mo ago
Evals let us agree on the baseline, measurement, etc, and compare if simple things others do perform just as well. For same reason, instead of 'works on my box' and 'my coding style', use one of the many community evals
80.
▲
by
lmeyerov
6mo ago
Evals or GTFO
81.
▲
by
lmeyerov
7mo ago
Speaking of embeddable, we just announced cypher syntax for gfql, so the first OSS CPU/GPU cypher query engine you can use on dataframes Typically used with scaleout DBs like databricks & splunk for analytical apps: security/f
82.
▲
by
lmeyerov
7mo ago
I would ban apps using unsafe ad platforms If I was simultaneously also the owner of the ad platform, I'd fix it & knock out the bad players, or get ready to be sued for a decade+ of knowing malpractice And if I was a US citizen se
83.
▲
by
lmeyerov
7mo ago
There are plenty of bad actors The interesting part is Google & Apple, as part of explaining to courts why their large app store fees are legit and not proof of monopoly positions, hid behind the security argument that they need to be t
84.
▲
by
lmeyerov
7mo ago
*legal in the US
85.
▲
by
lmeyerov
7mo ago
You can trace the big players If Google & Apple & friends refused to take a rake and opened distribution, then I'd agree, net neutrality etc, not their problem But they own so much, and so deep into the pipeline, and explain th
86.
▲
by
lmeyerov
7mo ago
Apple and Google are facilitating the data sales Specifically, these big companies revenue share with app companies who in turn increase monetization via selling your private information, esp via free apps. In exchange for Apple etc super h
87.
▲
by
lmeyerov
7mo ago
Once my code exists and passes test, I generally move on to having it iteratively hunt for bugs, security issues, and DRY code reduction opportunities until it stops finding worthwhile ones. This doesn't always work as well as I'd
88.
▲
by
lmeyerov
7mo ago
We can play that game - items like GIL-free interpreters and memory views are pretty relevant to folks on the more demanding side of scientific computing. But my point is this is a head-in-sand game when the community vastly outweighs any i
89.
▲
by
lmeyerov
7mo ago
The phenomena you're describing is why Cobol programmers still exist, and simultaneously, why it's increasingly irrelevant to most programmers The killer feature is ecosystem: Easily and reliably reusing other libraries and tools
90.
▲
by
lmeyerov
7mo ago
I liked they did this work + its sister paper, but disliked how it was positioned basically opposite of the truth. The good: It shows on one kind of benchmark, some flavors of agentically-generated docs don't help on that task. So naiv
More ›