2 ms·
Gotcha', but I'm just trying to see the audit for how claim X was rated and based on what sources. If we're looking at the Claude logs, we have huge files that
by dvt 7mo ago
Gotcha', but I'm just trying to see the audit for how claim X was rated and based on what sources. If we're looking at the Claude logs, we have huge files that have things like this[1]:
{"id": "claim_0081", "date": "2023-02-11", "claim": "Current Level 2 self-driving operates under easy conditions and is nowhere close to handling real-world complexity.", "type": "descriptive", "target": "Level 2 self-driving", "status": "supported", "horizon": null}
Why is this supported? How is this supported? Waymo would probably disagree, etc. Here's another one:
{"id": "claim_0083", "date": "2023-02-11", "claim": "Tesla's product naming ('Autopilot', 'Full Self Driving') misleads customers into thinking the cars are more capable than they are, potentially causing accidents and deaths.", "type": "causal", "target": "Tesla marketing", "status": "supported", "horizon": null}
I fully agree that TSLA engages in all kinds of deceptive marketing, but to fully support the stunning claim that it potentially causes deaths is, uh, a bit much. I mean, at least tell me who's saying this. What's the provenance?
If Claude itself rated the claims, which seems the be the case unless I'm totally off base, I fail to see how we're actually doing anything at all here. Right now I'm working on a local research agent, and I'm being absolutely meticulous about storing browsed webpages, snippets, etc. into short-term (session) LLM memory or a long-term (cross-session) SQLite db.
[1] https://github.com/davegoldblatt/marcus-claims-dataset/blob/main/claude/claude_claims_condensed.jsonl https://github.com/davegoldblatt/marcus-claims-dataset/blob/...
- davegoldblatt 7mo ago[flagged]