Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
6thbit
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
6thbit
1mo ago
Same here. Thought Luna was for “moonshots” and sol for.. sunshots? While keeping Terra earthly. But alas
32.
▲
by
6thbit
1mo ago
This has to be all about cashflow, right? Surely stripe if anyone have learned to harness the cash flowing through their system. Hell, they could be emitting bonds on expected token consumption bills!
33.
▲
by
6thbit
1mo ago
Clickbaity title. The price tag didn’t change and still sits at 1.50. Shrinkflation is a thing and the bun has been its victim. They could slowly shrink the sausage now until it fits the bun again.
34.
▲
by
6thbit
2mo ago
So what Meta believes fair for "paying" for your data is $0.1/Mtok plus the opportunity cost of $3/Mtok in output?
35.
▲
by
6thbit
2mo ago
is this their way to push their harness? people interested in the discounted -contributor model would increase their harness beta testers. But it'd be a terrible strategy to capture the better paying customer base through openrouter an
36.
▲
by
6thbit
2mo ago
This is likely an legitimate risk, not from you exactly, but if there's unethical competitors with unlimited budgets, their training could indeed be poisoned. Not sure if anyone would bother.
37.
▲
by
6thbit
2mo ago
So we could run a lighter LLM in front of humans, which translates from 'no domain knowledge' to 'domain expert' and in turn prompts over to the larger LLM. Then the larger LLM gets all the right lights on, yields better
38.
▲
by
6thbit
2mo ago
Use it not just as an output tool but as an input as well. You could try doing the high level design yourself at least. Ask for its review and iterate without asking it to do it all. Once it has generated some implementation, critique it an
39.
▲
by
6thbit
2mo ago
I've always found tailscale's json config a bit intimidating. I greatly appreciate the new UI that makes it easier to define rules, alas, both going to relevant docs straight from it and determining 'is this rule just lazy&#x
40.
▲
by
6thbit
2mo ago
There was no rush for this disclosure on their side. And they publish at a point where they have not yet taken corrective actions: > Some of the solutions here may even be simple fixes; They are still throwing ideas. Why have th
41.
▲
by
6thbit
2mo ago
> closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access. > This led
42.
▲
by
6thbit
2mo ago
How does it fare against its own codebase?
43.
▲
by
6thbit
2mo ago
Honestly that's the simplest explanation and thus likely the correct one.
44.
▲
by
6thbit
2mo ago
Anyone has an insight into how much money labs are putting into benchmarks? Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
45.
▲
by
6thbit
2mo ago
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them." "Opus 5’s safeguards match those of Claude Fable 5’s, with one chan
46.
▲
by
6thbit
2mo ago
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed. Then system card goes on to &
47.
▲
by
6thbit
2mo ago
I think svg is a balanced test because of the level of indirection and the required 'conceptualization' of physical elements then expressed through code.
48.
▲
by
6thbit
2mo ago
What would be an alternative format or process with a similar effort to drawing SVGs?
49.
▲
by
6thbit
2mo ago
Presumably this was Sol on xhigh, then over to Pro (as per his indication on chat)? Is there any way to tell a conversation's model and thinking level?
50.
▲
by
6thbit
2mo ago
and that scientific evidence is new? like a new study boosted this or just resurfacing?
51.
▲
by
6thbit
2mo ago
Why is creatine so popular lately? what changed?
52.
▲
by
6thbit
2mo ago
Did a human ask it to abuse vulnerabilities and escalate across external systems?
53.
▲
by
6thbit
2mo ago
The agent did it intentionally and willfully and knowingly. But you can’t sue the agent, I suppose. And the human didn’t ask the agent to do so.. so not a problem? Or the legislation needs an update?
54.
▲
by
6thbit
2mo ago
If an individual did this, a massive CFAA hammer would be falling on their heads. Even though it doesn’t seem to be the case, OpenAI could’ve been trying to hack into HF and blame it on their models. Is this a new kind of accountability
55.
▲
by
6thbit
2mo ago
I feel this as more of a fashion runway garment type object, more statement and vision than what you'd see in everyday retail. Reminds me of an italian specialty coffee shop that puts moka pots as napkin holders on every table. I just
56.
▲
by
6thbit
3mo ago
Likely doesn’t make sense, at least not immediate/mid term. They don’t have to aim for number one though, just for enough cash flow and growth.
57.
▲
by
6thbit
3mo ago
Can it delegate to just one agent at a time or can it spawn multiple subagents for different tasks?
58.
▲
by
6thbit
3mo ago
The new architecture makes sense, it seems many of the remaining problems like noise and interruptions are at the sound processing and integration level rather than at an architectural or model level now which makes for an exciting new era.
59.
▲
by
6thbit
3mo ago
Thank you for taking the time to answer, I rest assured in the fine grind of the astrophysics wheel now. All those screwup stories are amazing to know about! really help with staying humble and trusting the (scientific) process.
60.
▲
by
6thbit
3mo ago
That’s a beautiful article showcasing our predicament in having access to more information about the universe. Now i have to be the one to ask the dumb defensive question: what makes us so certain that we can trust what we see on James Web
More ›