Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ijk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
91.
▲
by
ijk
1y ago
Though Fireworks is one of the few providers that supports structured generation.
92.
▲
by
ijk
1y ago
One pattern that I've seen develop (in PydanticAI and elsewhere) is to constrain the output but include an escape hatch. If an error happens, that lets it bail out and report the problem rather than be forced to proceed down a doomed p
93.
▲
by
ijk
1y ago
I've been curious about grammar support for non-JSON applications. (i.e., I have some use cases where XML is more natural and easier to parse but Pydantic seems to assume you should only work with JSON.) Would guidance be able to handl
94.
▲
by
ijk
1y ago
Arcadia's biggest drawback was that it was dependent on Unity, which meant that it had all of the issues that Unity had plus some new ones. Without source code access it's difficult to exceed the features/performance of the b
95.
▲
by
ijk
1y ago
There are book scanners that don't require cutting the spine, though Anthropic doesn't seem to have used that approach.
96.
▲
by
ijk
1y ago
The judge's ruling from earlier certainly seemed to me to suggest that the training was fair use. Obviously, that's not part of the current settlement. I'm no expert on this, so I don't know the extent to which the earli
97.
▲
Optical Generative Models
(nature.com)
1 points
by
ijk
1y ago
|
0 comments
98.
▲
by
ijk
1y ago
Oh, OpenAI finally added it? Structured generation has been available in things like llama.cpp and Instructor for a while, so I was wondering if they were going to get around to adding it. In the examples I've seen, it's not somet
99.
▲
by
ijk
1y ago
I find casual knowledge particularly interesting, because it's the exact kind of thing that's most related to my day to day experience, but is simultaneously the exact thing that our encyclopedias and AIs omit.
100.
▲
by
ijk
1y ago
Exactly: at their core the only task they ate capable of is "complete this document." It just turned out that document completion was far more effective than anyone anticipated.
101.
▲
by
ijk
1y ago
I'm frustrated by the number of times I encounter people assuming that the current model behavior is inevitable. There's been hundreds of billions of dollars spent on training LLMs to do specific things. What exactly they've
102.
▲
by
ijk
1y ago
There's some evidence of valid relationships: you can build a map of Manhattan by asking about directions from each street corner and plotting the relations. This is still entirely referential, but in a way that a human would see some
103.
▲
by
ijk
1y ago
I think what trips people up is that LLMs and humans are both lossy, but in different ways. The intuitions that we've developed around previous interactions are very misleading when applied to LLMs. When interacting with a human, we&#x
104.
▲
by
ijk
1y ago
Ironically, scaling limits and evidence that quality vastly outweighs quantity suggests that all that web data is much less useful than buying and scanning a book. Most work with the Common Crawl data, for example, has ended up focusing on
105.
▲
by
ijk
1y ago
I'm still very confused by who is actually benefitting from the bots; from the way they behave it seems like they're wasting enormous amounts of resources on both ends for something that could have been done massively more efficie
106.
▲
by
ijk
1y ago
I am still a bit confused by what some of these crawlers are getting out of it; repeatedly crawling sites that haven't changed seems to be the norm for the current crawlernets, which seems like a massive waste of resources on their end
107.
▲
by
ijk
1y ago
Given the way attention works, it seems to me that AI has an even more concrete instantiation of cognitive load.
108.
▲
by
ijk
1y ago
I think the main difference is that sandboxing and simplifying the LLM's access to tools and data tends to be core functionality, whereas for XR it is more about performance and developer experience. I'm going to put a lot of work
109.
▲
by
ijk
1y ago
> I don't need to go on a VPN and be afraid to speak my mind here Well, so long as they don't crack down on VPNs here, like the UK is discussing.
110.
▲
by
ijk
1y ago
I know you're joking, but what the AI training lawsuits have said so far is that training and digitizing used books that you bought is fair use, but piracy isn't.
111.
▲
by
ijk
1y ago
Well, you can also batch your own queries. Not much use for a chatbot but for an agentic system or offline batch processing it becomes more reasonable. Consider a system were running a dozen queries at once is only marginally more expensive
112.
▲
by
ijk
1y ago
You mean that Americans used to be able to do that. Last year. For this coming year, it's over. Article: https://www.theverge.com/news/717308/irs-direct-file-gone-bi... HN discussion: https://new
113.
▲
by
ijk
1y ago
I got this just now on my first try with the free preview of ChatGPT (which isn't using the latest version, but is currently available on their site). I was surprised, I expected to have to work harder for it to fail like that.
114.
▲
by
ijk
1y ago
The sneaky thing is that the things we used to rely on as signals of verification and credibility can easily be imitated. This was always possible--an academic paper can already cite anything until someone tries to check it [1]. Now, someth
115.
▲
by
ijk
1y ago
I asked, just now: > How many 'r's are in strawberry? > ChatGPT said: The word "strawberry" has 2 'r's. It's going to be fairly reliable at this point at basic arithmetic expressed in an expected way
116.
▲
by
ijk
1y ago
True - though in the actual case of your examples, calcpay, process_user_input, and ProcessUserInput all encode into exactly 3 tokens with GPT-4. Which is the exact kind of information that you want to know. It is very non-obvious which one
117.
▲
by
ijk
1y ago
There's some libraries that might make this easier to implement: https://github.com/kanishkamisra/minicons
118.
▲
by
ijk
1y ago
It would probably be legal with digital copies; it's just that book publishers have been very zealous in preventing the existence of a market for used digital books. Copyright has been very silly in the digital realm from the beginning
119.
▲
by
ijk
1y ago
It has caused many community forums to close, past tense. Many cited the uncertainty about what is actually required, the potential high cost of compliance, the danger of failing to correctly follow the rules they're not certain abou
120.
▲
by
ijk
1y ago
Anthropic is arguing, among other things, that bankrupting the smallest major AI player while all the bigger ones do the same thing should be a reason to reduce the consequences.
More ›