Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tomberin
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
tomberin
3y ago
I was impressed with the Demo, ready to pay 10 and no option to sign up with email :(
2.
▲
by
tomberin
3y ago
This isn’t how any of this works. Manifests are built at boarding, visas work at entry. You are inventing problems to take the side of a corporation.
3.
▲
by
tomberin
4y ago
not tabloids, not a hoax, did you go on Twitter during that period? I have him blocked and saw his tweets
4.
▲
by
tomberin
4y ago
We can't comment on why, but there's no rational way to watch his behavior and assume he isn't obsessed with being liked. He constantly tweets about his own tweets performance, makes humiliating appearances on stages, and pre
5.
▲
by
tomberin
4y ago
It can't cite articles, if it told you it did and the link was gone that's because it was a hallucination.
6.
▲
by
tomberin
4y ago
There's a huge warning on the first page. This is a weird stance. Don't use it if you're at all concerned.
7.
▲
by
tomberin
4y ago
See the FAQ :)
8.
▲
by
tomberin
4y ago
This is true of any webscraper though, you need to santitize any content you collect from the web. If a person wanted a scraper to get something different from the browser, they could easily use UA sniffing to do so. (I've seen it thi
9.
▲
by
tomberin
4y ago
The author asked me to share this here: https://mastodon.social/@jamesturk/110086087656146029 He's looking for a few case studies to work on pro bono, if you know someone that needs some data that meets certain cr
10.
▲
by
tomberin
4y ago
FWIW, That's been my use case, when I saw the author post his initial examples pulling data from Wikipedia pages I dropped my cobbled together scripts and started using the tool via CLI & jq.
11.
▲
by
tomberin
4y ago
TIL, thanks!
12.
▲
by
tomberin
4y ago
No inference is needed. IME it can do a single page in ~10s, $0.01/page. Not practical for most use cases, great for a limited few right now.
13.
▲
by
tomberin
4y ago
These kind of one-shot examples are exactly where this hit for me. I was in the middle of some research when I saw him post this and it completely changed my approach to gathering the ad-hoc data I needed.
14.
▲
by
tomberin
4y ago
:) Agree, but the scraping arms race is way beyond that, if someone doesn't want their page scraped this isn't a threat to them.
15.
▲
by
tomberin
4y ago
I was most worried about #2 but surprised how much temperature seems to have gotten that under control in my cases. The author added a HallucinationChecker for this but said on Mastodon he hasn't found many real-world cases to test it
16.
▲
by
tomberin
4y ago
It requires API access, temperature=0 means completely deterministic results but possibly worse performance. Higher temperature increases "creativity" for lack of a better word, but with it, hallucination & gibberish.
17.
▲
by
tomberin
4y ago
Not the author, but it seems like the separation of system & user messages actually prevents page content from being used as an instruction. This was one of the first things I tried and IME, couldn't actually get it to work. I&#x
18.
▲
by
tomberin
4y ago
Perhaps not, the author mentioned on Mastodon that he was exploring simpler models.
19.
▲
by
tomberin
4y ago
It seems like he's setting temperature=0 which also means it is deterministic. Anecdotally, I've been playing with it since he posted an earlier link & it does shockingly well on 3.5 and nearly perfectly on 4 for my use cases.
20.
▲
Experimental library for scraping websites using OpenAI's GPT API
(jamesturk.github.io)
378 points
by
tomberin
4y ago
|
139 comments
21.
▲
Web Scraping with GPT-4
(jamesturk.net)
5 points
by
tomberin
4y ago
|
1 comments