Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ddod
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
91.
▲
by
ddod
13y ago
I'm glad to see a story like this getting some press as I've suspected that I've been dealing with something very similar for years now. Every so often I get an email from Facebook or some other service asking me to confirm a sign up I neve
92.
▲
by
ddod
13y ago
Eventually, yes, if there's enough demand. Right now I'm focusing on getting more data, as even though my sample size is rather large, my per-subreddit samples are still small.
93.
▲
Show HN: Reddit analytics
(crowdlistener.com)
5 points
by
ddod
13y ago
|
2 comments
94.
▲
by
ddod
14y ago
Node.js with nothing fancy. It's only about 80 lines of code serverside.
95.
▲
by
ddod
14y ago
I've upped all the sizes so hopefully it's a bit better now. I had only tested it on an old Macbook and a Nexus 7, where the original font, size, and color were all very readable to me. Thanks for the heads up.
96.
▲
by
ddod
14y ago
Sorry for the eye-strain; I switched it from #777 to #444 just now, so hopefully there's a marginal improvement. My scraper isn't currently set to get that deep, but the most upvoted per time period is a good idea that I'll look into.
97.
▲
by
ddod
14y ago
Thanks for the feedback! It's actually creating its own database so as time goes on, it will show everything since it's been running. I didn't want to scrape through the backlog as I've already gotten banned a couple times while making this
98.
▲
Show HN: Showing HN
(showinghn.com)
66 points
by
ddod
14y ago
|
40 comments
99.
▲
by
ddod
14y ago
Thanks Paul! I'm reluctant to try this in conjunction with developing any HN scrapers since I'm not sure what set it off in the first place and your language suggests it will only unban the IP once (I will, however, make sure the CMU IP I w
100.
▲
by
ddod
14y ago
I can't know what you mean by minimal load, but I'm guessing you could get by with 128mb RAM to run a few personal sorts of sites, but I've seen $7/m for 1-2gb RAM. Be sure to read comments to make sure the company is decent, as most are ru
101.
▲
by
ddod
14y ago
I'd assume most HN users would spring for a VPS, and if you're looking for a good price, check out lowendbox. You can easily get something for $12 a year.
102.
▲
Ask HN: Why do HN users still obscure their e-mail addresses?
42 points
by
ddod
14y ago
|
39 comments
103.
▲
by
ddod
14y ago
Could someone explain to me how this could be leveraged (or if it could be) to gather a sort of stream of messages, a la the Twitter streaming API or reddit.com/r/all/comments.json? I'd be interested in doing some language statistics and co
104.
▲
by
ddod
14y ago
There's no actual strict definition of terrorism; just a bunch of disparate attempts at one by various government agencies. The reason this is, is because it's so painfully, obviously subjective. If we use the definition provided by driveby
105.
▲
by
ddod
14y ago
I'm working on the assumption of statistical significance of people using "yea" pronounced "yeah" vs. the very few people who might have reason to take to Twitter to use "yea" in a voting or biblical context.
106.
▲
by
ddod
14y ago
Great points. I went into this knowing full-well that to many, many people, these misspellings were not considered as such. As the graph suggests, though, the majority of people still use the "older" spellings of "yeah" and "anyway". There
107.
▲
by
ddod
14y ago
There's no link between spoken and written English in the context of yea:yeah, as they're pronounced the same but either written correctly or incorrectly. As for anyways, I think your classification there will depend on who you speak to (wh
108.
▲
by
ddod
14y ago
Those might work for Twitter, but there may be some lurking variables on Reddit. While they never bother to catch yea and anyways, Redditers do have a tendency to pile on about the few grammatical/spelling errors they can spot. For example,
109.
▲
by
ddod
14y ago
That's why I'm not relying on any single metric, as well as framing it in the context of time. If Justin Beiber asks his fans to vote yea or nay on something, it might mess up the yea:yeah metric, but it should eventually normalize, and whi
110.
▲
by
ddod
14y ago
Sample size is the anyway|anyways|yeah|yea instances. I get those as they get posted, so that's why it fluctuates with time. Since they're so commonly used, it should incidentally give you an idea of all of Twitter's load. I'm also grabbing
111.
▲
by
ddod
14y ago
I chose those two to start things off for two reasons: The first is that they stand out to me as pretty basic misspellings that don't appear in any sort of popular literature, so it speaks to someone's reading experience. You can take that
112.
▲
by
ddod
14y ago
Thanks. It compares misspellings to correct spelling counterparts (anyways:anyway), so that, in itself, should account for sample size changes in Twitter. The Reddit sample size (should) stay constant, as I'm grabbing /all/comments every 15
113.
▲
by
ddod
14y ago
I made this over the last two days in Node.js. The analysis is still pretty simple, and I'd like to expand it over time. It updates every 30 minutes, and you can already see that there's some significant shifts in literacy between the dayti
114.
▲
Show HN: Automated writing analysis of Twitter and Reddit feeds
(anthropologize.com)
20 points
by
ddod
14y ago
|
19 comments
115.
▲
Show HN: Notebooth, an async Omegle-type thing
(notebooth.com)
1 points
by
ddod
14y ago
|
0 comments
116.
▲
by
ddod
14y ago
Instead of learning SPARQL, I wrote my own script to parse from their format into my own. I should probably devote some time to learning SPARQL, but dbpedia's documentation was really confusing for me and I was basically learning it all jus
117.
▲
by
ddod
14y ago
Their stuff really isn't very user friendly (or at least wasn't for me) so there's a lot of room for improvement. I'm sure if you made a simple (and useful) Wikipedia API, you would be loved by all.
118.
▲
by
ddod
14y ago
Thanks! https://github.com/benwasser/whodiedhere
119.
▲
by
ddod
14y ago
I just put it up on Github now: https://github.com/benwasser/whodiedhere After I built this I played around with some other information in dbpedia to see if I could get anything juicy out of it. No luck so far, but I'm sure there's a lot
120.
▲
by
ddod
14y ago
I had to use a couple databases from dbpedia to get all the information (the death location syntax wasn't uniform) to make my own more manageable database. As for the server, it's just Node.js and Socket.IO
More ›