Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jdrock
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
jdrock
14y ago
Austin, TX Client Project Developer Datafiniti is the world's first search engine for data. Datafiniti's search results are complete data sets taken and aggregated from the web. You can search for information on places, people, products a
32.
▲
by
jdrock
14y ago
No - inbound data transfer cost to AWS is free.
33.
▲
by
jdrock
14y ago
Being the CEO of a firm that offers web-crawling services, I found this post very interesting. On 80legs, the cost for a similar crawl would be $500, so it's nice to know we're competitive on cost.
34.
▲
Does Big Data Have a Story Beyond Geeks?
(siliconangle.com)
1 points
by
jdrock
15y ago
|
0 comments
35.
▲
by
jdrock
15y ago
That's not an accurate impression of SXSW. There are many companies and folks there doing fascinating things and solving hard problems. If you just go the panels and hang out at the convention center, you're going to miss out. If you hav
36.
▲
by
jdrock
15y ago
Houston/Austin TX - Full Time - http://www.datafiniti.net Datafiniti is a small startup building the first search engine for data. We crawl the entire web, collecting data on businesses, places, people and things and then normalize all o
37.
▲
Datafiniti builds a search engine for data
(gigaom.com)
1 points
by
jdrock
15y ago
|
0 comments
38.
▲
Datafiniti: The Search Engine for Data (public beta)
(datafiniti.net)
6 points
by
jdrock
15y ago
|
0 comments
39.
▲
Convoluted TOS and "Open" APIs Will Be the Death of Us
(readwriteweb.com)
15 points
by
jdrock
15y ago
|
3 comments
40.
▲
by
jdrock
15y ago
Houston, Texas (H1B accepted) 80legs is building a next-generation web data platform - a service that will allow anyone to run SQL-like queries on all data available from the web. We're looking for talented folks to help us :) Positions av
41.
▲
by
jdrock
15y ago
Not necessarily more traffic, though that's possible. More likely it was the fact that it was coming from multiple IP addresses.
42.
▲
by
jdrock
15y ago
True story: We crawled them a while back (before they expanded their engineering team) and because of our distributed system, "alarms" were going off. Rather then take time to tell their system we were not a DDOS attack, they put us in rob
43.
▲
by
jdrock
15y ago
Our default crawler obeys robots.txt and it looks like the /results URLs are not allowed. However.. I think you could start from a URL like http://www.youtube.com/watch?v=sAzhOSbxMiI and then crawl to the linked videos from there...
44.
▲
by
jdrock
15y ago
Er.. not true, the contest plan (which is different from the free plan) allows up to 10 MB download. Registered contestants should have access to their plans within an hour of registering. And if you don't have it, just contact us: http:/
45.
▲
by
jdrock
15y ago
Specific documentation for 80legs is available at http://wiki.80legs.com . To get fancy, you may want to check out http://wiki.80legs.com/80apps to make your own custom extractors. Note that any custom extractors you write will output
46.
▲
by
jdrock
15y ago
Hm.. can you try again? Seems to be working from our end!
47.
▲
by
jdrock
15y ago
Am I the only one that feels that this will impede the adoption of NFC-style payments?
48.
▲
80legs is hiring a data sales rep
(80legs.com)
2 points
by
jdrock
16y ago
|
0 comments
49.
▲
by
jdrock
16y ago
I'd recommend: 1. Blocking all the IP addresses you've found. You can also block the entire Amazon IP range. No real users will be coming from there. 2. Contacting Amazon and sending them your logs and evidence. 3. Contact competitor and
50.
▲
by
jdrock
16y ago
Shion from 80legs here. Please try contacting us again and I'll make sure that we adjust our crawl rate file accordingly for your domain(s) or IP(s).
51.
▲
5 Paths To The New Data Integration
(informationweek.com)
3 points
by
jdrock
16y ago
|
0 comments
52.
▲
by
jdrock
16y ago
What browser/OS are you using?
53.
▲
Extractiv Opens Semantic Web Crawling On-Demand
(siliconangle.com)
1 points
by
jdrock
16y ago
|
0 comments
54.
▲
by
jdrock
16y ago
Some direct links for more information: Live Demo: http://www.extractiv.com/demo.html Documentation: http://wiki.extractiv.com
55.
▲
We just launched Extractiv, our Semantic Web as a Service
(extractiv.com)
17 points
by
jdrock
16y ago
|
4 comments
56.
▲
by
jdrock
16y ago
I encourage anyone to use Extractiv ( http://www.extractiv.com ) for this challenge. We're rolling out an On-Demand (local document analysis similar to Calais) Semantic Conversion service that could be helpful.
57.
▲
by
jdrock
16y ago
I'm curious to see if this guy will get a letter from Facebook's legal department, like Pete Warden did.
58.
▲
by
jdrock
16y ago
So.. gaming the system equates to studying well and working hard? That's.. not what's usually meant by "gaming"..
59.
▲
by
jdrock
16y ago
My hope in making a code of conduct accepted by legitimate web crawling companies is that it lets customers and users more easily decide which products to use. It also helps in distancing legitimate companies from illicit ones.
60.
▲
Is it time for a web crawling code of conduct?
(readwriteweb.com)
4 points
by
jdrock
16y ago
|
2 comments
More ›