Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ziflex
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Test and wait on the availability of a remote resource
(github.com)
1 points
by
ziflex
7y ago
|
0 comments
2.
▲
by
ziflex
8y ago
Nope, just thought that ferrets are cool :)
3.
▲
by
ziflex
8y ago
Oh yes. The reason if this is that for now the language itself is DOM agnostic, it's just a port of an existing one. ( https://docs.arangodb.com/3.4/AQL/ ) . So, the entire DOM thing is implemented by standa
4.
▲
by
ziflex
8y ago
Great! ^_^
5.
▲
by
ziflex
8y ago
This package is more like a runtime. There are plans to create a dedicated server, where you would be able to store your queries, schedule them and set up output streams like Spark or Flink. For now, it does not respect robots.txt. But it c
6.
▲
by
ziflex
8y ago
And you can open as many pages as you want in a single query (or as your memory allows you :) )
7.
▲
by
ziflex
8y ago
That's what you can do right now :) https://github.com/MontFerret/ferret/blob/master/docs/exampl... Document, returned form DOCUMENT() function, represents an open browser tab which allows you
8.
▲
by
ziflex
8y ago
PRs are welcome :) There is gonna be a separate project within the organization that would do all these things and even more. It's just beginning :)
9.
▲
by
ziflex
8y ago
Yes! This is one of the reasons why I wanted to be able making these changes without redeploying the whole thing!
10.
▲
by
ziflex
8y ago
It can! :) Even more - it can interact with these pages! Here is an example of use of Google Search page: https://github.com/MontFerret/ferret/blob/master/docs/exampl...
11.
▲
by
ziflex
8y ago
Thank you very much for your valuable feedback and I'm glad that someone has finally got the idea :)
12.
▲
by
ziflex
8y ago
The idea is to create a high level abstraction that represents your web scraping logic. The project is still WIP. I will create a web server which will help you to store your queries, schedule them and set up output stream to other systems
13.
▲
by
ziflex
8y ago
You are fine, I totally understand your scepticism. And you are right, there are definitely issues in communication. First of all, I built it for myself. I needed a high level representation of scraping logic, which would run an isolated an
14.
▲
by
ziflex
8y ago
I could, if I knew Python pretty well :) But I've done it in the way I needed it to be done. I wanted to have an isolated and safe environment that would allow me to easily scrape the web without dealing with infrastructural code.
15.
▲
by
ziflex
8y ago
That's true. You can do it, of course. I'm not saying that this is the only way of doing it.
16.
▲
by
ziflex
8y ago
This is how it works under the hood. But everything is wired for you ;)
17.
▲
by
ziflex
8y ago
You definitely need to share. Web scraping is tedious. As more ideas we have, as more options we have to come up with a better solution for that.
18.
▲
by
ziflex
8y ago
That's true. The difference is how much efforts is needed to do that using API. What it brings is just a higher abstraction of that API which lets you easily to get work done.
19.
▲
by
ziflex
8y ago
Yes, puppeteer and ferret use same technology under the hood - Chrome DevTools Protocol. But, the purpose of the project is not to "be better than". No, the purpose is to let developers focus on the data itself, without dealing w
20.
▲
by
ziflex
8y ago
DOM is a representation of some data. Which means, you can extrapolate the data and then manipulate it. The language itself has nothing related to the DOM. All DOM operations are implemented via functions from standard library. "Good a
21.
▲
by
ziflex
8y ago
I'm sorry I do not fully understand what you mean. Imagine, that you need to grab some data from SoundCloud and, also imagine, they do not have a public API :) How would you do it without launching a browser?
22.
▲
by
ziflex
8y ago
The main purpose is to use scripts like SQL. Where you can write and modify your scripts for data without compilation. Plus, the project aims to simplify the process and hide technical complexity behind it. Moreover don't forget, that
23.
▲
by
ziflex
8y ago
Proxy support is not in place yet. I will add it in future releases. You are welcome to do a PR :)
24.
▲
by
ziflex
8y ago
Actually, it's not yet another language :) The syntax is taken from ArangoDB - AQL https://docs.arangodb.com/3.3/Manual/
25.
▲
by
ziflex
8y ago
One of the goals of the project is to hide technical details and complexities that follow modern web scraping, especially when you deal with dynamic web pages.
26.
▲
by
ziflex
8y ago
Well, I would say it's more declarative than imperative. There are few differences that make it less declarative - variables and ternary operator.
27.
▲
by
ziflex
8y ago
The main advantage of this over your approach with HtmlAgilityPack is that Ferret can handle dynamic web pages - those that are rendered with JS. And also, it can emulate user interactions. But anyway, thanks for your feedback :)
28.
▲
Ferret – Declarative web scraping
(github.com)
260 points
by
ziflex
8y ago
|
92 comments
29.
▲
Type-safe utility library for creating nested Immutable Records
(npmjs.com)
1 points
by
ziflex
8y ago
|
0 comments