6 ms·
This was announced originally early last year. It removes the requirement for TLD and nTLD (not ccTLD) operators to have a WHOIS service available, but doesn't
by hughesey 2y ago
This was announced originally early last year. It removes the requirement for TLD and nTLD (not ccTLD) operators to have a WHOIS service available, but doesn't mandate they must shut them down.
So far the sunsetting has had little effect with most TLDs still having their WHOIS services online. In reality, I think we'll see a period of time where many TLDs and nTLDs have both WHOIS and RDAP available.
Additionally, since ccTLD's aren't governed by ICANN, many don't even have an RDAP service available. As such, there's going to be a mix of RDAP and WHOIS in use across the entire internet for some time to come.
Disclosure: I run https://viewdns.info/ https://viewdns.info/ and have spent many an hour dealing with both WHOIS and RDAP parsing to make sure that our service returns consistent data (via our web interface and API) regardless of the protocol in use.
- jbverschoor 2y agoIt's funny to see that a lot of services are finally moving from a human-readable / plain text format towards structured protocols right at the point where we can finally have LLMS parse the unstructured protocols :-)
- francislavoie 2y agoBut isn't using LLM for that really expensive? Seems wasteful.
- genewitch 2y agoMy desktop GPU can run small models at 185 tokens a second. Larger models with speculative decoding: 50t/s. With a small, finetuned model as the draft model, no, this won't take much power at all to run inference. Training, sure, but that's buy once cry once. Whether this means it's a good idea, I don't think so, but the energy usage for parsing isn't why.
- smokel 2y agoA simple text parser would probably be 10,000,000 times as fast. So the statement that this won't take much power at all, is a bit of an overstatement.
- Hercuros 2y agoIt’s not just about the energy usage, but also purchase cost of the GPUs and opportunity cost of not using those GPUs for something more valuable (after you have bought them). Especially if you’re doing this at large scale and not just on a single desktop machine. Of course you were already saying it’s not a good idea, but I think the above definitely plays a role at scale as well.
- SlightlyLeftPad 2y agoYou’re right, I could be trying to get Crysis to run at 120 fps.
- tmtvl 2y agoIf you have spare GPU time you could donate it to projects like Folding@Home.
- boxed 2y ago50 tokens per second. Compared to a quick and dirty parser written in python or even a regex? That's going to be many many orders of magnitude slower+costlier.
- anthk 2y agoawk would run millions times faster, not to mention mawk and awka.
- berkes 2y agoIn order to make the point that > energy usage for parsing isn't why You'll need to provide actual figures and benchmark these against an actual parser. I've written parsers for larger-scale server stuff. And while I too don't have these benchmarks available, I'll dare to wager quite a lot that a dedicated parser for almost anything will outperform an LLM magnitudes. I won't be suprised if a parser written in rust uses upwards of 10k times less energy than the most efficient LLM setup today. Hell, even a sed/awk/bash monstrosity probably outperforms such an LLM hundreds of times, energy wise.
- chgs 2y agoHow many times would you need to parse to get an energy saving on using an lm to parse vs using an llm to write a parser, then using the parser to parse.
- GTP 2y ago> using an llm to write a parser You're assuming OP needs an LLM to write a parser, since they mentions writing many during their career they probably don't need it ;)
- berkes 2y agoI didn't use an LLM back then. But would totally do that today (copilot). Especially since the parser(s) I wrote were rather straightforward finite state machines with stream handling in front, parallel/async tooling around it, and at the core business logic (domain). Streaming, job/thread/mutex management, FSM are all solved and clear. And I'm convinced an LLM like copilot is very good at writing code for things that have been solved. The LLM, however, would get very much in the way in the domain/business layer. Because it hasn't got the statistical body of examples to handle our case. (Parsers I wrote were a.o.: IBAN, gps-trails, user-defined-calculations (simple math formulas), and a DSL to describe hierarchies. I wrote them in Ruby, PHP, rust and perl.)
- chgs 2y agoI was thinking more of when a sufficiently advanced device would be able to “decide” the task would be worth using its own capabilities to write some code to tackle the problem rather than brute force. For small problems it’s not worthwhile, for large problems it is. It’s similar to choosing to manually do something vs automate it.
- anthk 2y agoMy Atom n270 netbook with mawk and a few lines parsing the files with a simple regex will crush down your GPU+LLM's on both time and power usage.
- genmon 2y agoMy assumption is that models are getting cheaper, fast. So you can build now with OpenAI/Anthropic/etc and swap it out for a local or hosted model in a year. This doesn't work for all use cases but data extraction is pretty safe. Treat it like a database query -- a slow but high availability and relatively cheap call.
- szundi 2y ago[dead]
- Cthulhu_ 2y agoWhile it will become cheaper, it will never be as fast / efficient as 'just' parsing the data the old-fashioned way. It feels like using AI to do computing things instead of writing code is just like when we moved to relatively inefficient web technology for front-ends, where we needed beefier systems to get the same performance as we used to have, or when cloud computing became a thing and efficiency / speed became a factor of credit card limit instead of code efficiency. Call me a luddite but I think as software developers we should do better, reduce waste, embrace mechanical sympathy, etc. Using AI to generate some code is fine - it's just the next step in code generators that I've been using throughout all my career IMO. But using AI to do tasks that can also be done 1000x more efficiently, like parsing / processing data, is going in the wrong direction.
- relistan 2y agoI know this particular problem space well. AI is a reasonable solution. WHOIS records are intentionally made to be human readable and not be machine parseable without huge effort because so many people were scraping them. So the same registrar may return records in a huge range of text formats. You can write code to handle them all if you really want to, but if you are not doing it en masse, AI is going to probably be a cheaper solution. Example: https://github.com/weppos/whois https://github.com/weppos/whois is a very solid library for whois parsing but cannot handle all servers, as they say themselves. That has fifteen + years of work on it.
- 2y ago
- adrianmonk 2y agoI wouldn't use LLMs, but if I did, I would try to get the LLM to write parser code instead. If it can convert from one format to another, then it can generate test cases for the parser. Then hopefully it can use those to iterate on parser code until it passes the tests. In a sense, asking it to automate the work isn't as straightforward as asking it to do the work. But if the approach does pan out, it might be easier overall since it's probably easier to deploy generated code to production (than deploying LLMs).
- permo-w 2y agodeepseek API costs are quite literally pennies per million tokens
- ajnin 2y agoWell you can't really trust an LLM to give you reproducible output every time, you can't even trust it to be faithful to the input data, so that's nice to have a standard format now. And for like a millionth of the computing resources to parse it. Also Whois was barely human-readable, with the fields all over the place, missing or different from one registry to the other. A welcome change that should have come really sooner.
- axegon_ 2y agohttps://deviq.com/antipatterns/shiny-toy https://deviq.com/antipatterns/shiny-toy
- vrighter 2y agowe can't ever have LLMs reliably parse any form of data. You know what can parse it perfectly though? A parser. Which works perfectly, and consistently.
- TeMPOraL 2y agoOf course we can. Reliability is a spectrum, not a binary state. You can push it up however high you like, and stop somewhere between "we don't care about error rate this low" and "error rate is so low it's unlikely to show in practice". It's not like this is a new concept. There are plenty of algorithms we've been using for decades that are only statistically correct. A perfect example of this is efficient primality testing, which is probabilistic in nature[0], but you can easily make the probability of error as small as "unlikely to happen before heat death of the universe". -- [0] - https://en.wikipedia.org/wiki/Primality_test#Probabilistic_tests https://en.wikipedia.org/wiki/Primality_test#Probabilistic_t...
- kbolino 2y agoThere are two problems with this comparison. First, probabilistic prime generation has a mathematically proven lower bound that improves with iteration. There is no comparably robust tuning parameter with an LLM. You can use a different model, you can use a bigger variant of the same model, etc., but these all have empirically determined and contextually sensitive reliability levels that are not otherwise tunable. Second, the prime generation function will always give you an integer, and never an apple, or a bicycle, or a phantasm. LLMs regurgitate and hallucinate, which means that a simple error rate is not the only metric that matters. One must also consider how egregiously wrong and even nonsensical the errors can be.
- dcow 2y agoThe general point is not that the feature currently exists to dial down the LLM parse error rate, it’s that the abstract argument “we can’t use LLMs because they aren't perfect” isn’t a realistic argument in the first place. You’re probably reading this on hardware that _probably_ shows you the correct text most all of the time but isn’t guaranteed to.
- _ache_ 2y agoIf you job is to be a referent, to have authority. You absolutely don't want to make any error. Pretty safe isn't enough, you need to be absolutely sure that you control the output. You only have one job, don't delegate authority.
- klysm 2y agoWhich world would you rather live in: * structured protocols that can be parsed by machines * unstructured protocols that are unreliably parsed by LLMs that require significant power and latency
- mmooss 2y agoIn addition to ~determative machines and LLMs, what about humans reading the data?
- tephra 2y agoI think RDAP is going to be adopted by more and more ccTLDs as well. WHOIS is not a particularly well liked protocol (I was at an IETF meeting where ICANN did a presentation on the timeline and people were literally cheering for the demise of WHOIS). Disclosure: Work in the ccTLD space.
- hughesey 2y ago100% agree that there will be more ccTLD operators that will implement RDAP. The sooner we're on a consistent protocol the better!
- dubbel 2y agoSelf-plug: I run a little mastodon/activity pub bot that monitors DNS RDAP adoption according to the official bootstrap file: https://social.haukeluebbers.de/@stateofrdap https://social.haukeluebbers.de/@stateofrdap Last post from yesterday: > As of today 82.25% (1187) of all 1443 Top Level Domains have an authoritative RDAP service declared. > These TLDs were added: > .ye
- sidd_sarkar 2y agookay
- tecleandor 2y agoIt's kind of funny some operators have never had it in practice. For example, .es never had a public whois, and need to register with a national ID (and I think with a fixed IP address) to get access to it.
- berkes 2y agoThat need for a national ID hasn't been in place for a long time, AFAIK. I have a .es (my nickname berkes, domain berk.es) for almost 16 years now, and live in the EU, but not in Spain. In the beginning I used a small company that offered services for non-spanish companies to register .es through them (I believe they technically owned the domains?). But today it's just in my local domain registrar without need for an ID. That .es has no whois has struck me as somewhat of a benefit actually. Back in the days, it kept away a lot of spam from spammers that'd just lift email-addresses off the whois. My .com, .nl and other domains recieve(d) significant more such spam. Let alone phone-number and other personal details delivered over an efficient, decentralized network. Though recent privacy addons(?) have mitigated that a little.
- tecleandor 2y agoI meant for accessing the whois, not for registering. If you try any type of WHOIS request you'll be replied with a message sending you to nic.es site, where you'll be presented with a captcha if you try to get information about a registered domain. It's not very well documented, but you can register at a government site using a national ID and they'll open WHOIS access for a fixed IP address, for a maximum of 10 queries a minute. [0] Context for any of you not used to the .es ccTLDs: Until some years ago, and simplifying a bit, if you wanted to register a .es TLD you had to be an Spanish national or company, and be the legal holder of the domain name you wanted to register (or your name and surnames). -- 0: https://sede.red.gob.es/es/procedimientos/solicitud-de-acceso-servicio-de-whois-por-el-puerto-43
- reaperducer 2y agoFor example, .es never had a public whois, and need to register with a national ID (and I think with a fixed IP address) to get access to it. Is this new? I had an .es domain around 2011, and am not Spanish, or even European.
- RealStickman_ 2y agoOff topic thank you for runnig viewdns.info. I don't use it regularly, mainly for the occasional WHOIS information lookup and it has always worked perfectly.
- hughesey 2y agoThanks for the kind words and glad it's been useful :).
- spurgu 2y agoHey, I've been looking for a tool that can do reverse NS lookup for a nameserver pairs (ie. which domains have nameservers ns1.example.com and ns2.example.com) but all the services out there that I've found can only do one. Is this something you would consider implementing?
- notRobot 2y agoThank you so much for running your service. I've used it for years, and LOVE how functional and useful it is!