3 ms·
The article would be helped by using a full regex for ipv4 addresses - the one it uses would match invalid numbers (999.999.999.5 for instance), but the proper
by msoucy 8y ago
The article would be helped by using a full regex for ipv4 addresses - the one it uses would match invalid numbers (999.999.999.5 for instance), but the proper one is more complex (and would probably make for a better example as a result)
Also I think there's something wrong with this blog's formatting, it appears to be replacing underscores with italics even within code samples.
- wild_preference 8y agoThough to be fair it's not obvious where to encode validation. You're saying it should be intrinsic in the parsing rules of ipv4. But there are reasons why you would want to validate in a separate pass. For example, the more errors you parse, then the better error messages and introspection you have (structured data) when you want to validate. It's a classic trade-off. You're right though, doesn't make a very good example.
- setr 8y agoI feel like thats fine: syntactically valid, semantically not; Ideally syntax vs semantics should be decoupled in most parsing (hence the AST)
- samatman 8y ago256 is either 1[0-9][0-9] or 2[0-5][0-6], for the three digit case; that which can be syntactically detected, should be.
- setr 8y agoIt seems to me that validating on the AST would be a lot cleaner/saner (if int(ipv4addr) > 256: err), and generally allow for much better error messages; especially since you'll be doing almost all your validations on the AST anyways. It would ofc also be more useful in future parsing too (similar structure, but not ipv4), but thats a much more minor and rarer benefit
- diggernet 8y agoExcept you've missed 2x7-2x9...
- samatman 8y agoAh, you're right, that was careless of me here's the real deal courtesy of the URI spec: IPv4address = dec-octet "." dec-octet "." dec-octet "." dec-octet dec-octet = DIGIT ; 0-9 / %x31-39 DIGIT ; 10-99 / "1" 2DIGIT ; 100-199 / "2" %x30-34 DIGIT ; 200-249 / "25" %x30-35 ; 250-255