Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
nikic
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
14 ms
·
91.
▲
by
nikic
11y ago
I think the most interesting bit here is this: > Now, the more astute reader will point out that I just sent over 4 gigabytes of data over the internet; and that this can’t really be all that interesting - but that argument is readily co
92.
▲
Internal value representation in PHP 7, part 1
(nikic.github.io)
65 points
by
nikic
11y ago
|
11 comments
93.
▲
by
nikic
12y ago
The maximum load factor is 1. A lower load factor only makes sense if open addressing is used.
94.
▲
by
nikic
12y ago
Yes, pData and pDataPtr are no more. They have always been pretty pointless - both could have been dropped even retaining the rest of the previous implementation by using a struct hack layout. In the new implementation they aren't need
95.
▲
by
nikic
12y ago
I have no idea about engineering, but I don't see how a physicist would not know this. Basis transformations in infinitely dimensional Hilbert spaces are an integral part of quantum mechanics. The Fourier-transform in particular is use
96.
▲
by
nikic
12y ago
Ruby only needs symbols because it made the questionable choice of using mutable strings with by-object semantics. PHP does not have this problem. It will automatically intern literal strings, thus giving you the performance and memory usag
97.
▲
by
nikic
12y ago
The lifetime constraint in the enum is mandatory, the code won't compile without it. The tokenization function itself doesn't have to explicitly specify lifetimes like it does in the blog post. It would also work without them and
98.
▲
by
nikic
12y ago
Common mistakes: Testing with xdebug enabled in one version or testing with opcache enabled in only one version.
99.
▲
by
nikic
12y ago
I feel like you're giving a very one-sided representation of the issue. There were some very real concerns about the memory usage of the proposed modifications to phpng (you can see http://news.php.net/php.internals
100.
▲
by
nikic
13y ago
PHP has a rather strict no-BC-break policy for minor releases. As such the 5.x series improves the language mostly through additions. I think that's good. PHP has been missing a number of things that are now present. However that doesn
101.
▲
by
nikic
13y ago
While the gist behind this post is right, the article contains a number of factual errors. E.g. the version numbers in the "language features" section are pretty messed up: * Namespaces, closures and FPM were already in PHP 5.3 *
102.
▲
by
nikic
13y ago
It's not like you're writing the regular expressions and lookup tables per hand. You write the routes in whatever format you deem nice and it's automatically translated from there ;)
103.
▲
by
nikic
13y ago
You can use a lexer, yes - but any fast lexer implementation in a high-level language will actually use the very same technique described here. Generating a character-comparison + goto based FSM works great in C, but has bad performance in
104.
▲
by
nikic
13y ago
You won't hit catastrophic backtracking problems with the types of regular expressions you need for routing. Even if you do use some very weird expression, the worst that can happen is that you hit the backtracking limit and a route fa
105.
▲
by
nikic
13y ago
Did you read on to the end? I also had the problem that performance degraded with large route sets, but resolved it by matching in chunks of ten routes.
106.
▲
by
nikic
13y ago
Not sure I understand this comment. Your code snippet uses named subpatterns, but not within a (?| group. The interesting question is whether you can combined those two. Without (?| you can use named captures in PCRE as well (there's j
107.
▲
by
nikic
13y ago
PCRE actually provides this information. There's a feature called backtracking control verbs, and one of those is (* MARK:NAME). You could put a (* MARK:A), (* MARK:B), etc at the end of every path and PCRE would tell you which of the
108.
▲
Fast request routing using regular expressions
(nikic.github.io)
58 points
by
nikic
13y ago
|
41 comments
109.
▲
by
nikic
13y ago
To make sure you have proper context: With one or two exceptions, those proposals are just vague ideas of what we might want to do in PHP 6. It says very little about what will actually go in. Most of it won't, at least I dearly hope s
110.
▲
by
nikic
13y ago
It's a common misconception. Many people don't understand that the normal string functions are perfectly safe on UTF-8, as long as you don't use hardcoded lengths or offsets. I.e. substr($str, 0, 50) is not safe due to the
111.
▲
by
nikic
13y ago
Some languages make writing secure code easier than others. When it comes to web-related code of semi-good programmers, I'd conjecture that the amount of vulnerabilities is directly proportional to the amount of magic involved. Early P
112.
▲
by
nikic
13y ago
Please, if you do performance benchmarks, always run them with an opcode cache (PHP 5.5 + opcache would be the obvious choice). Without it your benchmark becomes completely meaningless. I also feel like HHVM devs should be aware of this
113.
▲
by
nikic
13y ago
Chances are good that there will be a PHP 6 release in the near future (~3 years?) that contains a non-trivial amount of BC breakage. At least the idea has gained some momentum recently. But no branch / timeline yet ;)
114.
▲
by
nikic
13y ago
Imho the main thing C++ offers are cheap abstractions. In C++ you can easily build abstractions that do not have any runtime overhead at all. You can write beautiful and general code and still not sacrifice a single cycle for it. Go doesn&#
115.
▲
by
nikic
13y ago
Not disagreeing there. Fast NFA implementations are very nice for matching the regular subset and falling back to a more general algorithm for the non-regular cases :)
116.
▲
by
nikic
13y ago
> They can't be. That's not exactly right. What "better" or "worse" here refer to is how much the advantage the attacker with specialized hardware has over you running on a standard server. This advantage is
117.
▲
by
nikic
13y ago
Obligatory comment: Nowadays nearly nothing uses actually "regular" regexes, which is also the reason why regex engines are typically using backtracking and not Thompson NFAs. For general-purpose applications (e.g. use in programm
118.
▲
by
nikic
13y ago
> Contrast this to PHP, where /e appears to be on by default and there's no mention of it whatsoever in the documentation, except in the change log to say it's deprecated in 5.5. The preg_replace documentation says "S
119.
▲
by
nikic
13y ago
Yes, the fact that 0 is returned ~1% more often than other digits is very significant. E.g. if the generated random number stream were used for XOR encryption, then you could just collect a large number of encryptions of the same text, for
120.
▲
by
nikic
13y ago
MT is not suitable for crypto because after observing 637 values you can predict all further values. 637 is the size of MTs internal state vector and the output of MT are basically just values from that state vector run through a tempering
More ›