On your suggestion, I read through http://www.chessvibes.com/plaatjes/rybkaevidence/RYBKA_FRUIT_Mar11.pdf http://www.chessvibes.com/plaatjes/rybkaevidence/RYBKA_FRUIT... .
It makes the clear and cogent statement:
> While a large (indeed, almost complete) match is found, it is presumably feasible to opine that the Fruit source code can be taken as a “manual” for chess programming (perhaps in the sense of a modern version of How Computers Play Chess), and if this paradigmatic view is accepted, then the re-use of the same evaluation components might arguably be less derelict.
That theme occurs elsewhere:
> This Fruit/Rybka overlap would already likely meet a “plagiarism” standard, for instance as used in the detection of non-original work in academia and/or book publishing (note that plagiarism is generally an ethical standard and not a legal one). There is also the question of how important this item is from a chess-playing standpoint, perhaps again viewing Fruit as a “manual” in some sense.
The issue, which is also that mentioned in the chessboard article, is that the standard for plagiarism "in the context of computer chess (or more generally, computer boardgames)" is extremely sensitive. It is not the same standard practiced in research, programming, arts, or any other field I can think of. It's so high that it's not reasonable.
As Wikipedia writes, "the notion [of plagiarism] remains problematic with nebulous boundaries." In this PDF I read of multiple cases where the copying is not copyright infringement but one of reimplementing an algorithm. Relevant quotes are "Rybka 1.0 Beta uses bitboards, making direct code comparison ineffective", "Rybka uses a look-up table of patterns, while Fruit does bit-scanning", and "the relative scaling for each rank-based bonus in Rybka is essentially 10-30-60-100, though in units of 256 as in Fruit."
This PDF does stress that the surface differences are not the issue:
> I might stress that the fact that Fruit 2.1 visibly computes these while Rybka 1.0 Beta just has an array is not really relevant for the discussion here. The content is of more import.
where I presume the context includes "everything must have independent origin."
The PST structures are similar, although it uses different weights in parts. Note also the scaling differences between the two code bases - the similarities are in normalized space. Hence, this is again not evidence of copyright infringement. It would not be plagiarism in the scientific research fields I work in; since "influenced by the work of XYZ" would suffice.
This PDF points out that the quad() function in Rybka uses a different scaling, rank-dependent values, and more cases than Fruit.
This is again not a case where "parts of the code of Rybka were copied verbatim from Fruit" but where the approach from Fruit was modified. The PDF author then says:
> Not all of these terms have exactly the same meaning in Rybka 1.0 Beta, and discussing any differences would diverge from my focus on the re-use of the quad() function. Perhaps the main difference is with FreePasser, as to whether the pawn’s path is met by a friendly or enemy piece, which uses SEE in Fruit and “attacks” bitboards in Rybka, and further is split into 3 parts in Rybka.
> As with the PST comparison, it seems that there is a structural similarity between Fruit and Rybka, and the question of “originality” therein allows multiple approaches.
I think these statements are enough to establish that there was no copyright violation for this section. Your question "Does any of that excuse either violating the GPL" is therefore not relevant - there does not appear to be a copyright violation.
Could you clarify what you mean by "violating the GPL"? Does you refer to things like the file parsing code, which shows idiomatic similarities between Rybka and Fruit?
Over and over again I see that the issue is not outright copyright infringement of the chess engine nor lack of attribution, but that the definition of "plagiarism" as used in chess competition is extremely sensitive; sensitive enough that "structural similarity" even with attribution is considered excessive. It's much more stringent than any other field I can think of. As used here, it has lost its moral meaning and become more of a technical term.
Speaking as a complete outsider, it appears that the Rybka code base went through several iterations where it was based on ideas in different, existing programs. This is not uncommon, and is both legal and moral. My Minix example is quite relevant; it's meant to be used as a reference for understanding operating systems, which means people who use it as a reference will tend to create similar OSes. A question (rightly pointed out earlier) is, do open source chess programs serve as a similar manual?
The Rybka author was sloppy though. He didn't use version control until very late, and he followed too closely some of the more boring parts, like file parsing. There may be copyright infringement, and the remedy under the GPL is to request that the author either apply the GPL to the entire program, or remove the infringing parts. That this hasn't happened (I'm only guessing that it hasn't) tells me that the copyright holders aren't concerned enough to ask for help from the SFLC or other organizations which help enforce the GPL.
However, the core part shows signs of creative thought and improvement, which means it was not copied verbatim from Fruit. ("Creative" here in the legal sense related to copyright law.)
That's why Rybka still exists as a commercial program. But the chess competition arena has a different criteria for originality. While they use the term "plagiarism", they do so with a different meaning than used by nearly the rest of the world.
The entire point of the chessbase essay was to stress that this "originality criterion" is increasingly at odds with how software, including chess programs, are developed. Not only does it need to be changed, but it should have been changed years ago ("updating WCCC Rule 2 to reflect contemporary reality would be a years-overdue positive step"). At the end of that essay the author quotes:
> A fair group of participating programmers present have expressed they want the rules to be updated. One line of thinking is that attribution plus added value should be sufficient to compete, instead of 100% originality.
You say the documents against Rybka are convincing. I have read a couple of them now, and I am convinced that Rybka is in technical violation of WCCC Rule 2. I am not convinced that it's plagiarism. For that I would want to see lack of evidence of attribution, which is hard given that there is attribution. Nor am I convinced that there's wholesale copyright violation. For that I would want to see large spans of code which are not just functionally identical but which use the same values, same implementation, and same function call order. Here too the strongest evidence shows "structural similarity" but definitely not copyright infringement.
What would it take for you to be convinced otherwise? What was not persuasive in the chessbase article?
Thanks for taking the time to put together such a good argument. I'm afraid that I won't have time to write another reply like this, so if it's not at all convincing, we'll just need to agree to disagree.
The first point I'd make is that from the point of view of the investigators, what mattered were not copyright issues but a possible tournament rules violation. So they'd certainly not want to muddle the issue with issues of whether something was copied over verbatim or transcribed. On the other hand my interest is more in the GPL exploitation, since that actually matters outside the insignificant scope of computer chess politics.
It should be absolutely clear e.g. from the comparisons between pre-1.0 Rybka and Crafty that there was verbatim code copying going on. And not only in things like parsing code, but in actual game playing code. There is no other reasonable explanation for having exactly the same dead code around in exactly the same places (for example the double-zeroing bugs, comparisons to funny magic numbers that could never be true).
Also in places where arbitrary decisions needed to be made in the code, they were done exactly the same way as in Crafty (e.g. the numbering of pieces, the ordering of operations during evaluation). Now, this might not be proof of those parts of the code having been copied. It would be totally reasonable to argue that the author, having read the original source code, would naturally make the same arbitrary decisions.
But of course nobody cares about the code of that version of Rybka. It's mainly useful as context for what happened after Fruit was released, followed by a new version of Rybka.
First, there is again evidence of object code that exactly matches that of certain parts of Fruit. For example the command parsing, the decision of when to stop searching, or what to do when a result has been found. These are not as strong evidence of verbatim copying though as for the earlier copying from Crafty, since this code is at best idiosyncratic rather than clearly buggy or useless.
Likewise all of the earlier arbitrary decisions start to be made differently. Piece numbering changes from what was used in Crafty previously what Fruit uses. The main evaluation routine stops doing things in exactly the same order as Crafty did them, and starts doing them in exactly the same order as Fruit did them.
It's not really any longer a reasonable defense that this is just how he'd naturally do things after having read the source code and seen an example. Clearly he already had intimate knowledge with another source code base with different conventions. And even if this was just a matter of being inspired by the Fruit code why would he even be
rewriting all of this non-essential code rather than adding these concepts to his existing codebase.
From a copyright / GPL point of view, I think the argument for copying is fairly strong already at this point, and the question of how large the rewrites to the other code were is irrelevant. From a tournament rules viewpoint, it's the opposite.
So why isn't there equally strong evidence for verbatim copying in the game playing code as in the earlier Crafty case? Because the underlying board representation also changed from the representation used by Fruity to that used by Crafty, making establishing 1:1 correspondence between source and object code harder. So of course we can't reliably tell what kind of process produced this new code, e.g:
(a) modifying the Fruit code in-place
(b) using the Fruit code as a constant guide when writing a new version using a different data structure
(c) reading the Fruit code and then at some later point writing entirely new code
(d) at some point modifying his old code to use the concepts learned from Fruit
From a copyright / GPL point of view I don't think there's much difference between (a) and (b), but I could be wrong. In either case it doesn't seem like a very creative endeavor. Case (d) should be acceptable to anyone, but seems like a remote possibility at best, since the engine lost a number of game playing features of Crafty exactly at the same time as gaining a set matching those of Fruit very well. From a copyright point of view case (c) seems totally fine to me, but from a chess playing creativity point less so. Further, given the established flagrant pattern of copying, it does seem a lot more likely that he would have been taking the expedient road out in this particular instance as well.
I did not find the chessbase article very persuasive due to a few things.
First of all, it was framing the situation of a witch-hunt where a group of jealous sub-peers saw that their only chance of ever being successful again would be to destroy the superior competition by legal means. To an outsider this kind of naked emotional appeal seems very suspicious. Further, trying to suggest that e.g. Ken Thompson would have been motivated by these ulterior motives is just ridiculous.
Second, it's making the argument that really plagiarism is just the accepted custom of computer chess, or at least should be, and that this is therefore selective application of the rules. Knowing whether this is true would require much more intimate knowledge of the computer chess politics than I have any interest in acquiring. However, some of the evidence such as the evaluation feature comparison is certainly suggestive that the level of copying was much more significant in this case than was the norm. Certainly for this verdict to be fair, the reverse engineered and slightly tweaked versions of Rybka should not be allowed to compete in the tournament either.
Third, it tries to establish non-relatedness by the decision dendogram. That's clearly a fallacious argument. A graph like this could show evidence of copying, assuming they were using a sufficiently large amount of positions as input (I presume they did). But it can't possibly be used as evidence of non-copying. That's because a few small changes could easily have an effect on the evaluation results of a large number of positions. Sure, making those small changes are a creative act. But that doesn't mean that the combined work is any less derivative of the original.
Fourth, when discussing the actual evidence, it's considering every bit of evidence totally isolated. "Oh, but it could have happened like this in a totally innocent way". But that's clearly nonsense, you have to consider the totality of the evidence. At some point there are too many coincidences to explain away.
What would convince me to change my mind? It depends on what exactly. I don't think that anything could convince me otherwise of the copyright violations, beyond somebody showing that the investigators lied, and that reported similarities don't exist at all. Of the copyright violations extending to the core evaluation routines? The missing source code would be the best proof, since it could show whether the similarities extended to things like ordering and naming of functions. Of Rybka being unfairly singled out, since everybody was doing exactly the same copying from Fruit? Somebody would need to show that this really happened, and that the other rules violators are getting a free pass.