9 ms·
Determination of the fifth Busy Beaver value
- ape4 1y agoI can't quite understand - did they use brute force?
- olmo23 1y agoYou can not rely on brute force alone to compute these numbers. They are uncomputable.
- PartiallyTyped 1y agoThey are at the very boundary of what is computable!
- arethuza 1y agoIsn't it rather that the Busy Beaver function is uncomputable, particular values can be calculated - although anything great than BB(5) is quite a challenge! https://scottaaronson.blog/?p=8972 https://scottaaronson.blog/?p=8972
- IsTom 1y ago> particular values can be calculated You need proofs of nontermination for machines that don't halt. This isn't possible to bruteforce.
- QuadmasterXLII 1y agoIf such proofs exist that are checkable by a proof assistant, you can brute force them by enumerating programs in the proof assistant (with a comically large runtime). Therefore there is some small N where any proof of BB(N) is not checkable with existing proof assistants. Essentially, this paper proves that BB(5) was brute forcible! The most naive algorithm is to use the assistant to check if each length 1 coq program can prove halting with computation limited to 1 second, then check each length 2 coq program running for 2 seconds, etc till the proofs in the arxiv paper are run for more than their runtime.
- schoen 1y agoIn this perspective, couldn't you equally say that all formalized mathematics has been brute forced, because you found working programs that prove your desired results and that are short enough that a human being could actually discover and use them? ... even though your actual method of discovering the programs in question was usually not purely exhaustive search (though it may have included some significant computer search components). More precisely, we could say that if mathematicians are working in a formal system, they can't find any results that a computer with "sufficiently large" memory and runtime couldn't also find. Yet currently, human mathematicians are often more computationally efficient in practice than computer mathematicians, and the human mathematicians often find results that bounded computer mathematicians can't. This could very well change in the future! Like it was somewhat clear in principle that a naive tree search algorithm in chess should be able to beat any human player, given "sufficiently large" memory and runtime (e.g. to exhaustively check 30 or 40 moves ahead or something). However, real humans were at least occasionally able to beat top computer programs at chess until about 2005. (This analogy isn't perfect because proof correctness or incorrectness within a formal system is clear, while relative strength in chess is hard to be absolutely sure of.)
- QuadmasterXLII 1y agoNot quite. There is some N for which you can’t prove BB(N) is correct for any existing proof assistant, but you can prove that BB(N) by writing a new proof assistant. However, the problem “check if new sufficiently powerful proof assistant is correct” is not decidable.
- IsTom 1y agoLength of proof / machine for proving it can be much bigger than BB(n) itself. Or even there can be specific machines that don't halt, but there is no proof for this at all – you can encode problems independent of ZFC in a few hundred states.
- gus_massa 1y agoYou can try them with simple short loops detectors, or perhaps with the "turtle and hare" method. If I do that and a friend ask me how I solved it, I'd call that "bruteforce". They solved a lot of the machines with something like that, and some with more advanced methods, and "13 Sporadic Machines" that don't halt were solved with a hand coded proof.
- CaptainNegative 1y agoOnly finitely many values of BB can be mathematically determined. Once your Turing Machines become expressive enough to encode your (presumably consistent) proof system, they can begin encoding nonsense of the form "I will halt only after I manage to derive a proof that I won't ever halt", which means that their halting status (and the corresponding Busy Beaver value) fundamentally cannot be proven.
- karmakurtisaani 1y agoOnce you can express Collatz conjecture, you're already in the deep end.
- jerf 1y agoYes, but as far as I know, nobody has shown that the Collatz conjecture is anything other than a really hard problem. It isn't terribly difficult to mathematically imagine that perhaps the Collatz problem space considered generally encodes Turing complete computations in some mathematically meaningful way (even when we don't explicitly construct them to be "computational"), but as far as I know that is complete conjecture. I have to imagine some non-trivial mathematical time has been spent on that conjecture, too, so that is itself a hard problem. But there is also definitely a place where your axiom systems become self-referential in the Busy Beaver and that is a qualitative change on its own. Aaronson and some of his students have put an upper bound on it, but the only question is exactly how loose it is, rather than whether or not it is loose. The upper bound is in the hundreds, but at [1] in the 2nd-to-last paragraph Scott Aaronson expresses his opinion that the true boundary could be as low as 7, 8, or 9, rather than hundreds. [1]: https://scottaaronson.blog/?p=8972 https://scottaaronson.blog/?p=8972
- adgjlsfhk1 1y agohttps://people.cs.uchicago.edu/~simon/RES/collatz.pdf https://people.cs.uchicago.edu/~simon/RES/collatz.pdf Generalized colatz is uncomputable.
- Veserv 1y ago
- karmakaze 1y agoThe busy beaver numbers form an uncomputable sequence. For BB(5) the proof of its value is an indirect computation. The verification process involved both computation (running many machines) and proofs (showing others run forever or halt earlier). The exhaustiveness of crowdsourced proofs was a tour de force.
- arethuza 1y agoI think you have to exhaustively check each 5-state TM, but then for each one brute force will only help a bit - brute force can't tell you that a TM will run forever without stopping?
- bc569a80a344f9c 1y agoNot quite, I think this is the relevant part of the paper: > Structure of the proof. The proof of our main result, Theorem 1.1, is given in Section 6. The structure of the proof is as follows: machines are enumerated arborescently in Tree Normal Form (TNF) [9] – which drastically reduces the search space’s size: from 16,679,880,978,201 5-state machines to “only” 181,385,789; see Section 3. Each enumerated machine is fed through a pipeline of proof techniques, mostly consisting of deciders, which are algorithms trying to decide whether the machine halts or not. Because of the uncomputability of the halting problem, there is no universal decider and all the craft resides in creating deciders able to decide large families of machines in reasonable time. Almost all of our deciders are instances of an abstract interpretation framework that we call Closed Tape Language (CTL), which consists in approximating the set of configurations visited by a Turing machine with a more convenient superset, one that contains no halting configurations and is closed under Turing machine transitions (see Section 4.2). The S(5) pipeline is given in Table 3 – see Table 4 for S(2,4). All the deciders in this work were crafted by The bbchallenge Collaboration; see Section 4. In the case of 5-state machines, 13 Sporadic Machines were not solved by deciders and required individual proofs of nonhalting, see Section 5. So, they figured out how to massively reduce the search space, wrote some generic deciders that were able to prove whether large amounts of the remaining search spaces would halt or not, and then had to manually solve the remaining 13 machines that the generic deciders couldn't reason about.
- ufo 1y agoLast but not least, those deciders were implemented and verified in the Rocq proof assistant, so we know they are correct.
- lairv 1y agoWe know that they correctly implement their specification*
- meithecatte 1y agoNo, they are correct, because the deciders themselves are just a cog in the proof of the overall theorem. The specification of the deciders is not part of the TCB, so to speak.
- deleted 1y ago[deleted]
- Xcelerate 1y agoThere are a few concepts at play here. First you have to consider what can be proven given a particular theory of mathematics (presumably a consistent, recursively axiomatizable one). For any such theory, there is some finite N for which that theory cannot prove the exact value of BB(N). So with "infinite time", one could (in principle) enumerate all proofs and confirm successive Busy Beaver values only up to the point where the theory runs out of power. This is the Busy Beaver version of Gödel/Chaitin incompleteness. For BB(5), Peano Arithmetic suffices and RCA₀ likely does as well. Where do more powerful theories come from? That's a bit of a mystery, although there's certainly plenty of research on that (see Feferman's and Friedman's work). Second, you have to consider what's feasible in finite time. You can enumerate machines and also enumerate proofs, but any concrete strategy has limits. In the case of BB(5), the authors did not use naive brute force. They exhaustively enumerated the 5-state machines (after symmetry reductions), applied a collection of certified deciders to prove halting/non-halting behavior for almost all of them, and then provided manual proofs (also formalized) for some holdout machines.
- fedeb95 1y agowhat's most interesting to me about this research is that it is an online collaborative one. I wonder how many more project such as this there are, and if it could be more widespread, maybe as a platform.
- arethuza 1y agoThe BB Challenge site is really well structured: https://bbchallenge.org/13650583 https://bbchallenge.org/13650583
- marvinborner 1y agoIn recent years there has been a movement to collaborate on math proofs via blueprints (dependency graphs) in the Lean language, which seems related. For example: https://teorth.github.io/equational_theories/ https://teorth.github.io/equational_theories/ https://teorth.github.io/pfr/ https://teorth.github.io/pfr/
- fedeb95 1y agothanks, these are interesting indeed!
- longwave 1y agoThis comment reminded me to check whether https://www.distributed.net/ https://www.distributed.net/ was still in existence. I hadn't thought about the site for probably two decades, I ran the client for this back in the late 1990s back when they were cracking RC5-64, but they still appear to be going as a platform that could be used for this kind of thing.
- schoen 1y agoI was also excited about those projects and ran DESchall as well as distributed.net clients. Later on I was running the EFF Cooperative Computing Award (https://www.eff.org/awards/coop https://www.eff.org/awards/coop), as in administering the contest, not as in running software to search for solutions! The original cryptographic challenges like the DES challenge and the RSA challenges had a goal to demonstrate something about the strength of cryptosystems (roughly, that DES and, a fortiori, 40-bit "export" ciphers were pretty bad, and that RSA-1024 or RSA-2048 were pretty good). The EFF Cooperative Computing Award had a further goal -- from the 1990s -- to show that Internet collaboration is powerful and useful. Today I would say that all of these things have outlived their original goals, because the strength of DES, 40-bit ciphers, or RSA moduli are now relatively apparent; we can get better data about the cost of brute-force cryptanalytic attacks from the Bitcoin network hashrate (which obviously didn't exist at all in the 1990s), and the power and effectiveness of Internet collaboration, including among people who don't know each other offline and don't have any prior affiliation, has, um, been demonstrated very strongly over and over and over again. (It might be hard to appreciate nowadays how at one time some people dismissed the Internet as potentially not that important.) This Busy Beaver collaboration and Terence Tao's equational theories project (also cited in this paper) show that Internet collaboration among far-flung strangers for substantive mathematics research, not just brute force computation, is also a reality (specifically now including formalized, machine-checked proofs). There's still a phenomenon of "grid computing" (often with volunteer resources), working on a whole bunch of computational tasks: https://en.wikipedia.org/wiki/List_of_grid_computing_projects https://en.wikipedia.org/wiki/List_of_grid_computing_project... It's really just the specific "establish the empirical strength of cryptosystems" and "show that the Internet is useful and important" 1990s goals that are kind of done by this point. :-)
- nathanrf 1y agoHere's a high level overview for a programmer audience (I'm listed as an author but my contributions were fairly minor): [See specifics of the pipeline in Table 3 of the linked paper] * There are 181 million ish essentially different Turing Machines with 5 states, first these were enumerated exhaustively * Then, each machine was run for about 100 million steps. Of the 181 million, about 25% of them halt within this memany step, including the Champion, which ran for 47,176,870 steps before halting. * This leaves 140 million machines which run for a long time. So the question is: do those TMs run FOREVER, or have we just not run them long enough? The goal of the BB challenge project was to answer this question. There is no universal algorithm that works on all TMs, so instead a series of (semi-)deciders were built. Each decider takes a TM, and (based on some proven heuristic) classifies it as either "definitely runs forever" or "maybe halts". Four deciders ended up being used: * Loops: run the TM for a while, and if it re-enters a previously-seen configuration, it definitely has to loop forever. Around 90% of machines do this or halt, so covers most. 6.01 million TMs remain. * NGram CPS: abstractly simulates each TM, tracking a set of binary "n-grams" that are allowed to appear on each side. Computes an over-approximation of reachable states. If none of those abstract states enter the halting transition, then the original machine cannot halt either. Covers 6.005 million TMs. Around 7000 TMs remain. * RepWL: Attempts to derive counting rules that describe TM configurations. The NGram model can't "count", so this catches many machines whose patterns depend on parity. Covers 6557 TMs. There are only a few hundred TMs left. * FAR: Attempt to describe each TM's state as a regex/FSM. * WFAR: like FAR but adds weighted edges, which allows some non-regular languages (like matching parentheses) to be described * Sporadics: around 13 machines had complicated behavior that none of the previous deciders worked for. So hand-written proofs (later translated into Rocq) were written for these machines. All of the deciders were eventually written in Rocq, which means that they are coupled with a formally-verified proof that they actually work as intended ("if this function returns True, then the corresponding mathematical TM must actually halt"). Hence, all 5-states TMs have been formally classified as halting or non-halting. The longest running halter is therefore the champion- it was already suspected to be the champion, but this proves that there wasn't any longer-running 5-state TM.
- tsterin 1y agoIn Coq-BB5 machines are not run for 100M steps but directly thrown in the pipeline of deciders. Most halting machines are detected by Loops using low parameters (max 4,100 steps) and only 183 machines are simulated up to 47M steps to deduce halting. In legacy bbchallenge's seed database, machines were run for exactly 47,176,870 steps, and we were left with about 80M nonhalting candidates. Discrepancy comes from Coq-BB5 not excluding 0RB and other minor factors. Also, "There are 181 million ish essentially different Turing Machines with 5 states", it is important to note that we end up with this number only after the proof is completed, knowing this number is as hard as solving BB(5).
- DarkNova6 1y agoAs somebody not familiar with this field, is there any tangible benefit to this solution or is it purely academic?
- twosdai 1y agoI am also unfamiliar, but a naive guess I can make is helping solve other math proofs, and maybe better encryption for the public. Basically everytime I see some progress in math, it usually means encryption in some way is impacted.
- schoen 1y agoThe Busy Beaver problem is one of the most theoretical, one of the most purely mathematical, of all theoretical computer science questions. It is really about what is possible in a very abstract sense. Working on it did make people cleverer at analyzing the behavior of software, but it's not obvious that those skills or associated techniques will directly transfer to analyzing the behavior of much more complex software that does practical stuff. The programs that were being analyzed here are extraordinarily small by typical software developer standards. To be more specific, it seems conceivable to me that some of these methods could inspire deductive techniques that can be used in some way in proof assistants or optimizing compilers, in order to help ensure correctness of more complicated programs, but again that application isn't obvious or guaranteed. The people working on this collaboration would definitely have described themselves as doing theoretical computer science rather than applied computer science.
- arethuza 1y ago"extraordinarily small by typical software developer standards" That's the incredible thing about these TMs - they are amazingly small but still can exhibit very complex behaviours - it's like microscopy for programs.
- non_aligned 1y agoThe research into BB numbers is purely academic and unlikely to be more than that, unless some other part of mathematics turns out to be wrong. Our current understanding is that the numbers are essentially guaranteed to be useless. This particular proof is doubly-academic in the sense that the value was already known, this is just a way to make it easier to independently verify the result. It's a part of a broader movement to provide machine proofs for other stuff (e.g., Fermat's last theorem), which may be beneficial in some ways (e.g., identifying issues, making parts of proofs more "reusable").
- jsjsjxxnnd 1y agoBB(5) was discovered last year, so why is this paper coming out now? Is there a new development?
- nathanrf 1y agoThis is the (pre-print) paper produced from that project - i.e. it is a self-contained, complete description of the work completed so that it can be reviewed, and then cited. The underlying result has not changed, but the presentation of the details is now complete.
- ks2048 1y agoAn interesting thing about this is that they figured out a high-level description of what this Busy Beaver program is doing - it's computing a Collatz-like sequence until it terminates. I'm not sure if that is described in this paper, but I learned about it in this Scott Aaronson talk, https://www.youtube.com/watch?v=VplMHWSZf5c https://www.youtube.com/watch?v=VplMHWSZf5c e.g see the slide at 31:40.
- gowld 1y agoCollatz-like except on one critical aspect: it is known to terminate!
- bobbylarrybobby 1y agoInteresting, it seems like a possible contender for BB6 (“Antihydra”) also does something Collatz-like. Is Collatz just a good blueprint for constructing long, complex, finite sequences?
- matusp 1y agoI find it very interesting as well. One would think that there might be many problems with similar properties, but people have already discovered the one that surfaces in Turing machines as well. Does this suggest that there are not that many problems with this type of property? Or does it suggest that we stumbled upon the simplest case.
- onraglanroad 1y agoI was wondering if this had any implications for BB(6) but it seems that can't be written down exactly (not enough room in this universe) so BB(5) is the last one we'll see an exact value for.
- sapiogram 1y agoWikipedia give a lower bound for BB(6) of 2^2^2^2^2^2... repeated 33554432 times, that's definitely a fast-growing function!
- jojomodding 1y agoAnd it gets that figure from this paper and the authors of this paper continuing to work on BB(6).
- mostly_a_lurker 1y agoHmm, I think you might be a bit underestimating? My understanding of up-arrow notation is: 2 ^^^ 5 means 2 ^^ (2 ^^ (2 ^^ (2 ^^ 2))) (with five 2's) = 2 ^^ (2 ^^ (2 ^^ (2 ^ 2))) = 2 ^^ (2 ^^ (2 ^^ 4)) = 2 ^^ (2 ^^ 65536), since Wikipedia calculates 2 ^^ 4 = 65536. Now 2 ^^ 65536 means 2 ^ (2 ^ (2 ^ ...))) with 65536 two's. Call that number N (too large to really describe). Then 2 ^^ N is 2 ^ (2 ^ ( ...)) with N two's.
- bravura 1y agoQuestion: Is the computational cost of verifying the proof significantly less than the computational cost of creating the proof?
- crdrost 1y agoLike I haven't run Coq but I assume that in this particular instance, the answer is a rather trivial "yes"? You'd think about, creating this proof is a gigantic research slog and publishable result; verifying it is just looking at the notes you wrote down about what worked (this turing machine doesn't halt and here's the input and here's the repetitive sequence it generates, that turing machine always halts and here's the way we proved it...). But taken more generally, not for this specific instance, your question is actually a Millenium Prize problem worth a million dollars.
- LegionMammal978 1y agoIn principle, you could reduce the problem into running a few very costly programs, so that the proof is completely verified as soon as someone eats that cost. (Some 6-state machines are already looking like they'll take a galactic cost to prove, and others will take a cost just barely within reason.) But this definitely isn't the case for the BB(5) project.
- LegionMammal978 1y agoThe compiled proof does take a few hours to finish verifying on a typical laptop. But obviously the many experiments and enumerations along the way took much more total processing power.
- tsterin 1y ago45 minutes on 13 cores on a standard laptop :)
- Kranar 1y agoThis is true in general for every mathematical proof in ZFC (and even in more powerful theories). The decision problem "Given a formula F and an integer n, is there a ZFC proof of F of length <= n?" is NP complete, meaning that verifying the proof can be done in polynomial time while deriving the proof can require an exponential amount of time.
- TheAmazingRace 1y agoI can't help but immediately think of Busy Beaver stores from the tri-state (PA-OH-WV) area. https://busybeaver.com/ https://busybeaver.com/
- purplejacket 1y agoThis paper is delightfully readable.
- tsterin 1y agoBest compliment ever
- 1970-01-01 1y agoTL,DR; 47,176,870 :}
- fndhope10 1y agoLe trouble neurologique fonctionnel est fréquent en pratique neurologique. Une nouvelle approche du diagnostic positif de ce trouble met l’accent sur des schémas reconnaissables de symptômes et de signes réellement ressentis, qui présentent une variabilité au sein d’une même tâche et entre différentes tâches au fil du temps. Les facteurs de stress psychologiques sont des facteurs de risque courants du trouble neurologique fonctionnel, mais ils sont souvent absents. Quatre entités — les crises fonctionnelles, les troubles fonctionnels du mouvement, le vertige postural perceptuel persistant et le trouble cognitif fonctionnel — présentent des similarités sur le plan de l’étiologie et de la physiopathologie, et constituent des variantes d’un trouble à l’interface entre la neurologie et la psychiatrie. Ces quatre entités possèdent des caractéristiques distinctives et peuvent être diagnostiquées à l’aide d’études neurophysiologiques cliniques et d’autres biomarqueurs. La physiopathologie du trouble neurologique fonctionnel comprend une hyperactivité du système limbique, le développement d’un modèle interne de symptômes dans le cadre d’une approche de codage prédictif, ainsi qu’un dysfonctionnement des réseaux cérébraux qui confèrent au mouvement son caractère volontaire. Les données disponibles soutiennent une prise en charge multidisciplinaire adaptée, pouvant inclure des approches thérapeutiques physiques et psychologiques.
- IAmBroom 1y agoThat seems wildly off-topic. Êtes-vous sûr de poster dans le bon fil de discussion?
- casey2 1y agowhat program if any of the 5-state TMs produces the same halting configuration as BB(4) in the fewest steps? What about just the halting tape?