11 ms·
Just one technical comment - that parser will break on many FASTA files. They can also include '.' & * characters ('.' is often output by sequence aligners to i
by coriny 12y ago
Just one technical comment - that parser will break on many FASTA files. They can also include '.' & * characters ('.' is often output by sequence aligners to indicate sections outside of the model, * is there for aa but not dna/rna and I have seen them in sequencer output for uncertain positions). I think this kind of edge case variance is possibly why the various Bio-* tools have not been very successful in all, you end up spending hours trying to work out why it's not working with your particular inputs.
On a personal viewpoint: "Take a look at the REPL dump below and see if you can work out where we are implicitly using $seq as a Str rather than Seq:" - this section pretty much sums up why I've always hated working with Other Peoples' Perl. To accurately follow someone else's perl requires either an encyclopedic knowledge of the language and it's edge cases or very heavy documentation and a set standard of coding practices. It's a shame that this looks true for Perl 6 as well.
Great for personal use, I guess. But by and large I'm glad to see more projects switch over to python.
- kamaal 12y ago>>this section pretty much sums up why I've always hated working with Other Peoples' Perl. I think that statement was really more of 'The remainder is left for the users as an assignment' kind of statement. >>To accurately follow someone else's perl requires either an encyclopedic knowledge of the language and it's edge cases or very heavy documentation and a set standard of coding practices. To accurately follow anything non-trivial written by anybody in any programming language will require you a good grip over the language and its corner cases. There is no magic here, if a language doesn't offer facilities to heavy lift a certain thing you will end up writing those yourself.
- coriny 12y ago>> To accurately follow anything non-trivial written by anybody in any programming language will require you a good grip over the language and its corner cases. There is no magic here, if a language doesn't offer facilities to heavy lift a certain thing you will end up writing those yourself. I'm going to have speak subjectively and relative to my own skills here. Without rigorous use of standards Perl 5 is a nightmare to read. With very clear and consistent style then it can be readable, but this has to be enforced extremely strictly for any code base beyond trivial. Perl's 100 ways to do things is not a positive in this context. Case in point, I was given a bit of perl two weeks ago by a bioinformatics lecturer. Did not use strict or warnings; variables weren't declared in scope; no subroutines, let alone Moose objects. And this is typical of ~90% of bioinformatics perl that gets written. What's more, it was a trivial script that was simply parsing some BLAST matches from CSV, checking for overlaps and assigning a type based on the match gene combination - and was completely indecipherable. It took me two days of refactoring to pick out all the details. And he managed to use a couple of magic symbols I'd never seen before. The first time I was given a bit of python, without learning the language at all, I was able to extend the script and get it do some new things (it was considerably more complex than the aforementioned perl script). And even when it is well written and documented - I did work with some very good perl programmers - it's tended to end up as very fragile and difficult to extend or even substantially redevelop due to the heavy use of context for determining behaviour (and the actual context can be difficult to determine). Unit testing had to be very extensive since it was difficult to isolate changes. I'm going to have to be precise here - I'm attacking Perl 5. It just seemed from that section that Perl 6 doesn't solve the major problem (difficult to deduce the context and work out the resulting effect) I had with Perl 5. I may be completely missing the mark and actually readability is much better. I really hope so. >> Perl 6 isn't production ready yet. But even a cursory glance over the features tells, Python isn't even in the same league of tools Perl 6 will be. Vapourware is vapourware. Perl 6 was meant to be production ready a decade ago. On the other hand, I'm interested in knowing what you think Perl 6 brings that will give it the edge over Python? Do you see it displacing Python as rapidly as Python replaced Perl 5 as the science "glue" language of choice? With the apparent rise of Julia as well, could it just be too late?
- kamaal 12y agoI'm not talking about Perl 5 per se. But I've worked with C, Java, Pig, Python, Perl and a bit of shell scripting so far. And I've never had a situation where I could simply glance over something and know all of it(This is for reasonably big applications). I'm currently working on a big Java application where logic is often deep buried below Factor classes, Interfaces, Enums and layers of classes simply passing control to level below. And when you find the code that does something finally, you will see it wrapped will kinds of exception handling code. Its simply too verbose, with pointless ceremonial code. In pig, any reasonably large workflow has had me hooked for hours working with stub data trying to figure out how data flows and comes out. In fact I see this nothing new to Perl. If you are working with any serious application. A good touch over the language plus experience with handling things with patience is a must. Python is no exception to this rule, neither is any language I know of. >>I may be completely missing the mark and actually readability is much better. I really hope so. It depends on how you define readability. I would say verbose code is equally unreadable especially if you work with large applications. If you simply don't like regular expression you have a different problem. In some cases use of regular expression is simply unavoidable. IMO Python works very well for applications that interact with standard data sources(Databases, API's, XML's). Once you get the data to some very easily ingestable form, and there is little work to do from there on, Python works just fine. But Perl works fine for other use cases, heavy lifting and other kinds of data work. Perl may not be the right language for all use cases, but neither is Python or Java. >>Vapourware is vapourware. Perl 6 was meant to be production ready a decade ago. No, I don't think they ever gave out a date when they wanted a finished version out. Besides Rakudo is feature ready already. They are just working out last bit of performance issues and good to have things. And plan to have it out by 2015: https://fosdem.org/2015/schedule/event/get_ready_to_party/ https://fosdem.org/2015/schedule/event/get_ready_to_party/ http://perl6.org/compilers/features http://perl6.org/compilers/features >>On the other hand, I'm interested in knowing what you think Perl 6 brings that will give it the edge over Python? Grammars, Macros, Concurrency, Traits, Types, Improved regexes, Better OO, Functional programming, User defined operators, Mutable grammar, Junctions, Currying, Continuations, Multiple VM backends to name a few. The full list is at : http://perlcabal.org/syn/ http://perlcabal.org/syn/ >>Do you see it displacing Python as rapidly as Python replaced Perl 5 as the science "glue" language of choice? Perl and Python both co exist and will continue to, albeit the web frameworks market is dominated by Python and Ruby at this point in time, Perl is still pretty big in a lot of places. This isn't Perl or Python, people use tools which work for them within their project limitations. But I guess people will use what they find useful. And that only individual projects can decide as to what works best for them. >>With the apparent rise of Julia as well, could it just be too late? Reading the Perl 6 spec will tell how different Perl 6 is with regards to everything else. Even with the delay Perl 6 has suffered, I don't see any other tool that has the feature set Perl 6 has to offer.
- DonHopkins 12y agoPerl is nothing but corner cases and magic context sensitive punctuation. If you think Python is as hard to understand as Perl, then you've never used Python.
- Ultimatt 12y agoSure but the nice thing is you can just add an extra catch all to the Grammar. Or better yet create a new Grammar for that specific sequencing machine and override the token. For example just adding: rule sequence:sym<seq> { <-[\>]>+ } Would mean you can just grab anything without worrying about if it's correct. Because a tonne of people abuse FASTA doesn't mean you shouldn't at first attempt to enforce something valid. The nice part of that is you know within the parser if you have clean data or something with a badly defined character like * in DNA when it should just be X or N or one of the other confusion characters. Also it's an advent post not some formal public codebase ;P The focus was on showing some of the more obscure features of Perl 6 that aren't like other languages. Python can look as crazy as you like, it's just the majority of people using it fall into the camp of avoiding the crazy. For example here's (https://gist.github.com/MattOates/6195743 https://gist.github.com/MattOates/6195743) some crazy Python I've written. The six line list comprehension is especially readable :P
- coriny 12y agoSorry, I was just pointing it out in case for anyone copying the code as they are both very common cases - especially dots as gap characters. The Grammar system does look quite nice. Parsing is definitely a Perl forte - though I seriously hope bioinformatics in general will continue to abandon these crazy formats that can't be read easily by humans or machine.
- Ultimatt 12y agoYeah I would rather all the awful formats would go away. But the process of removing them tends to just birth new awful formats into existence :(
- coriny 12y agoI'm increasingly of the opinion, after having to deal with biologist data format hackery for many years (they have no shame!), that every biologist should be taught relational database normalisation. That would make them actually stop and think about the structure of the data they're recording and how to make it computationally tractable. Even more, they should not even start their experiment until they've cleared their data recording structure with a bioinformatician. Then we ban them from using flat files. Related tables or nothing. I've wasted too much time, and they're not competent enough to work with JSON/XML/Excel.
- cjfields1 12y agoThe FASTA parsers for the various Bio* all support additional characters; if not then that would definitely be a bug. Saying that, 'support' completely depends on what the FASTA is and where it comes from (assembly, sequence database, alignment tool, etc), which is something the parser/grammar can't define up front from such a simple format. Much of the problem comes when validating the sequence for a specific alphabet or symbol set. If the alphabet isn't explicitly defined (again impossible to determine from the format without guessing) then you can certainly run into problems. This is also an issue when using FASTA for both regular sequence data and for alignments (e.g. length of the sequence would have to take into account possible gap characters generated from various tools like '-?.', or stops in protein seqs like '*'). These are tools that have been in widespread use for ~15 years, so if you have run into problems you should be more explicit in what they are so they can be addressed.
- coriny 12y agoBioPerl has probably got a lot better over the years, I haven't touched it in a long time (5 years+), so my views may be very out of date. I don't really want to get into reviewing Bio* libraries generally :) There is a general problem I think though, and other people have experienced the same. I've been trying to reason why, but I think it's due to multiple factors. By the time you've managed to get that bit working reliably with BioPerl you could have written it yourself. And it would be faster and use less memory. BioPerl has the problem of trying to solve every case, whereas normally I'm working with a limited subset of possibilities. So why? I think it's a combination of dodgy, manually hacked inputs (e.g PDB files!), the learning curve, poor backwards compatibility preventing upgrading, and, ah I don't know. Every time it's seemed like a good idea, and yet every time I've ended up abandoning BioPerl. Maybe it's too integrated with itself as a library of functions? Take parsing a FASTA. By the time I've read the BioPerl documentation, I could have already written that one line split string statement (because I e.g. know in advance the sequence string is always on one line). It's hard to overcome that laziness and make the commitment to learn it and become fluent in it.
- niels_olson 12y ago> glad to see more projects switch over to python. Just getting into bioinformatics and focusing on python. Glad to hear some validation about that decision.
- collyw 12y agoPeople, (bioinformatics in particular) are capable of writing truly awful Python code in my experience.
- coriny 12y agoBioinformaticians are generally capable of writing truly awful code... that's why I like to call myself a Computational Biologist. I still write crap code, but no one's expecting any better from a biologist ...
- niels_olson 12y agoThanks for keeping the bar low for us entry-level types :)