5 ms·
Bioinformatican here with couple years of bench experience. Nice writeup, but what would be the use case here? I would imagine something like automatically anno
by daemonk 10y ago
Bioinformatican here with couple years of bench experience. Nice writeup, but what would be the use case here? I would imagine something like automatically annotating motifs/restriction sites and then searching for them would be more useful than searching for near-exact sequence?
For example, finding all plasmids with a certain primer sequence, specific restriction sites, promoter motif...
- dekhn 10y agoRegular expression-based searching for bio sequences has been used for a long time but I don't think anybody thinks regular expressions are the write language for this, although I guess it depends a lot on the usecase. i think most people want a sensitive search that has a probabilistic model for substitutions and for indels, ideally with a heuristic so it's fast. BLAST does this, and it also includes support for "profiles" which are probabilistic models (and HMMER has an even more powerful version of said models).
- theophrastus 10y agoProtein Structural Chemist (i guess) here and I'd like an even higher level 'regular expression' which would allow one to search DNA sequences likely to code for protein domains with likely secondary structure, (something like a string to search for "Secstruc" within "Exon Structure" [1]). The "likely" will always be a problem in these sorts of attempts; any genomics regular expressions should allow for the option of the statistical best hit. [1] http://www.rcsb.org/pdb/protein/P00918 http://www.rcsb.org/pdb/protein/P00918
- jonstewart 10y agoCan you elaborate on the types of expressions you want a bit more, and how they should match?
- kusmi 10y agoHave you looked at T-coffee? They have this new feature called expresso which might do what you want.
- joshma 10y agoPosting on behalf of Hannah, since her new account is getting blocked from replying: hi! resident scientist at Benchling here! searching near-exact sequence can be really useful! for example, sometimes i would like to find if any of my plasmid has a certain signal peptide, it would be so hard without this search algorithm because so many DNA sequences can be translated to the same signal peptide. it would be great if i could just paste the amino acids and i could identify which plasmids contain the signal peptide.
- daemonk 10y agoI see. I guess near exact match functionality would be great for plasmids where there is a custom sequence inserted. But wouldn't a pre-processing pipeline that either automates annotation of the plasmid based on commonly used motifs (m17 primers, t7, sp6 promoters, fluorescent protein sequences, antibiotic resistence...etc) or requiring the user to annotate their plasmid during submission be more ideal? I would imagine the more common use case is to search for these elements?
- mbreese 10y agoBlastx? Or are your motifs too short? I know you mention that this isn't a use case for blast, but it's simple to keep a private blast index to use for cases like this.