3 ms·
The semantics of such a language would the the primary challenge. If you want to be able to compile a statement like "implement a protein that cleaves proinsuli
by hyperion2010 8y ago
The semantics of such a language would the the primary challenge. If you want to be able to compile a statement like "implement a protein that cleaves proinsulin to produce insulin" there is SO much context that the real challenge is not in managing the compilation to ACTG, but the surrounding environment.
At what temperatures?
In what pH range?
In what species?
With what condon bias?
When should it be expressed?
At what rate should it mutate?
Do you want any other nearby 'functions' from mutation?
Is this a membrane protein?
Do you want restriction sites?
Is this for insertion into a plasmid or for insertion directly into a genome?
What bacterial cell line will you use to maintain the plasmid?
If you are expressing the protein for purification what cell line or bacterial strain will you be amplifying?
Do you want a single sequence that will work for all of these or are you willing to compile a different version for each combination? etc. etc. etc.
The list goes on and on.
The sequence you compile to will depend on those environmental parameters and the semantics of that 'environmental ISA' will likely be highly specific for many high level descriptions. I imagine you could produce sequences that were more robust, but the simulation time required to generate and validate them would grow accordingly.
All of this not even mentioning the fact that you also absolutely must specify _all_ the things it should not do, such as cleave a bunch of other proteins, or bind non-specifically and form aggregates in the cytoplasm, etc. A language level list of defaults here would certainly be a requirement, and that means the space that you will be optimizing in is absolutely massive. I don't even want to imagine how slow it would be. Probably faster to synthesize a bunch of variants and test them all in real cells.
The simpler case of taking an NCBIGene identifier and smashing it into an Addgene plasmid identifier and setting codon biases and optimizing for expression is a much more manageable task, and would probably be a building block for the more complex version.
- etrautmann 8y agoGreat considerations. Also what chaperone proteins are necessary for folding
- xrd 8y agoDo you have suggestions on learning the basics of all that you discussed here? I know nothing about it but it sounds fascinating.
- DoreenMichele 8y agoIf you search on "protein folding" you should be able to find some basics.