4 ms·
I thought of making a tool for language-agnostic verification and/or refactoring advisory. That would be a modular system with a design looking something like:
by Eli_P 7y ago
I thought of making a tool for language-agnostic verification and/or refactoring advisory. That would be a modular system with a design looking something like:
AST -[basis]-> Normalized AST -[to sequence]-> Common Representation [graph; embeddings] -> [Code Structure DB; Temporal DB(reflects what's a code actually doing, runtime info involved)] <-[query for anomalies]
The idea is simple, we teach some database, feeding with lots of source codes. We may want also to define some rules, like, what's a deadlock, what's an event, etc. I've no clear idea how to implement the rules, but it should be possible to make an embeddings for the temporal data and map it onto a space where checking for correctness can be done in linear time, fingers crossed. But here comes the question, what's a temporal db should be? Should it be a separate data definition language, or code is the data?
As for basis, we're supposed either to use some generic calculus and its subset for the particular programming language, or somehow make it possible to aggregate into the same database with different bases.
Obviously this approaches ask for machine learning instead of traversing the graphs. I think there's good reason for that. For example, if I've written many chunks of code which actually do the same thing, it'd have taken enormous time to prove that using combinatoric approach (graph self-similarities, NP).