10 ms·
Launch HN: MutableAI (YC W22) – Automatically clean Jupyter notebooks using AI
Hi HN, I’m Omar the Founder and CEO of MutableAI (YC W22) (https://mutable.ai https://mutable.ai). We transform Jupyter notebook code into production-quality Python code using a combination of AI (OpenAI codex) and PL metaprogramming techniques.
I'm obsessed with clean code because I've written so much terrible code in the past. I went from being a theoretical physics PhD dropout -> data scientist -> software engineer at Google -> research engineer at DeepMind -> ML engineer at Apple. In that time I've grown to tremendously value code quality. Clean code is not only more maintainable but also more extensible as you can more readily add new features. It even enables you to think thoughts that you may have never considered before.
I want to reduce the cost of clean, production-quality code using AI, and am starting with a niche I'm intimately familiar with (Jupyter), because it's particularly prone to bad code. Jupyter notebooks are beloved by data scientists, but notorious for having spaghetti code that is low on readability, hard to maintain, and hard to move into a production codebase or even share with a colleague. That’s why a Kaggle Grandmaster shocked his audience and recommended that they do not use Jupyter notebooks [1].
MutableAI allows developers to get the best of both worlds: Jupyter’s easy prototyping and visualization, plus greatly improved quality with our AI product. We also offer a full featured AI autocomplete to help prototyping go faster. I think the quadrant of "easy to develop in" and "easy to create high quality code" has been almost empty, and AI can help fill this gap.
Right now there are two ways of manipulating programs: PL techniques for program analysis and transformation, and large scale transformers from OpenAI/DeepMind, which are trained on code treated as text (tokens) and don't look at the tree structure of code (ASTs). MutableAI combines OpenAI Codex / Copilot with traditional PL analysis (variable lifetimes, scopes, etc.) and statistical filters to identify AST transformations that, when successively applied, produce cleaner code.
We use OpenAI's Codex to document and type the code, and for AI autocomplete. We use PL techniques to refactor the code (e.g. extract methods), remove zombie code, and normalize formatting (e.g. remove weird spacing). We use statistical filters to detect opportunities for refactoring, for example when a large grouping of variable lifetimes are suddenly created and destroyed, which can be an opportunity to extract a function.
Some of the PL techniques are similar to traditional refactoring tools, but those tools don’t help you decide when and how to refactor. We use AI and stats to do that, as well as to generate names when the new code needs them.
A tool that reduces the time to productionize code can be compared to having an extra engineer on staff. If you take this seriously, that’s a pretty big market. Stripe Research claims that developer inefficiency is a $300B problem [2]. Just about every tech company would become more efficient through increased velocity, fewer errors, and the ability to tackle more complex problems. It may even become unthinkable to write software without this sort of tool, the same way most people don't write assembly and use a compiler.
You can try the product by visiting our website, https://mutable.ai https://mutable.ai and creating an account on the setup page https://mutable.ai/setup.html https://mutable.ai/setup.html. License keys are on the setup page once you’ve signed up (check your mailbox for an email verification link). I’ve bumped up the budget for free accounts temporarily for the day, I hope you enjoy the product !
In addition to inviting the HN community to try out the product, I’d love it if you would share any tips for reducing code complexity you’ve come across and of course to hear your ideas about this problem and tools to address it.
[1] https://youtu.be/tsGGpe-onZI?t=1067 https://youtu.be/tsGGpe-onZI?t=1067
[2] https://stripe.com/files/reports/the-developer-coefficient.pdf https://stripe.com/files/reports/the-developer-coefficient.p...
- hervature 5y agoI think bad code in Jupyter is a symptom of both the tool and the user. Jupyter's flexibility for running cells out of order should never have been allowed. Code developed in this way will not be able to be transformed into a sequential program. One example of this is illustrated in your demo, how did the algorithm decide input_shape was an input but num_classes should remain a global? For me, it is obvious that the user's intent was that these are the variables that define the model. Personally, I wouldn't pay those prices to hide the data scientist's problems. However, if you developed extensions to prevent and encourage the user to overcome the tool's problems, I would buy that for myself. Also, you probably don't a demo that would fail to run in the second cell (in both unclean and clean code) .
- oshams 5y agoThanks, founder here. I personally try not to run cells out of order. But I feel like almost anyone at some point will want to run cells out of order. The extraction is based partly on statistical edge detection. We are working on training transformers on actual diffs on GitHub that would be more natural.
- hervature 5y agoKeep up the work, I hope you get traction. This space is definitely in need of something. In particular, I will be keeping tabs to see your upcoming tests feature.
- oshams 5y agoThank you! I completely agree. I certainly don't think people should give up using Jupyter because of the frustration of keeping code quality high enough to port to other environments. Test feature is coming very soon. Please feel free to email me at omar@mutable.ai if you have more thoughts on this.
- version_five 5y ago> Personally, I wouldn't pay those prices to hide the data scientist's problems I have data scientists that work for me that have this problem. I'd rather that everyone I hire wrote the kind of code I want out of the box, but for lots of reasons it's not always possible. I have tried enforcing frameworks and ended up wasting more money trying to get people to change their workflow. A developer or data scientist (even an inexperienced one) is expensive enough that $30/mo is nothing if it makes them more productive. I'd guess this is what the company is banking on. (You'll see I made another comment where i say i don't like their pricing model, this is a matter of principle because I don't believe in just trying to wrap a program in SaaS in order to get more money without offering something beyond "you can run my code if you pay me". But the cost / benefit is still easy to justify)