5 ms·
Adam for Kite here. We learned a lot over the past year and have worked hard to build several new features to Kite based on detailed user feedback (much of whic
by adamsmith 8y ago
Adam for Kite here. We learned a lot over the past year and have worked hard to build several new features to Kite based on detailed user feedback (much of which came from the HN community). We're excited to release two new features of special importance today:
* Line-of-Code Completions - Kite's completions engine can now predict several tokens of code at a time, powered by the most sophisticated AI code models available.
* Cloudless Processing - Kite now performs all processing locally on users' computers, instead of in the cloud. No need to upload your code to our servers. You don't even have to sign up for a Kite account.
We know privacy is a big concern for users, so that's why we decided to bring Kite off the cloud. You can learn more about this decision as well as the full Kite release on our blog post, linked here.
Our core belief is programmers spend too much time on repetitive work like copying and pasting from StackOverflow, fixing simple errors, and writing boilerplate code. That's why Kite uses AI to make writing code less repetitive and more fun.
Speaking of fun, we've set up a playground for you to try our Line-of-Code Completions out in your browser. We hope you enjoy it!
As always, Kite is free to download and use. And we no longer require user accounts now that we've moved off the cloud.
If you already use Kite (thank you for your support!), you now have these features via auto-update.
We're really looking forward to your feedback. The detailed feedback we've received in the past has been immensely helpful in getting us to this point. We'll be here all day to answer questions, too.
- drcongo 8y agoQuick question: Is it doing much differently from TabNine[1]? For me TabNine has been incredibly accurate at guessing what I'm wanting to type. [1] https://tabnine.com https://tabnine.com
- adamsmith 8y agoGreat question! TabNine is a different approach that uses far less semantic information than Kite. It really shines when you're talking about syntactic repetition, e.g. if you check out the screenshots on their homepage. They have a page about semantic completions, but the semantics are very shallow -- basically what attributes are on the instance you're accessing. In contrast, we've spent ~50 eng-years semantically indexing all the code on Github, building statistical type inference, and rich statistical models that use this semantic information in a very deep way. The result is that Kite can help more often, in ways that reflect a deep understanding of the semantics of the code you're writing.
- deleted 8y ago[deleted]
- shittyadmin 8y agoNot to mention by pushing malware to people via IDE addons, pretty good strategy for training those datasets.
- Drdrdrq 8y agoInteresting - is the model that learns from (semantically indexed) code bound by its license? One could argue it is derivative work.
- adamsmith 8y agoInteresting question. I’m not a lawyer, but I presume it’s derivative enough not to be copyrighted. For example, one could argue Google’s giant n-gram corpus is derived from many copyrighted webpages.
- tyingq 8y agoCurious how much better your approach is versus fairly vanilla "most popular ngrams".
- adamsmith 8y agoN-grams is the first thing you try when you want to statistically model code. We tried it in 2014 and the results are disappointing, if you want to do anything beyond identifying low level syntactic patterns. The intuition here is that in natural language, context is defined locally. If you want to know whether a 'they' refers to a male or female, for example, you look at the nearby text. In contrast, the flow of data and control through code is highly non-local, which is a lot of why techniques like N-grams don't work. This has been a very active area of research in academia since we started Kite in 2014. (Suggested google searches: [big code], [ml on code], etc.) All that said I don't have data on how our approach would compare to N-grams, but I'm guessing if you look at some of the academic research, you'll find that NLP techniques were abandoned, despite the early papers' focus on them.
- tyingq 8y agoThoughtful answer, thank you. People complain about HN having gone downhill...but answers like this one still pop up fairly frequently.
- hasperdi 8y agoCould you explain to us in details what kind of analytics data are you collecting? Also for which purposes are these collected? I am interested (and I bet more people here) in getting a comprehensive list minus the marketing BS.
- dhung 8y agoGreat question. We're in the midst of updating our privacy policy, but I'll give a full explanation here. It may be useful to differentiate between what data we collect and what data we don't collect in order to clear up any confusion. 1. What we collect We collect a variety of usage analytics that help us understand how you use Kite and how we can improve the product. These include: * Which editors you are using Kite with * Number of requests that the Kite Engine has handled for you * How often you use specific Kite features e.g. how many completions from Kite did you use * Size (in number of files) of codebases that you work with * Names of 3rd party Python packages that you use * Kite application resource usage (CPU and RAM) You can opt-out of sending these analytics by changing a Kite setting (https://help.kite.com/article/79-opting-out-of-usage-metrics https://help.kite.com/article/79-opting-out-of-usage-metrics). Additionally, we collect anonymized "heartbeats" that are used to make sure the Kite app is functioning properly and not crashing unexpectedly. These analytics are just simple pings with no metadata, and as mentioned, they're anonymized so that there's no way for us to trace which users they came from. We also use third party libraries (Rollbar and Crashlytics) to report errors or bugs that occur during the usage of the product. 2. What we don't collect * Contents (partial or full) of any source code file that resides on your hard drive * Information (i.e. file paths) about your file system hierarchy * Any indices of your code produced by the Kite Engine to power our features - these all stay local to your hard drive
- save_ferris 8y agoHow do you guys plan on monetizing this? This list of analytics seems fine at a quick glance, but I'm concerned that there's not a transparent path to profitability here. Just like any other service, there's no guarantee that you won't start collecting snippets of source code or other metadata to start selling once the VC's start applying pressure to generate income. What's your strategy?
- projectramo 8y agoThanks for moving it off the cloud and making it local. That is a huge step towards making us comfortable with the potential privacy aspects. Do you collect any information? And how will you make money? (I want you to be around, pay staff, do well etc)
- mattnewport 8y ago> Our core belief is programmers spend too much time on repetitive work like copying and pasting from StackOverflow, fixing simple errors, and writing boilerplate code. Do you have data backing this core belief up or is it just an intuition? I don't feel any of these things are problems I need a solution for. Maybe this depends on the type of programming someone does or the language they use but as a working programmer this is not a compelling sales pitch to me. You're offering solutions to problems I don't have. They're also not problems that I see to any great extent in junior programmers who report to me.
- adamsmith 8y agoOn an intuitive level, I would guess there are many, many developers debugging a "you forgot to cast your int to a string" error message as I'm writing this. More concretely, the StackOverflow thread for "Parsing values from a JSON file?" has almost 3 million views. We take this as evidence that programming is repetitive. (We should probably clarify that when we say 'repetitive' we mean on a global scale, not an individual scale.)
- mattnewport 8y ago> I would guess there are many, many developers debugging a "you forgot to cast your int to a string" error message as I'm writing this. I've always avoided dynamically typed languages because of these types of issues so I guess this seems like a solved problem to me - just use a statically typed language. Clearly there are many people who are using dynamically typed languages for whatever reason though so I can see how better tooling might be useful to them. > More concretely, the StackOverflow thread for "Parsing values from a JSON file?" has almost 3 million views. I certainly make use of Stack Overflow, I was taking issue with the idea that significant time is spent copying and pasting code from the site though. I often refer to it to answer a programming related question but I almost never then copy and paste code. I'm looking for a pointer to a library or API, to understand some confusing or poorly documented behavior or for a workaround to a bug etc. and getting an answer to my question rarely leads to copying and pasting code in my experience.
- cjhanks 8y ago"Our core belief is programmers spend too much time on repetitive work like copying and pasting from StackOverflow, fixing simple errors, and writing boilerplate code." Sounds like you will likely be introducing bugs into people's code. Or at the very least, regressing code quality towards the mean.
- adamsmith 8y agoKite aside, I think both of these things will happen as developers adopt these technologies. We're seeing similar effects with autocorrect on your phone and self driving cars (i.e. there will still be accidents, and they tend to drive more slowly), but on balance these technologies are positive.
- thecatspaw 8y agoFor me autocompletion is the second thing I turn off on a phone (right after vibrate on touch). There is nothing more annoying than writing a word correctly, and autocomplete "fixing" it, just because it was slang word, or a word in another language because I forgot the english word in that moment.
- LeifCarrotson 8y agoA lot of that is because the minimal phone keyboard doesn't give you a good UI to accept or decline a suggested autocompletion. With a keyboard and mouse and large monitor, you can display the completion prompt and let the user select or ignore from the various options with Ctrl-Space or similar. When "Select the first autocomplete over what I typed" is the default behavior, yeah, that sucks. I often end up using the "key combo" Space,Backspace,Space to accept the autocomplete on my phone. But I definitely see utility in offering a correction for, eg. 'DateTime.Format("yyyy-mm-dd" -> yyyy-MM-dd' or any of the million other idioms that we have to remember or look up. What I don't understand is how Kite can offer useful local-only suggestions - Is the default dictionary that comes with it preloaded with a million lines of open source, or is it literally just my code being suggested?
- 8y ago
- sneak 8y agoHere’s some feedback: pretending you’re grateful for users screaming at you for breaching their trust is not a good way to regain that trust. > If you're already a user, Kite has been auto-updated and is now working locally without sending code to the cloud. If you previously uploaded code to our servers, you can remove your data via our web portal. > We're grateful to our users for helping us reach this point. Why aren’t you scrubbing all surreptitiously uploaded data proactively? Why are you pretending you are glad you got caught doing sketchy things? You’re going to have to do a lot better than that to convince us that your company suddenly operates with integrity now simply because you got caught.
- bitL 8y agoThis is excellent! Thanks for allowing me to play with it! Did you publish any papers on how you approach intelligent code completion at Kite (the parts you can make public or just rough areas I should look into? I understand you have many emerging competitors). Thanks!
- adamsmith 8y agoYes, this may be a good starting point — https://github.com/src-d/awesome-machine-learning-on-source-code https://github.com/src-d/awesome-machine-learning-on-source-... Thanks for checking it out!