19 ms·
Act-1: Transformer for Actions
- tasdfqwer0897 4y agoHey, I helped make this! Happy to answer any questions.
- aaaaaaaaaaab 4y agoWhat was the training data?
- version_five 4y agoAlso, are there benchmark tasks that you either created or that already exist that you evaluated the model on? PS - please don't let this me used at a way to prevent human interaction. Chatbots are a disaster and literally the worst possible application of ML, as a shitty interface to a menu system. I hope this will be used in a way that is not consumer-hostile and that the company actively resists ignorant business attempts to use it to avoid paying for customer support.
- tasdfqwer0897 4y agoYeah, we did have to custom-build our own benchmarks. And we are not building a chatbot, we're building something collaborative that you can work with to accomplish the stuff you want to do!
- tasdfqwer0897 4y agoWe used a combination of human demonstrations and feedback data! You need custom software both to record the demonstrations and to represent the state of the Tool in a model-consumable way.
- tux3 4y agoOn an unrelated note, I imagine this can solve recaptchas and other simple non-visual challenges. Can I make ACT-1 Sybil a few thousand people on mechanical turk? Can I submit CVs with ACT-1 for entry-level full remote jobs and have it work for legacy companies, if those companies cannot setup ACT-1 themselves but provide a traditional human jobs interface? Can I put an interface that interracts with the real world through controls and text on a webpage and have ACT-1 take a physical presence?
- woojoo666 4y agoI thought captcha was designed specifically so that AI image recognition couldn't perform well at it
- TuringTest 4y agoIt was, for previous methods of image recognition. Generative adversarial networks and models trained on huge amounts of data might change that.
- learndeeply 4y agoThanks for answering questions! Are the example given in the blog post considered zero-shot learning? Was the model trained on the websites in the examples given (e.g. on the Redfin site)? How much labeled data was used?
- lee101 4y agoI'd be interested to see if we can work together in some way, i'm founder of https://text-generator.io https://text-generator.io which also crawls links to the web/images to generate better text: https://text-generator.io/blog/text-generator-now-researches-via-crawling https://text-generator.io/blog/text-generator-now-researches... Some people self host it, so it can solve some subset of the questions, the image understanding part is important, both it can understand a question with image e.g. if a user drops a fridge to search by image, or multiple images (e.g. which of the 10 images is the nicest looking fridge in 1 API request) as well. Also supports getting shared embeddings for images/text/code, which can be important for the information retrieval/question answering example where it needs to first find the relevant context on wikipedia then feed to the reader model to read it out Also do other custom stuff like retraining etc. Thanks, Lee https://leepenkman.appspot.com/ https://leepenkman.appspot.com/
- blueblimp 4y agoIn the Salesforce example, it's modifying database contents. Suppose the model misunderstands your request and modifies data in an unintended way (e.g. adding garbage data or, worse, deleting data). What's the recovery plan?
- WiggleGuy 4y agoThis is super cool; I'd want to stay updated. Do you think you guys could add an rss feed to your blog? I'd want to add it to my rss feed aggregator.
- codekansas 4y agoThis is incredible :) Will you be releasing more information about how the system was designed / how data was collected / how actions are executed?
- __sy__ 4y agoSecond that. Also what guardrails look like for it :)
- tasdfqwer0897 4y agoYes! We plan on putting out a more detailed technical post soon.
- holoduke 4y agoHow does the AI alter it's models during a process? I thought the weight models are pregenerated and not altered once used in a real-life app.
- visarga 4y agoCould be prompt based memory, or fine-tuning a small part of the big model.
- mrits 4y agoNatural language interfaces are very limited and certainly not the next generation of computing. Granularity of functionality and composable input will always be more efficient as long as the original source is a human. I think the natural language part of your product is the lease interesting and certainly not the most impressive.
- amilios 4y agoOn the flipside, natural language interfaces have the potential to be extremely easy to use for anyone, including non-experts. Anyone can type a message to the computer, without having to learn the specifics of an interface's custom controls. There are different types of efficiency. I'm assuming you're referring to something like 'operational efficiency', while NLI wins on 'ease of adoption' per se.
- angrais 4y agoConsider search engines and searching for content more generally: are you certain that natural language is most often used? When I search for content, I use key terms to produce refined and better results. If you don't use such terms then what you're looking for may be difficult to find.
- dwrodri 4y agoSo, I actually thought the opposite: the growth in userbase that most big tech companies have seen over the 2010s probably meant the "death" of people creating queries á la AltaVista. But it turns out, I think I'm wrong, Google's own research says the average query is between 2-3 words in size, but apparently it has in fact been going up over the years. Here's my source: https://dl.acm.org/doi/pdf/10.1145/1753326.1753333?casa_token=vVa95Yql8LQAAAAA:LXHeBXPq3m5fhM5oi-gwk0Y62TxvCzZr0lbGn59zDYKCdJPM7g5teO4vZ2I8aXj_jwaOePlnVHD3 https://dl.acm.org/doi/pdf/10.1145/1753326.1753333?casa_toke...
- averylamp 4y agoI think this is true for everyone. Without ever knowing how google search works, somehow over time you figure out how to do prompt engineering and figure out that the results you want are more highly correlated with certain phrasing. Wouldn’t something similar apply here though, where after using it for some time you get an inherent understanding of what it works well with and does not? There’s always some sort of mix of things to learn and make available to the end user though, so that they understand how to use the tool to be successful. Disclaimer: I also work at Adept
- colemannugent 4y agoSo here's the main problem I see with this: >Anyone who can articulate their ideas in language can implement them I'd be shocked if even 10% of the users who can't navigate a GUI could accurately describe what they want the software to do. To the user who doesn't know they can use Ctrl-Z to undo, the first half dozen times the AI mangles their inherited spreadsheet might be enough to put them off the idea.
- ffhhj 4y agoBut those who can articulate will have a very quick automation tool to scrap data from the web.
- tartoran 4y agoI’ve been thinking for a while about a common people programming language able to interface with machines with pure casual conversation ( not exact commands) and I feel something it’s coming in the next decades even if not earlier. Imagine the ability to casually chat with a widget which understands flawlessly and where most devices would be able to communicate as well. This could eventually be used in psychotherapy, everything automation around humans and in nefarious ways as well. I’m only hopeful of a human augmentation scenario but there are countless ways it could become totally different.
- 1MachineElf 4y agoThey don't need to explain what they want the software to do, they just need to explain what they want ACT-1 to do. I agree with you that it won't be basic users, however, use anything long enough and you will become an expert. This vision would fundamentally change how people interact with computers.
- jacobr1 4y agoCertainly there is a huge middle ground. Vague, but common, use cases might have more articulate versions of the commands inferred. I find myself learning new tools all the time - I certainly have enough domain knowledge of many things to express intent without describing implementation. I suspect plenty of people are similar enough - just operating at different levels of abstraction. What I find more concerning would be people operating under misconceptions, or being more precise than needed, thus not actually accomplishing their objective with the introduction of irrelevant detail.
- i_am_toaster 4y agoI look forward to seeing the progress made on this in the future, but at this time I don’t see any potential in this product.
- anigbrowl 4y ago- OK here's my email - Please select all pictures of taxis to prove you are not a robot ಥ_ಥ Seriously though, the potential is good. I see several things they're doing right that have the potential to distinguish them from competing offerings.
- bluecoconut 4y agoWow! Love it, this is the most exciting thing I've seen in a while. I'm working on something similar, and it's so great to see others who seem to get-it and are chasing generalization in AI systems! A few questions: 1. I'm curious if you're representing the task-operations using RL techniques (as many personal assistant systems seem to be) or if this is entirely a seq2seq transformer style model for predicting actions? 2. Assumption: Due to scaling of transformers, I assume that this is not directly working on the image data of a screen, and instead is working off of DOM trees; (2a) is this the case? and (2b) if so, are you using purely linear tokenization of the tree or are you using something closer to Evoformer (AlphaFold style) to combine graphs-neural nets and transformers? 3. Have you noticed that learning actions and representations of one application transfers well to new applications? or is the quality of the model heavily dependent on app domain? I noticed multiple references to data applications (Excel, tableau, etc.). My challenge is that large language models and AI systems in general are about to hit a wall in the data domain because they fundamentally don't understand data [1] [2], which will ultimately limit the quality of these capabilities. I am personally tackling this problem directly. I'm tying to prove more coherent data-aware operations in these systems by building a "foundation model" for tabular data that connects to LLMs (think RETRO style lookups of embeddings (representing columns of data)). I have been prototyping conversational AI systems (mostly Q/A oriented), and have recently been moving towards task oriented operations (right now, transparently, just SQL executors). There seem to be good representations of DOM tree/visual-object models that you all are working with to take reasonable action, however I assume these are limited in scale (N^2 and all), and so I am wondering if you have any opinions on how to extend these systems for data (especially as the "windowed context grows" (eg. an excel with 100k+ rows))? [1] https://arxiv.org/abs/2106.03253 https://arxiv.org/abs/2106.03253 "Tabular Data: Deep Learning is Not All You Need" [2] https://arxiv.org/abs/2110.01889 https://arxiv.org/abs/2110.01889 "In summary, we think that a fundamental reorientation of the domain may be necessary. For now, the question of whether the use of current deep learning techniques is beneficial for tabular data can generally be answered in the negative"
- tasdfqwer0897 4y agoThanks - glad you like it! I probably won't get to all of these but let me try a couple: 1. There's a spectrum (sort of) between using full on RL techniques and just doing sequence modeling. We're trying to pick a reasonable place on that spectrum that lets us model whether things have gone well without doing too much fiddling. 3. It really depends on how closely related the domains are. I think it's safe to say that you should expect more transfer of abstract/high-level capabilities than nitty-gritty things related to the specific domain - that's part of why we're excited about training one big model to use all software tools.
- skybrian 4y agoWow, what a great way to make a mess online! I can see spammers using it, but who's going to trust this with access to any accounts they care about?
- visarga 4y agoI think it's going to be expensive to run and require an account with AdeptAI. Most usage will be to automate office work, RPA style. They could also detect malicious usage as the model sees the pages and takes actions in those pages.
- visarga 4y agoRelated - GPT-3 with a Python interpreter can solve many tasks. It is also language model + a computer, but on a different level. https://mobile.twitter.com/sergeykarayev/status/1569377881440276481 https://mobile.twitter.com/sergeykarayev/status/156937788144...
- bilsbie 4y agoI’m confused how it actually uses the interpreter? Or it just returns the syntax and you run it.
- teraflop 4y agoIt's right there in the middle section of the image: the Python program just appends your question to a prompt, sends it to GPT-3, extracts the code from the response, and calls exec() on it.
- TuringTest 4y agoWhat could possibly be wrong?
- blind666 4y agoIf this scales up, it can be thought of as "actionable Google search", and if taken to extreme, has the potential to make internet query-able for better or worse
- deleted 4y ago[deleted]
- midislack 4y agoWe gotta stop this lazy "X for Y" marketing crap. Seriously, if your product is just "X for Y" it doesn't even sound like a good pitch.
- quickthrower2 4y agoSometimes X for Y just literally means X for Y, like "Milk for lactose intolerant people". I think this is such a case.
- frenchie4111 4y agoOne thing I have noticed with heavy use of copilot/dall-e, is that it's great at getting you most of the way there. But a big thing it's not great at is repeatability. When relying on something like ACT-1 to do data entry in salesforce, I need to do roughly exactly the same thing every time, even if the context is slightly different or I tell it something slightly differently. How well will it be able to do that? Also this is very very cool, I love copilot, I hope I get to use this thing very soon.
- tasdfqwer0897 4y agoYeah this is a good point! We are spending a lot of time thinking about reliability and it's true that existing models fall a little flat here. I think ultimately the key to making this work really well is some combination of a) collecting and training on human feedback and b) doing intelligent things to samples from the model after the fact
- machiaweliczny 4y agoI think it would be better if these tools simply generated UI functions that could be named and used ala command list in editor. I think future UIs will be just talking and asking to find from big list of commands. Where AI will be able to navigate API from graphQL/OpenAPI description and maybe automatically plot data you want.
- lee101 4y agoAwesome, I wonder if the app is recording what we do so it can replicate what we do, or maybe if not it should have a training mode where we tell it what we do then do it so it can learn. I feel like some of this could one day be built using a shared model that understands HTML and JavaScript code etc with a few example prompts. Or maybe something that understands intent+a browser automation language like Selenium, if not then some custom input output language+training as adept alludes to. If interested in building something like this also checkout https://text-generator.io https://text-generator.io which already pulls down links and images to analyse to generate better text so has a lot of the required parts
- JediLuke 4y ago
- FeepingCreature 4y agoAct-1, please fix the hideous low contrast on the adept.ai website.
- deleted 4y ago[deleted]
- leetrout 4y agoSo many folks standardizing on swagger/openapi opens the door to training on structured api definitions... this never occurred to me before.
- d--b 4y ago“Open the pod bay doors, Act-1”
- atemerev 4y agoMany people worry that these things will take over our jobs. Worry not! Imagine how much work will be needed to fix things when these models will screw up, and how much we will charge for hour.
- joaquincabezas 4y agoI remember a long time ago, reading about Semantic Web and intelligent agents and dreaming of a Natural Language interface for planning journeys… “I want to travel from Seville to Berlin next October, avoiding weekends, for a two or three nights stay in a hotel by the river. Direct flights preferred.”
- codpiece 4y agoOK! You want Seville burning red October, avoiding weekends forty two or three nights/days....
- rajnathani 4y agoThe founders of this company are the main authors of the Transformer architecture: https://techcrunch.com/2022/04/26/2304039/ https://techcrunch.com/2022/04/26/2304039/