8 ms·
Stanford A.I. Courses
- redeux 3y agoCan I take these courses online for free or is this an ad for Stanford?
- spmurrayzzz 3y agoSome of the courses have been available for free on YouTube for quite a while: https://www.youtube.com/@stanfordonline https://www.youtube.com/@stanfordonline There's also Coursera courses that are much of the same content (taught by Andrew Ng as well in many cases). They have specializations for Machine Learning [1], Deep Learning, etc. These are paid via Coursera subscription, but financial assistance is available [1] https://www.coursera.org/specializations/machine-learning-introduction https://www.coursera.org/specializations/machine-learning-in...
- pmulard 3y agoI recently completed the specialization with Andrew Ng and think it’s a fantastic introduction to ML. It has a good blend of theory, practical tips, and coding. If anyone is interested, I’ve published detailed notes and my submissions for the lab assignments: https://github.com/pmulard/machine-learning-specialization-andrew-ng https://github.com/pmulard/machine-learning-specialization-a...
- ecks4ndr0s 3y agoI am in no way affiliated to Stanford. I don't think you can take the courses for free, but you can sure as hell read through the slides for many of the courses. Cheers!
- CamperBob2 3y agoHonestly, Andrej Karpathy's video series on YouTube is good enough to keep me from even looking at for-profit courses. That attitude might change as I get further along in them, but for now I'm a big fan of his pedagogical approach.
- ayhanfuat 3y agoIndeed most of these are not available to the public.
- JimmyRuska 3y agoYou might be better off looking at MIT OCW search, and selecting video lectures, looking at standford youtube, checking out the 2019 videos https://ai.stanford.edu/stanford-ai-courses/ https://ai.stanford.edu/stanford-ai-courses/ Most of these look like they're just the slides and syllabus, correct me if I'm wrong.
- leminimal 3y agoAre there project-based tutorial that talks more about neural net architecture, hyperparameters selection and debugging? Something that walks through getting poor results and make explicit the reasoning for tweaking? When I try to use transformers or any AI thing on a toy problem I come up with, it never works. Even Fizz-Buzz which I thought was easy doesn't work (because division or modulo is apparently hard to represent for NNs). And there's this blackbox of training that's hard to debug into. Yes, for the available resources, if you pick the exact same problem, the exact same NN architecture and exact same hyperparameters, it all works out. But surely they didn't get that on the first try. So what's the tweaking process? Somehow this point isn't often talked about in courses and consequently the ones who've passed this hurdle don't get their experience transferred. I'd follow an entire course on this if it were available. An HN commenter linked me to this https://karpathy.github.io/2019/04/25/recipe/ https://karpathy.github.io/2019/04/25/recipe/ which is exactly on point. But it'd be great if it were one or more tutorials with a specific example, wrapped in code and peppered with many failures.
- jwilber 3y agoThere’s an interactive neural network you can train here, which can give some intuition on wider vs larger networks: https://mlu-explain.github.io/neural-networks/ https://mlu-explain.github.io/neural-networks/ See also here: http://playground.tensorflow.org/ http://playground.tensorflow.org/
- candiodari 3y agoThere's no great answer to this question. It is a bunch of tricks. Fundamentally: If you're saying FizzBuzz doesn't work, presumably you mean that encoding the n directly doesn't work. Neither does encoding n from 0 to 1 or between -1 and 1 (and don't forget: obviously don't use relu with -1 to 1). It doesn't. Neural networks can do a LOT of things, but they cannot deal with numbers. And they certainly cannot deal with natural or real numbers. BUT they can deal with certain encodings. Instead of using the number directly, give one input to the neural network per bit of the number. That will work. Just pass in the last 10 bits of the number. Or cheat and use transformers. Pass in the last 5 generations and have it construct the next FizzBuzz line. That will work. Because it's possible. To make the number-based neural network for FizzBuzz "perfect" think about it. The neural network needs to be able to divide by 3 and 5. They can't. You can't fix that. You must make it possible for the neural network to learn the algorithm for dividing by 3 and 5 ... 2, 3 and 5 are relative primes (and actual primes). So "cheat" and pass in numbers in base 15 (by one-hot encoding the number mod 15 for example). PM me if you'd like to debug whatever network you have together over zoom or Google meets or whatever. https://en.wikipedia.org/wiki/One-hot https://en.wikipedia.org/wiki/One-hot This may be catastrophically wrong. I only have a master's in machine learning (a European master's degree, meaning I've written several theses on it (didn't pass first time, had to work full time to be able to study), and I was writing captcha crackers using ConvNets in 2002. But I've never been able to convince anyone to hire me to do anything machine learning related.
- ripvanwinkle 3y agoLooking for guidance here. There are a lot of courses out there on AI from esteemed institutions at that. What do people recommend as a curriculum for someone with a formal univ education in CS albeit from a while ago and who has programmed extensively though not in Python. The goal at the end is to have a deep understanding of the LLM space and its adjacencies.
- azmodeus 3y agoI would start with a fastai course such as practical deep learning for coders. After doing one of the fastai courses you will have some applied Python project experience and you can hone in deeper on a particular part of the project you are more interested in intellectually.
- kulikalov 3y agoDefine “deep understanding” here? You certainly have to lean python, at least because you are gonna need it for data manipulation and cleaning no matter what you do in this field.
- tonmoy 3y agoAlthough I myself am not related to the industry or academia pertaining to AI, I have heard many people speak highly of the zero to hero course by Andrej Karpathy: https://youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ https://youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9Gv... I myself loved it and learned a lot, but YMMV
- itissid 3y agoI think the way courses are taught can give you some needed grounding, like you should always take a good linear regression class. But I think that is as far as it gets you, a theoretical base. Honestly the issue is that most ML programs are taught as being some kind of additive skill set: the more courses you take the better or selection of the right kind of courses gets you some where. In reality: 1. most real world problems are also about subtraction knowing what not to try and why it might not work. Like when I ask people about Recommendtaion engines for recommending colocated things, people pile on embeddings, in reality its about finding good false negatives to train datasets, calibration of classifier output and those are really hard problem. Embeddings may be necessary but are the least of your worries. 2. Most companies will not teach you about the fundamentals of stats; you will be lucky if you can get a mentor in a company that has both the theoretical rigour and the practical implementation skill to solve problems. 3. Most ML problems require engineering to work as well, for example you can't use Bayesian MCMC to do most things at scale. Its why Topic models that used statistical models like simulating posterior were crazy expensive on large datasets.
- itissid 3y ago4. Models are taught like an end, but courses don't teach you to mix them for debugging. They are usually a means to an end for example say you are using decision trees and your models are acting up, you could still try some debugging techniques from linear regression like residual analysis or plotting variable slopes of each variable vs Y to debug before jumping for shapley values. The reason is not that using shapley values is bad, they are great, but you can get a lot of insight by having some base models that are simpler to debug.
- lolinder 3y ago> most real world problems are also about subtraction knowing what not to try and why it might not work This is true in most fields. I view school as giving you a broad overview of everything that you might need in your field, but for any given problem it will be on you to narrow it down to the solutions you actually need and then to learn that specific set of solutions well enough to apply it. People fresh out of college will usually try to apply everything all at once until they learn—either from a mentor or their own hard experience—to filter it down. It might be that ML has it worse than other fields right now not because it's taught wrong but because it's new enough that there aren't enough mentors with decades of war stories.
- antegamisou 3y agoWhy is Convex Optimization (EE364a) not included? https://stanford.edu/class/ee364a/ https://stanford.edu/class/ee364a/ https://www.youtube.com/playlist?list=PL3940DD956CDF0622 https://www.youtube.com/playlist?list=PL3940DD956CDF0622 It's one of the best courses to take if you want to obtain some fundamental understanding of the mathematical concepts behind AI. Yes there's much more to it than NNs/transformers/'Attention is all you need' paper/whatever else is trendy right now. No, don't expect to do important, as in employable, work if you won't be spending some time truly understanding the mathematical foundations.
- ecks4ndr0s 3y agoAgreed.
- mhh__ 3y ago> Yes there's much more to it than NNs/transformers/'Attention is all you need' paper/whatever else is trendy right now. No, don't expect to do important, as in employable, work if you won't be spending some time truly understanding the mathematical foundations. People only seem to want the former.
- caddemon 3y agoI hate academic trend following as much as anyone, but is it really true people are not employable in this space if they don't understand mathematical foundations? Sure, they're not getting a job at DeepMind, but it feels like there are many successful ML grifters these days too. Maybe I'm just on Twitter too much.
- BigElephant 3y agoIf I search the intro to robotics course online, I see there is a playlist from Stanford but the videos are 14 years old. Does anyone know if there are more recent videos? Or are there other courses that are good for robotics?
- imranq 3y agoI would not start with any course unless you had a project in mind that would take the knowledge from the course to get started. Otherwise you risk wasting a lot of time for knowledge that won't help you in any way and will get outdated in a few months anyway
- lolinder 3y agoI took a deep learning course in late 2019, during which we implemented transformers as described in Attention is All You Need and fine-tuned GPT-2. The output was amusing, but useless, but I still remember the basic principles. Now, a few years later, transformers are the tech, and GPT-2's successors are the most hyped technologies of the century so far. All of which is to say that I wouldn't assume that coursework without immediate application is useless. I'm in a much better position to jump in on the latest AI stuff than I would be if I hadn't taken that course.
- victor106 3y ago>I took a deep learning course in late 2019, during which we implemented transformers as described in Attention is All You Need and fine-tuned GPT-2. Which course was this?
- imranq 3y agoI mean if someone likes the subject, then they should invent a project that forces them to build something or write something up. Otherwise its easy to "fake" learning by doing the motions on a bunch of tutorials and quizzes.
- caddemon 3y agoIt really depends how good the course is, with well-designed problem sets/project prompts you can't really fake learning (assuming you actually complete them). Is it going to be totally exhaustive of everything you may need to know in practice? Obviously not, but no single project will be either, especially not for such a broad field as machine learning. Independent projects can definitely be a great way to learn, and yes many courses are shitty. But it is also very possible to take a good course and walk away with new knowledge you didn't even realize you needed. Some of my favorite projects actually started with an idea from a course, and then I learned even more in order to further expand on it. Synergy between project-driven and course-driven education can be a powerful iterative process.
- peter_retief 3y agoAndrew Ng is excellent.
- Exuma 3y agoCare to explain cross entropy simply? That’s where I paused currently
- mellavora 3y agoShooting from the hip: entropy of a single signal, say a sequence of letters, "ababababab" is the scaled "average" surprise per letter. So if they are uniformly distributed, each letter is equally likely/unlikely to come next in the sequence, where if instead one letter only 1/1000th of the time (aaa....aaa...aa..a.z.aaaa), then when the rare beast shows up, it is a big surprise, so the total amount of surprise available in the sequence is high. That's entropy. The same thing would be true for a sequence of numbers. But what if there is some relationship? if aaabaa occurs frequently with 111211, if you line up the sequences by timestamp? In this simple case, if you know the letters and you can spot the relationship, then there is zero surprise in the number sequence. The cross entropy "letters plus numbers" has the same entropy as "letters" or "numbers" in isolation. And as you move away from the 1:1 correspondence, you'll see the cross entropy increase until it reaches its max at "entropy(letters) + entropy(numbers)" -- no information shared between the two systems. To bring it home, I think of cross entropy as the amount of information shared between two signals. Others might think of it slightly differently.
- tczMUFlmoNk 3y agoMostly yes, but to your second paragraph: > if instead one letter only 1/1000th of the time (aaa....aaa...aa..a.z.aaaa), then when the rare beast shows up, it is a big surprise, so the total amount of surprise available in the sequence is high …when a Bernoulli distribution is skewed, the maximum surprise is high, yes, but the average surprise (= entropy) is low. The entropy of a Bernoulli distribution is maximized when p = 0.5 and falls off to either end: https://en.wikipedia.org/wiki/Binary_entropy_function https://en.wikipedia.org/wiki/Binary_entropy_function For your examples, if the sequence is uniformly distributed (Bernoulli(1/2)), the entropy is log(2) ≈ 0.693 bits per symbol; if instead one letter occurs 1/1000th of the time, the entropy is about 0.0079 bits per symbol.
- Exuma 3y ago
- AJRF 3y agoI've moved from "traditional" software engineering to a role of working with ML (building + deploying models used in product features) and of the team I work with - and my extended communication with developers at other companies making the same transition - every single person has said the Francis Chollet book (Deep Learning with Python) is all they really needed. It walks a very thin line between too little info and *just* enough to get you to the point where you know what you don't know (the productive point) and it explains the Math in code samples. It really is a very good way of teaching. When I was reading, I thought the theory covered was too far from the Mathematical base, but I found my self being surprised at how I could hold my own in discussions that moved in to theory. That said, this likely won't be enough for you to be a researcher - but I imagine for a lot of people tempted by courses like the OP - that isn't the actual end goal anyway.
- ra7 3y agoThanks for the book recommendation! I’m interested in making the same transition. Can I ask what you did to be considered for a role in ML coming from a software engineering background? Did you showcase any personal projects in your resume?
- AJRF 3y agoSo this might not work for you, but I will tell you my path anyway. 1. Was one of the first members of the #ai slack channel inviting some people I had in person conversations about AI with. 2. I posted _a lot_ in there. Stuff about regulatory updates, people using co-pilot, cool github repos, little demo projects I was working on. 3. Now this was pure luck and probably the best thing to push me over the boundary, there was a hackathon. I thought "Hmm if I make a kick ass demo showcasing generative AI here, a lot of high up people will see it" - that 100% happened, CTO reached out to me saying demo was great and that people will be in touch. 4. I started really digging in to how I could provide value to our existing data team - be that code, deploying things, bringing some of my engineering know how to that team. This point the #ai channel really started to grow and the head of data and engineering started talking to me and directing people my way based on what they saw at the hackathon. 5. Did a demo of my hack in the company all hands which the CEO was MC'ing. 6. Started having fortnightly 1 to 1s with head of data at this point 7. Floated idea of team taking a little subset of good and motivated people from other teams for a short time to investigate and implement LLMs in some small way into our apps. That team has now grown to effectively investigate any and all use cases (internal and external) for generative AI. 8. I started reading more theory and also following a bit of a road map for things I should learn to have a better picture of how to actually bring LLMs in some form to production (fine-tuning, vector dbs, functions, guard rails). 9. Now I am just building some quick feature in the mobile app to show case the value of the team to exec as quick as I can, which should give us few months cover to work on the thing I am really interested in - multi-arm bandit LLM that uses our existing models. This was pretty much it. Seems trivial, but in between each points was lots of reading, tinkering, working on weekends, but its totally possible. The ML + AI focused PhD's in your company likely need help from engineering but don't know it - bringing those two groups together quickly shows how you can be useful. This post was helpful; https://blog.gregbrockman.com/how-i-became-a-machine-learning-practitioner https://blog.gregbrockman.com/how-i-became-a-machine-learnin...
- deleted 3y ago[deleted]
- ecks4ndr0s 3y agohttps://stanfordasl.github.io//aa274a/ https://stanfordasl.github.io//aa274a/ (as the CS237A Principles of Robotic Autonomy I link seems to be broken)
- Havoc 3y agoI get that it’s basics but is a ~4 year old course still the way to go given pace?
- BossingAround 3y agoAre the lectures not available publicly?