6 ms·
New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
- luciana1u 3mo ago[flagged]
- boulos 3mo agoDo you have a larger study planned for the Fall? It definitely seems promising. I'm curious how well you feel this worked because the subject was Statistics (objective grading) versus something more subjective like Civics or Literature. PS - I'd say this qualifies for Show HN, too! Do you
- ilaksh 3mo agoThey were using Sonnet 4.6 for some fre form responses so that could be applied to something subjective.
- boulos 3mo agoBut it's not clear that using Sonnet or any other LLM as a "grader" would result in the same improvement. For objective grading, you could be sure that the additional adaptive support is helping. For subjective things like writing style, literature, poetry, you end up with whatever Sonnet thinks is good (and randomly so). It still could be better for students, but it's not obvious that it would be (or maybe not as strongly?).
- albinahlback 3mo agoVery nicely typeset.
- kubb 3mo agoToo bad the educational use case doesn't make any money. Good LLMs are a game changer for people motivated to learn.
- TheLML 3mo agoI don't want to learn from hallucinations where it will change its answers based on me questioning their teachings. I use it for conversations in a language I'm learning, but I quickly learned that asking it grammar questions for example is not a wise decision.
- afro88 3mo agoCurious whether you were just bare asking it questions, or whether you provided it with lessons one by one with instruction that the lesson is the baseline truth etc
- treis 3mo agoAre we talking about human teachers or LLMs here?
- sumeno 3mo agoThis is such a lazy response to every LLM criticism
- Robotbeat 3mo agoWikipedia doesn’t make much money but is still helpful. LLMs don’t need to make a whole bunch of money to be helpful.
- jasondigitized 3mo agoNot sure if its education, but there is huge money in the college admissions process, e.g., SAT prep.
- Rperry2174 3mo agoHonestly whether or not this was effective seems less important to me than the adoption numbers. Text book reading in this course was 10-15% at baseline ... but this AI thing got 90% voluntary usage ungraded. Even if its worse per-hour than a textbook, you're now teaching 6x as many students _something_ instead of teaching a small minority everything. So really it just becomes an optimization problem at that point because most students are at least in the funnel/in the running to learn something. The paper kind of proves this itself ... they tweaked the quize formats mid-semester and where able to iterate which you can't do on a textbook that nobody opens in the first place
- baq 3mo agoI'd argue the results are even better: just reading a textbook doesn't really teach you much. You have to do exercises, but they're expensive to create and grade. LLMs with a proper harness (see paper) tackle both.
- deleted 3mo ago[deleted]
- rusbus 3mo agoThis is exciting because the effect size is so large. But as the author's acknowledged, selection bias is nearly impossible to control for in this non-randomized study: > and lacks randomized controls. Self-selection is the central threat: students who complete more quizzes may be more motivated or higher-performing generally But this is still a strong result. I'm excited to see more in this space.
- rahimnathwani 3mo agoThey tried to control for this. It's described in the first paragraph of section 4.
- syou1024 3mo ago[flagged]
- constantius 3mo agoInteresting, congrats. Are you planning on opening access to Phosphor?
- thadk 3mo agomaybe they did already via the "formerly known as" comment in the paper?: https://www.spongium.org https://www.spongium.org
- baq 3mo agoI'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of the student - agentic tutor pairs, a referee when the student and model disagree, etc. and most importantly still being the human teacher you can just talk to in the human education process. I'm convinced this is the future of education - models are there, we need the classroom tech to catch up. The alternative is obvious and quantified in the paper - students just use models to do their work for them and learn nothing.
- Buttons840 3mo agoI would add that somewhere in there should be a spaced repetition algorithm. Spaced repetition is very effective, but it's really really clunky to use. My unpopular opinion is that we all have Stockholm syndrome when it comes to creating "cards", and people talk about how valuable creating cards is; but I think it stucks, it takes a lot of time. If AI is already teaching me math (let's say), it would be nice to tell the AI/app "quiz me on this periodically", and then the AI makes up a fresh polynomial to factor (or whatever) and presents that to you according to a spaced repetition algorithm. Behind the scenes, the AI should have access to what has happened the last several times a specific topic has been quized, so the AI can watch to see that certain mistakes are resolved, and the AI might also know better how to correct the user if it has context about previous quizzes of that topic.
- tired-turtle 3mo agoBut the very act of making and organizing your card deck is part of the SRS! It “sucks” because you get no dopamine hit from a fresh desk, as the reward system is not yet in place.
- Buttons840 3mo agoAgain, I really think this is a viewpoint we've talked ourselves into to help us feel better about how cumbersome creating the cards are. I'm willing to grant that there is some value in choosing what to put in the cards, but most of the awkwardness around making cards is UI related. Nobody creates cards on their phone, or while they're walking (AI could do both of these) - people create cards sitting at their computer (like cavemen!) usually clicking through a clunky UI and managing thousands of cards with thousands of clicks. That sucks, and people probably wont realize it sucks until something better comes along.
- ilaksh 3mo agoShocking that a well executed AI tutor improves outcomes. Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI? Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet? How can interactive AI driven individually paced learning _not_ be obviously dramatically more effective?
- dominotw 3mo agoits like anything else. benifits students that are already motivated to learn. very few are actually motivated to learn and are just there to get a job or its just next thing that they have to do in life.
- hajile 3mo agoI fear even a lot of bright, motivated people will be so discouraged by AI doomers that they won’t bother trying to learn.
- skybrian 3mo agoSelection effects are extremely important in education. Dartmouth students have already had a large selection effect. If you try to apply this more broadly then it might not work. Motivation is also a huge part of the problem. I'm wondering if the novelty of the AI tutoring gets more people to try it and whether it would wear off? It's surprising to me that many students at Dartmouth don't read the textbook. You'd think college admissions would select for that? It seems promising but, as they say, more research needed.
- dghlsakjg 3mo agoLots of people in education will happily tell you how the past 15 years of tech integration has been a net negative. There ARE technologies that have improved things, but so much high-cost useless tech has been shoved into every level of education that many educators are incredibly leery of new tech. The issue is that while the underlying technology is useful, the way it gets integrated is frequently not. An administrator cuts a deal for a product they never have to use to an ed-tech giant for a huge amount. Because the ink is dry and a huge sum of money has been spent admins pressure educators to use the technology as much as possible regardless of outcome. In that context it makes a lot more sense why there is pushback and FUD among educators.
- mmarian 3mo agoConflicted about this study. On one hand, LLMs have been incredible for my personal learnings of new concepts. On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper. I'd also caution against attaching the tool to grading. That means students have to put more effort into the course, which increases the chances that they will use LLMs to save time rather than make the investment.
- zerobees 3mo ago> LLMs have been incredible for my personal learnings of new concepts. Mind if I ask what did you learn and how you're using it? The reason I'm asking is that I repeatedly felt excitement only to realize down the line that the explanations didn't actually translate into practical skills. I'm not sure it's even an AI problem, it's a "doing versus reading" problem. Same as with reading a pop-science article and thinking to myself that I learned something about physics or medicine or mathematics.
- mmarian 3mo agoVarious concepts when I joined new teams in domains I've never worked in before. And system design. So very practical, and where stakes were high.
- MoneyBurning 3mo ago[flagged]
- isomorphic_duck 3mo agoWhy did you make a new account to spam AI comments?
- rictic 3mo agoYes! Very exciting to see this. Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem
- Herring 3mo agoThe story around Bloom's two sigma is a bit complex https://nintil.com/bloom-sigma/ https://nintil.com/bloom-sigma/
- SgtBastard 3mo agoThank you for this and for parent’s comment - I know what rabbit hole I’m going down today.
- Herring 3mo agoHmm, it might be better to just go straight to what leading countries currently do (Finland/Estonia/Japan/Singapore/etc). Other factors probably play a bigger role, like school funding, school autonomy, teacher professionalism (high education level + continuing education + good pay), free school meals/healthcare/transportation etc
- guai888 3mo agoOne major limitation is the cost. It costs lots of money to train teachers and pay them good salaries. What leading countries are doing currently might not be sustainable for the long run. Cognitive abilities also different greatly from people to people in different discipline. There is no silver bullet.
- naishoya 3mo agoThe long term sustainability in at least one of those leading countries (Japan) has been shown effective for at least the period from the end of WWII to now. The principle difficulty they have is not maintaining high levels of educational success but in having enough children to keep the schools open and the teachers fully employed. But, that is a different concern than successful educational outcomes. The average level of capability and comprehension in fundamental disciplines for students completing primary is a direct result of some fundamental differences in the way they approach classroom organization: for one middle school and high school students do not change classrooms during a school day. Teachers rotate rooms while students remain in place and this maintains a less disjointed, less distracted transition between classes. There are still very little to zero technological advances to the teaching methods: blackboards/whiteboards, overhead projectors, hand written paperwork, textbooks. The reasoning is simple. The fundamentals which students in primary education are in need of learning do not change much. Last years' textbooks are perfectly useful and the cost of replacement is directly carried by each student. If they damage a book, they must replace it. Salaries for teachers are not unusually high, but they are also not low. The building and administrative costs are kept low, for one, by there not being a significant non-instructor labor expense of janitorial and maintenance workers. Schools share the pool of municipal HVAC and other trades for serious infrastructure, but janitors? Nope. The students are required to actually clean up after themselves, and to actually clean the whole school. It makes a difference, and those avoidable labor costs can be directed to proper compensation for the teaching staff. Anecdotal observation of the overall efficacy of this approach reaches areas not usually measured as 'educational success' but also includes one noteworthy artifact. The common clerk, shop keeper, cook, or gas tank delivery driver all know that i = sqrt(-1) and that a complex number is a pair of numbers, one of which is the coefficient of (i). Let that sink in. How many people graduate with a B.A. from western schools without that? There is not a 'silver bullet' but there are a full compliment of approaches which when applied together, consistently, and persistently, yield excellent sustainable outcomes. Back to the falling population problem. There are pluses and minuses from the choice to provide separate boys and girls middle and high schools; one of which is lower teenage pregnancy and it's inverse, lower adult birthrates. (It's not the only reason for the lower birthrates, but it's not having zero impact.) You cant win 'em all.
- KaiserPro 3mo agoI'm not an expert, but how much of this is down to novelty, ie https://en.wikipedia.org/wiki/Hawthorne_effect https://en.wikipedia.org/wiki/Hawthorne_effect ? (ie changing the environment can lead to short term productivity gains because either participants are aware they are being watch, or it breaks up the monotony and makes people work a bit harder. )
- radioactivist 3mo agoI am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial. Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform. Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript. If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.
- computerdork 3mo agoThis is a helpful explanation - am not a researcher so I have little idea how to run an unbiased, meaningful experiment (except that it takes a lot of effort and thought to run one). Useful analysis
- yorwba 3mo agoWorse, because students complained about the difficulty of the AI-graded quizzes, they switch to multiple-choice questions only, which increases engagement, but after analyzing the exam results they determine that multiple-choice questions don't seem to help and add AI-graded questions back, after which engagement drops again. That means their experiment design is partially caused by their results instead of the other way around, which is a bad situation to be in. Their statistical analysis is completely inadequate for dealing with this. And the change in engagement suggests that there's strong selection involved. Their attempt to use midterm scores to control for selection effects is unconvincing. Why not control for whether students used the platform more when there were only multiple-choice questions? Those are the ones who self-selected out of using the AI grader.
- wxw 3mo agoThe title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder. > constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria > Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient. They specifically call out that the "RAG chat assistant" part of Phosphor (the platform) wasn't used much. I commend the effort here, but I don't think these results are particularly noteworthy. The conclusion is essentially that people who do practice quizzes will do better on exams.
- fragmede 3mo ago> a practice quiz platform with an AI autograder. What do you think tutoring is?
- mkl 3mo agoDefinitely not just grading. Tutoring is explaining and back and forth discussion to impart knowledge, in context and in response to specific difficulties/confusion the student is having.
- sumeno 3mo agoWhat do you think tutoring is? Because it's not just extra quizzes
- woolion 3mo agoThe role of a tutor is to find your weakest point (or points) and give you personalized advice on how to improve on them. It is important that you can put trust on the tutor, as your weak point is likely due to a blind spot. When given criticism, it can be viewed as objective or subjective, i.e. a "question of taste", and thus the criticism would be viewed as invalid and not acted upon. As to whether the tutor should point your weakest point or some combination of weak point, it should depend on what helps you learn the most efficiently; it might be to focus on one aspect, or work in a more integrated fashion. Last (although that may be first), they should consider what are really your learning objectives to tailor that advice.
- RA_Fisher 3mo agoThis is super, but students will have access to AI during the test in real life, so it's ironically less realistic to remove it (thinking of the "... GPT-4 actually harmed subsequent performance by 17% when the tool was removed ..." part). I'm more curious how students perform on the test with vs. without AI.
- glenstein 3mo agoIn mice! Jk, but the skepticism is inevitable. I think we can be dubious about how AI mobilizes global capital while also appreciating tutoring as one of its best targeted use cases.
- or_am_i 3mo agoThe article explicitly calls out selection bias (this is entirely based on 90% that opted into using the tutor, there was no control group), I wish the headline did as well. "Engaged students score 0.71 - 1.30 SD better in tests" sounds like a much simpler explanation.
- cgearhart 3mo agoI used to TA a graduate level CS math class at Georgia Tech. We regularly saw that the students who self-organized study groups did dramatically better in the course than average. One semester they told us to put everyone in study groups to see if it helped. The effect disappeared. Turns out that it was the self-selection of the most engaged students into a small group that mattered, not the study group itself.
- sebastiennight 3mo agoSo there might be zero effect? If it's purely a correlation, then maybe those students would be more successful than average even without the study group. They're already the most motivated kids. Maybe they just do "motivated kid stuff" and would still outperform.
- cgearhart 3mo agoYes, that’s what I think at this point. There is no effect of the study group except as a support group. (That’s all it was for me when I was a student and joined the self-organized study group.)
- sebastiennight 3mo agoMy self-organized study group was quite a mix of characters, and while I don't doubt that a couple of my friends's grades did benefit from having me there to help them out, I do definitely think that most of the stuff I remember learning there was outside the boundaries of what any of our parents would have approved :)
- 3mo ago
- zerobees 3mo agoWhile there's some skepticism in the thread, I'm not particularly surprised if this is true. Children who can get human tutoring do a lot better. An LLM that can answer questions and patiently explain likely offers some benefit. What creeps me out about bringing LLM into early education is that it's a period where kids learn to socialize and cope with problems, and I do worry about forming substitute relationships with chatbots that are engineered for sycophancy / enablement. But I guess that's a problem either way, because almost every student will try an LLM at some point.
- sarchertech 3mo agoFrom my understanding the actual AI that was barely used. What was used was a quiz with an AI grader.
- NeutralForest 3mo agoInteresting article, wonder where we're going with this though, I find it's very difficult to keep LLMs on track and critical enough to be useful. Just want to say that: >In our deployment, student-reported reading completion baselines for MATH 010 were approximately 15%, with instructors estimating 10%. Individual student reports of reading compliance ranged from "literally no one does that" to "is this being recorded?" is hilarious
- tancop 3mo ago[dead]
- klustregrif 3mo agoA lot of pessimism in the comments, but I am just happy that we are seeing some work towards bridging the 2 Sigma gap for regular education vs. elite private tutoring. I can't imagine that people assume it's the physical presence of the tutor that is making the difference, it has to come down to the personalisation and expertise which is exactly what AI can provide in a form. And yea it might not be "there" yet. But if we don't start trying and studying then it'll never get there.
- johnfn 3mo agoTell me if I am oversimplifying, but I never understood the noise about the two sigma problem. Like, of course if you have a private tutor to immediately answer any question that pops into your head at the immediate moment you get confused, you are going to learn vastly more efficiently than in a large classroom where once you get confused you are likely to stay confused. To say nothing of how the pace will likely either drag way behind what you'd like, or accelerate too fast ahead of it. The environment is just obviously two sigma better. This just... seems obvious to me? In the same way that I will get stronger much faster if I have a physical trainer to tell me exactly what I am doing wrong when I do it? And it seems obviously unsolvable other than by getting everyone a private tutor (or AI..?). Asking from a place of curiosity.
- SgtBastard 3mo agoThe “problem” is exactly as you’ve framed it: individual 1-1 tutoring unambiguously improves education outcomes by a significant (2 StDev’s in Blooms study) vs any other understood method but delivering this universally is (was?) infeasible. Bloom looked at other methods used in concert to achieve similar improvements to “solve” the problem that could be delivered at scale. LLMs may provide a new path to the 2-Sigma improvements without the same delivery problems.
- klustregrif 3mo agoThe noice isn’t about the effect, it’s about trying to scale the outcome. We essentially have proof of a method that is significantly better than what is used with the broad population, but we cannot scale it to the broad population. That’s the “problem” in the two sigma problem. And in general ongoing from “it’s obvious to me” to a quantized effect is always an effort. For instance I don’t believe you are correct in assuming that one would see a two sigma effect in physical exercise when comparing someone who attends a regular exercise course va someone who spends the same time with a personal trainer. Two sigma is a lot, and you won’t be lifting significantly heavier weights from hav a PT vs doing starting strength training. In my opinion you would most definitely see a two sigma effect from doing steroids though. But this is all pure speculation which underlines that part of the “noice” is about the documented aspect of the two sigma problem, where we have a body of data to work against not just personal assumptions. That said if anyone has studies on personal trainers va course work in fitness I’d love to see my world image challenged by data.
- bobajeff 3mo agoEven if the research is flawed I'm happy they are trying this. They are taking advantage of LLMs to have less rigid tests and also give feedback. I think there is more potential applications possible with combining LLMs with reference/text books. Like how about an assistant that points you to the correct books/chapter/paragraph for the concept you need to understand better for a project you are working on? Or clarify any confusion you are having? Like a human tutor but infinitely patient and non-judgy + search engine.
- delis-thumbs-7e 3mo agoI currently study Multivariate Calculus by using very new and nodern method: I read the text book, while solving the examples of the general for,ulas, or try to come up with my own. Then I do a s$ht-ton of exercises. I only use LLM’s to quickly clarify confusing topics or notation, but not really much else. I cancelled my Claude subscription. Now I use just Mistral and local Vibethinker-3B, but they work just fine. Earlier I used Claude by giving it the course material and asking it to generate me exercises (our cpurse work went way over my head) and yeah i learned to differentiate a gradient or Jacobian, but it was very shallow - I knew the formulas, but not what they meant or how to apple them correctly. After I just filled glaring holes I had in Univariate Calculus by readong and doing, I actually started to understand something. Lon story short, in my experience Learning with LLM’s is ok with very unfamiliar material that is not too complex (there’s obvious problems of LLM’s themselves being pretty ghastly with maths sometimes), but at least it os not better than the traditional method of just putting your nose on the grimd stone.
- raegis 3mo agoIn the Matrix movies you can upload complex knowledge to a person in seconds. If LLMs could help with this even a little it would be useful. But alas, people like myself who are slow learners still get overwhelmed when trying to learn too much or too fast without rest breaks.
- usernametaken29 3mo agoWow, putting effort into training material, thoughtfully designing it, and relating the material to the final exam, will increase performance on said exam. So much AI so much wow. Like seriously. Most university statistics courses suck big time, so literally any effort put into them will Improve the field. I’m happy the authors want to improve education but they don’t seem to understand that preparing questionare style material is a confounding factor which could very well explain the better performance too… instead of cramming AI into the next thing. I’m generally opposed to AI on basic textbooks. You don’t want hallucinations imprinted on students who have no idea and can’t judge the quality of the generated text. Some things require effort, reading intro to statistics is one of them, and it’s for a reason, the effort IS the learning
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- zkmon 3mo ago> the platform was adopted by 90.2% of enrolled students Then it is "effectiveness", not "efficacy". Prefer simpler and more specific words when possible, to reduce effort for the reader.
- yashthakker 3mo ago[flagged]
- riddlemethat 3mo agoOn a Dartmouth related note, I still can’t believe they abandoned Blitzmail for Office 365. What a loss.
- red_admiral 3mo agoThere's a famous post by Erik Hoel that calls the human version of this Aristocratic Tutoring [1] (Scott Alexander is unconvinced [2]). In the 1980s, a researcher called Benjamin Bloom claimed a z=2.0 (that is 2σ) advantage for a combination of mastery learning (don't move on the the next topic until you've mastered the current one) and 1-on-1 tutoring. Later replications show there is definitely something going on, but the effect size is much lower, for example around z=0.7 in a 2020 paper [3]. I'm still open on AI tutoring, though the Dartmouth results look impressive. Someone please try and replicate this. There's a saying that AI helps the best students get better, and the worst ones get worse. (Anthropic sort-of agrees [4].) It'll be interesting to see how that turns out. [1] https://www.theintrinsicperspective.com/p/why-we-stopped-making-einsteins https://www.theintrinsicperspective.com/p/why-we-stopped-mak... [2] https://www.astralcodexten.com/p/contra-hoel-on-aristocratic-tutoring https://www.astralcodexten.com/p/contra-hoel-on-aristocratic... [3] https://www.nber.org/papers/w27476 https://www.nber.org/papers/w27476 [4] https://www.anthropic.com/research/AI-assistance-coding-skills https://www.anthropic.com/research/AI-assistance-coding-skil...