50 ms·
Phi-3 Technical Report
- anticensor 2y agoThis paper broke ArXiv's HTML generator: https://github.com/arXiv/html_feedback/issues/1090 https://github.com/arXiv/html_feedback/issues/1090
- oersted 2y agoIncredible, rivals Llama 3 8B with 3.8B parameters after less than a week of release. And on LMSYS English, Llama 3 8B is on par with GPT-4 (not GPT-4-Turbo), as well as Mistral-Large. Source: https://chat.lmsys.org/?leaderboard https://chat.lmsys.org/?leaderboard (select English in the dropdown) So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? Kinda? Wild. (I'm sure there's a lot of nuance to it, for one these benchmarks are not so hard to game, we'll see how the dust settles, but still...) Phi-3-mini 3.8b: 71.2 Phi-3-small 7b: 74.9 Phi-3-medium 14b: 78.2 Phi-2 2.7b: 58.8 Mistral 7b: 61.0 Gemma 7b: 62.0 Llama-3-In 8b: 68.0 Mixtral 8x7b: 69.9 GPT-3.5 1106: 75.3 (these are averages across all tasks for each model, but looking at individual scores shows a similar picture)
- crakenzak 2y agoCan’t wait to see some Phi-3 fine tunes! Will be testing this out locally, such a small model that I can run it without quantization. Feels incredible to be living in a time with such neck breaking innovations. What are chances we’ll have a <100B parameter GPT4/Claude Opus model in the next 5 years?
- stavros 2y agoIs it released?
- Deverauxi 2y ago5 years? 5 years is a millennia these days. We’ll have small local models beating gpt-4/Claude opus in 2024. We already have sub 100b models trading blows with former gpt-4 models, and the future is racing toward us. All these little breakthroughs are piling up.
- refulgentis 2y agoAbsolutely not on the first one. Not even close.
- ashirviskas 2y agoWhy not? There's still 7 months left for breakthroughs.
- refulgentis 2y agoSmall leaves wiggle room, but it's extremely unlikely trad small, <= 7B, will get there this year even on these evals. UX matching is a whole different matter and needs a lot of work: Worked heavily with Llama 8B over last days, and Phi 3 today, and the Q+A benchmarks don't tell the full story. Ex. It's nigh impossible to get Llama _70_B to answer in JSON; when Phi sees RAG from search it goes off inventing new RAG material and a new question.
- bugglebeetle 2y agoWe already do. It’s called LLama 3 70B Instruct.
- vitorgrs 2y agoLlama 3 is awful in non-English. 95% of their training data is in English.... GPT is still the king when talking about multiple languages/knowledge.
- nl 2y ago> What are chances we’ll have a <100B parameter GPT4/Claude Opus model in the next 5 years? In 5 years time we'll have adaptive compute and the idea of talking about the parameter count of a model will seem as quaint as talking about the cylinder capacity of a jet engine.
- regularfry 2y agoIt feels like it's going to be closer than that. People always forget that GPT4 and Opus have the advantage of behind-the-curtain tool use that you just can't see, so you don't know how much of a knowledge or reasoning leg-up they're getting from their internal tooling ecosystem. They're not really directly comparable to a raw LLM downloaded from HF. What we need is a standardised open harness for open source LLMs to sit in that gives them both access to tools and the ability to write their own, and that's (comparatively speaking) a much easier job than training up another raw frontier LLM: it's just code, and they can write a lot of it.
- moralestapia 2y ago>And on LMSYS English, Llama 3 8B is well above GPT-4 Source?
- oersted 2y agoRight thanks for the reminder, I added it
- moralestapia 2y agoThanks, I don't see them being "well above GPT-4", merely 1 point? Also, no idea why one would want to exclude GPT-4-Turbo, the flagship "GPT-4" model, but w/e. I also don't think they "beat Llama 3 8B"; their own abstract says "rivals that of models such as Mixtral 8x7B and GPT-3.5", "rivals" not even "beats". Great model, but let's not overplay it.
- oersted 2y agoIn the English category: GPT-4-0314 (ELO 1166), Llama 3 8B Instruct (ELO 1161), Mistral-Large-2402 (ELO 1151), GPT-4-0613 (ELO 1148). You are right, I toned down the language, I got a bit overexcited, and I missed the difference in the versions of GPT-4. And LMSYS is a subjective benchmark for what users prefer, which I'm sure has weird inherent biases. It's just that any signal of an 3.8B model being anywhere in the vicinity of GPT-4 is huge.
- moralestapia 2y agoYeah, GPT3.5, in a phone, at ~1,000 tokens/sec ... nice!
- mlyle 2y ago> at ~1,000 tokens/sec 12 tokens per second.
- 2y ago
- ignoramous 2y ago> Phi-3-mini 3.8b: 71.2 Per the paper, phi3-mini (which is english-only) quantised to 4bit uses 1.8gb RAM and outputs 1212 tokens/sec (correction: 12 tokens/sec) on iOS. A model on par with GPT-3.5 running on phones! (weights haven't been released, though)
- coder543 2y ago> (weights haven't been released, though) Phi-1, Phi-1.5, and Phi-2 have all had their weights released, and those weights are available under the MIT License. Hopefully Microsoft will continue that trend with Phi-3. > outputs 1212 tokens/sec on iOS I think you meant "12 tokens/sec", which is still nice, just a little less exciting than a kilotoken/sec.
- jph00 2y agoWeights will be realised tomorrow, according to one of the tech report authors on Twitter.
- ignoramous 2y ago> you meant 12 tokens/sec Thanks! The HTML version on archive.is has messed up markup and shows 1212 instead: https://archive.is/Ndox6 https://archive.is/Ndox6
- intellectronica 2y agoWeights are coming tomorrow.
- homarp 2y agotomorrow is now: https://huggingface.co/microsoft/Phi-3-mini-4k-instruct https://huggingface.co/microsoft/Phi-3-mini-4k-instruct
- jxy 2y agoThis inductive logic is way overblown. > Incredible, beat Llama 3 8B with 3.8B parameters after less than a week of release. Judging by a single benchmark? Without even trying it out with real world usage? > And on LMSYS English, Llama 3 8B is on par with GPT-4 (not GPT-4-Turbo), as well as Mistral-Large. Any potential caveat in such a leaderboard not withstanding, on that leaderboard alone, there is a huge gap between llama 3 8B and Mistral-Large, let alone any of the GPT-4. By the way, for beating benchmark, "Pretraining on the Test Set Is All You Need"
- oersted 2y agoIt's easy to miss: select English in the dropdown. The scores are quite different in Overall and in English for LMSYS. As I've stated in other comments, yeah... Agreed, I'm stretching it a bit. It's just that any indication of a 3.8B model being in the vicinity of GPT-4 is huge. I'm sure that when things are properly measured by third-parties it will show a more sober picture. But still, with good fine-tunes, we'll probably get close. It's a very significant demonstration of what could be possible soon.
- saretup 2y agoFirstly, English is a highly subjective category. Secondly, Llama 3 usually adds first sentences like ‘What a unique question!’ or ‘What an insightful thought’, which might make people like it more than the competition because of the pandering. While Llama 3 is singular in terms of size to quality ratio, calling the 8B model close to GPT4 would be an overstretch.
- YetAnotherNick 2y agoYes, I don't know how people don't realize how much cheap tricks works in Chatbot Arena. A single base model produces 100s of ELO difference depending on the way it is tuned. And on most cases, instruction tuning heavily slightly even decreases reasoning ability on standard benchmark. You can see base model scores better in MMLU/ARC most of the times in huggingface leaderboard. Even GPT-4-1106 seems to only sounds better than GPT-4-0613 and works for wider range of prompt. But in a well defined prompt and follow up questions I don't think there is an improvement in reasoning.
- zone411 2y ago> So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones? No, we don't. LMsys is just one, very flawed benchmark.
- oersted 2y agoAgreed, but it's wild that even one benchmark shows this. Based on what we knew just a few months ago, these models should be so far from each other in every benchmark.
- ukuina 2y agoWhy is LMsys flawed? Many people treat LMsys as gospel because it's the only large-scale, up-to-date qualitative benchmark. All the numeric benchmarks seem to miss real-world applicability.
- viraptor 2y agoOn par in some categories. Phi was intended for reasoning, not storing data, due to small size. I mean, it's still great, but the smaller it gets, the more facts from outside of the prompts context will not be known at all.
- candiodari 2y agoI wonder if that's a positive or negative. How does it affect hallucinations?
- viraptor 2y agoIt depends what you want to do. If you want a chat bot that can replace most Google queries, you want as much learned data as possible and the whole Wikipedia consumed. If you want a RAG style system, you want good reasoning about the context and minimal-or-no references to extra information. It's neither positive nor negative without a specific use case.
- alecco 2y agoAt a glance, it looks like Phi-3 was trained on an English only, STEM-strong dataset. See how they are not as strong in HumanEval, Trivia, etc. But of course it's very good.
- karmasimida 2y agoWhere did you get this from? > So we now have an open-source LLM approximately equivalent in quality to GPT-4 that can run on phones No, not even close ... Even Gemini has huge UX gap comparing to GPT4/Opus, 8B I won't even attempt this argument.
- infecto 2y ago"But still"? Lets be realistic, all of these benchmark scores are absolute garbage. Yes, the open source community is making great strides, they are getting closer but the gap is still wide when comparing to commercially available models.
- blackeyeblitzar 2y agoIt’s not open source, but is open weight - like distributing a precompiled executable. In particular what makes it open weights rather than just weights available is that it is licensed using an OSI approved license (MIT) rather than a restricted proprietary license. I really wish these companies would release the training source, evaluation suites, and code used to curate/filter training data (since safety efforts can lead to biases). Ideally they would also share the training data but that may not be fully possible due to licensing.
- simonw 2y agoI'm getting a bit skeptical of MMLU at this point. As far as I can tell it's a set of multiple choice questions that hasn't been updated since 2020. We have to trust the model providers not to deliberately or accidentally train on it for those scores to be useful.
- minimaxir 2y agoAt the least, there's multiple benchmarks noted in the paper (21!) and the results are consistent across all of them. I'd trust Microsoft to do decontamination testing, although the paper doesn't explicitly mention it other than "The prompts and number of shots are part of a Microsoft internal tool to evaluate language models, and in particular we did no optimization to the pipeline for the phi-3 models."
- brcmthrowaway 2y agoIf I was Apple I'd be quaking in my boots. They are getting too far behind to ever catch up. Nokia in 2010 vibes.
- esafak 2y agoDid they ever claim to be a powerhouse in foundation models? Did your MacBook or iPhone become obsolete or stop working? They use the models, they don't release them because they don't hoard data.
- moralestapia 2y agoI don't recall Nokia being a 3 trillion dollar company. Your vibes may vary, though.
- thoughtegting 2y agoIf I were apple, I would be developing something in total secrecy and then release something ahead of the rest of competition when people least expect it. very big ifs but siri can be updated everywhere overnight and I dont see them rushing into anything like this
- golergka 2y agoIf I were apple, I would just buy one of the major LLM companies. They have the cash.
- bingbingbing777 2y agoThey've been buying AI companies and have nothing to show for it.
- hackerlight 2y agoLess tokens than Llama 3 (3.3T vs 15T) yet better outcome. No doubt more information dense training data. The interesting thing is the use of synthetic data which they don't talk about.
- minimaxir 2y agoYes, "chinchilla optimal" is a meme, but 15T might turn out to be too many tokens.
- wrsh07 2y agoMy understanding from this tweet thread [1] is that chinchilla probably overspecified some of the hyperparameters to the model tl;dr I'm looking forward to having lots of models (ideally models) trained with a wide range of parameters to narrow down "what is actually optimal" I think there is an interesting tradeoff of data quality and data volume, though (Eg if we train with the highest quality 10% of our data, does the model improve if we use the other 90%? What if we increase our data size by 10x?) [1] https://twitter.com/tamaybes/status/1780639257389904013 https://twitter.com/tamaybes/status/1780639257389904013
- vessenes 2y agoActually the original Phi papers did talk about their synthetic data strategy, and it's very cool -- essentially invert high quality textbook text using GPT-4 to create prompts, where the textbooks supply the answers. There may be more undisclosed, but it remains in my mind as one of the best ideas of the last twelve months -- so smart, and interesting, and apparently, it works well.
- xarope 2y agoperhaps that's the best path forward? Text and reference books (hopefully unbiased) for answers, and web scraped data for conversational tone.
- astrange 2y agoI feel like literal dictionaries would make good training data; wonder if any of them have done that. LLMs are good at faking so it's hard to tell by asking them.
- blackoil 2y agoHas anyone used these/similar with fine tune and RAG? How is the performance over a narrow domain for simple queries? Is it good enough for say an informational chat bot?
- modeless 2y agoEveryone needs to take these benchmark numbers with a big grain of salt. According to what I've read, Phi-2 was much worse than its benchmark numbers suggested. This model follows the same training strategy. Nobody should be assuming these numbers will translate directly into a high ranking on the LMSYS leaderboard, or usefulness in everyday tasks. Let's not dethrone Llama 3 until some real world testing can be done. That said, I don't think it's impossible for a small model to be very good. I see their "synthetic data" as essentially a way of distilling GPT-4 into smaller models. It would be exciting if a large fraction of the performance of huge models could be transferred to small ones! If true, then Chinchilla-optimal training could make sense again, as you could optimally train a ginormous model and then distill it afterward for efficient inference.
- refulgentis 2y agoPhi-2 wasn't chat/instruct tuned, so it didn't act good in chat, it was a base model. But the benchmark #s were real.
- irjustin 2y agoI'm pretty naive so please forgive it's a stupid question. To me, what the parent comment is saying is that even though the benchmarks are cool, it's not super helpful to the every day person. Because if you can't chat with it very well (even for a narrow context) what utility does it have with great benchmarks?
- svnt 2y agoBoth are saying the same thing: in order for the base model that is phi to perform well as a chat agent, it would need to be tuned for that purpose before its benchmark results could have real-world value.
- imjonse 2y agoFrom this report. Phi-2 was not instruct tuned indeed. "Our models went through post-training with both supervised instruction fine-tuning, and preference tuning with DPO. We have worked on generating and curating various instruction and preference data. This has improved the model chat capabilities, robustness, as well as its safety."
- visarga 2y agoThis shows the power of synthetic content - 3.3 trillion tokens! This approach can make a model even smaller and more efficient than organic text training, and it will not be able to regurgitate NYT articles because it hasn't seen any of them. This is how copyright infringement claims can be placated.
- ur-whale 2y agoThat's a whole lot of Zhangs!
- mythz 2y agoI'll believe it till I try it for myself, Phi-2 was the clear worst of the 20 LLMs we evaluated (was also smallest so was expected). But it was slow for its size, generated the longest responses with the most hallucinations, as well as generating the most empty responses. It was also the model ranked with the lowest quality answers.
- abidlabs 2y agoHugging Face Paper Page and Discussion: https://huggingface.co/papers/2404.14219 https://huggingface.co/papers/2404.14219
- whereismyacc 2y agothey're just spamming "weights or it didn't happen" i mean, fair
- smartmic 2y agoHm, roundabout 84 authors of one "scientific" paper. I wonder if this says something about (a) the quality of its content, (b) the path were academic (?) paper publishing goes to, (c) nothing at all, or (d), something entirely else.
- a_bonobo 2y agoI have been on far larger author lists :) There's probably a whole team for the training data generation and assessment, a whole team for the safety assessment (section 4), that stuff adds up.
- lysecret 2y agoJust means you need a big machine and a lot of capital to make advancement. Take a look at any paper coming out of cern.
- samus 2y agoIt's a tech report. Fair enough to include the whole lab.
- 0cf8612b2e1e 2y agoYou should see physics. Stuff involving the large hadron collider can be pages of authors. It costs so little to share the credit if someone was an asset.
- maximsicora 2y agoinsane
- Havoc 2y agoBoth precious phi have been epic letdowns when I actually tried them myself so quite low confidence in this being reflective of real world. Will try it anyway though
- m3kw9 2y agoPhi-2 was useless for practical purposes except if you want to show your friends that it can write a poem, llama3 8b was slightly better but is still same category, it’s complete trash with coding vs gpt4. Llama3 400b “iS OPen SoURce!” But no you will need to pay to access because most one can not practically afford an A100 and set it up properly. What I’m trying to say is that user experience is now as key as the model smarts and these barely touching gpt4 models cannot beat OpenAI right now as a whole package.
- azinman2 2y agoI just tried to give gpt4 a scrape of a menu page and asked it to reformat it to csv. It hallucinated several parts and missed others. Llama3-70b hasn’t done that. So far it’s been more reliable. You can run a quantized version on consumer hardware or pay significantly less ($10 vs $1 in, $30 vs $1 out) on hosted platforms.
- pkoiralap 2y agoThey have started putting some models in huggingface: https://huggingface.co/collections/microsoft/phi-3-6626e15e9585a200d2d761e3 https://huggingface.co/collections/microsoft/phi-3-6626e15e9...
- minimaxir 2y agoAnd with a MIT license!
- Patrick_Devine 2y agoAnd of course if you want to try it out locally, `ollama run phi3`.
- homarp 2y agothe weights have been released, 4k https://huggingface.co/microsoft/Phi-3-mini-4k-instruct https://huggingface.co/microsoft/Phi-3-mini-4k-instruct and 128k context https://huggingface.co/microsoft/Phi-3-mini-128k-instruct https://huggingface.co/microsoft/Phi-3-mini-128k-instruct
- ein0p 2y agoTried it: as soon as you ask something outside the head of the likely training data distribution it starts hallucinating like crazy. This isn’t surprising to me as a researcher: you need the associative memories of a larger model to cover the tail with at least something. That said, it’ll likely work well at specific narrow tasks once fine tuned. Just don’t expect it to really “beat GPT-3.5” at the general chat use case
- deleted 2y ago[deleted]