9 ms·
OpenAI GPT-4 vs. Groq Mistral-8x7B
- huqedato 3y agoCan somebody explain why this Grok is more performant than Microsoft infrastructure ? LPU better than TPU/GPU ?
- tosh 3y agoThe Mistral Mixed Expert model has way fewer parameters active during inference and Groq has special purpose hardware (and probably less concurrent demand).
- kkielhofner 3y ago> probably less concurrent demand This is a significant understatement. ChatGPT has an estimated 100m monthly active users. Groq gets featured on HN from time to time but is otherwise almost completely unknown. According to their stats they have done something like 15m requests total since launch. ChatGPT likely does this in hours (or less).
- naiv 3y agoIt's a totally different approach for interference In short: Groq - Ai Chip Microsoft etc. - Nvidia Gpu
- kkielhofner 3y agoLLM performance is about parallelism but also memory bandwidth. Groq delivers this kind of speed by networking many, many chips together with high bandwidth interconnect. Each chip has only 230mb of SRAM[0]. From the linked reference: "In the case of the Mixtral model, Groq had to connect 8 racks of 9 servers each with 8 chips per server. That’s a total of 576 chips to build up the inference unit and serve the Mixtral model." That's eight racks with ~132GB of memory for the model. A single H100 has 80GB and can serve Mixtral without issue (albeit at lower performance). If you consider the requirements for actual real-world inference serving workloads you need to serve multiple models, multiple versions of models, LoRA adapters, sentence embeddings models (for RAG), etc the economics and physical footprint alone get very challenging. It's an interesting approach and clearly very, very fast but I'm curious to see how they do in the market: 1) This analysis uses cloud GPU costs for Nvidia pricing. Cloud providers make significant margin on their GPU instances. If you look at qty 1 retail Nvidia DGX, Lambda Hyperplane, etc and compare it to cloud GPU pricing (inference needs to run 24x7) break even on hardware vs cloud is less than seven months depending on what your costs are for hosting the hardware. 2) Nvidia has incredibly high margins. 3) CUDA. There are some special cases where tokens per second and time to first token are incredibly important (as the article states - real time agents, etc) but overall I think actual real-world production use or deployment of Groq is a pretty challenging proposition. [0] - https://www.semianalysis.com/p/groq-inference-tokenomics-speed-but https://www.semianalysis.com/p/groq-inference-tokenomics-spe...
- malux85 3y agoSorry to be nit-picky but thats the essence of these benchmarks - Mistral putting "N/A" for not available is weird - N/A is not applicable, in every use I have ever seen, and they DONT mean the same thing. I would expect null for not available and N/A for not applicable Impressive inference speed difference though
- mewpmewp2 3y agoI have always known N/A as not available.
- malux85 3y agoCurious, where are you from? If I Google N/A every single hit on the first page is explaining it means "Not applicable" are you from a non-english country? Maybe its cultural?
- selcuka 3y agoThe first entry on Google is Wikipedia [1] for me: > N/A (or sometimes n/a or N.A.) is a common abbreviation in tables and lists for the phrase not applicable, not available, not assessed, or no answer. [1] https://en.wikipedia.org/wiki/N/A https://en.wikipedia.org/wiki/N/A
- malux85 3y agoThats interesting, wikipedia is not on the first page for me, my first hit is Cambridge dict: (and then a bunch of other dicts) - Im flying right now but IP geolocation puts me in the US Meaning of n/a in English written abbreviation for not applicable: used on a form to show that you are not giving the information asked for because the question is not intended for you or your situation: If a question does not apply to you, please put N/A in the box provided. COMMERCE. TIL
- Jensson 3y agoIn a data table "not available" is usually the right word for it, like if you have a list of national statistics then some of the values wont be available due to political reasons etc. But all of those means basically the same thing to the end user, this value isn't there.
- retrac98 3y agoThere are so many applications for LLMs where having a perfect score is much more important than speed, because getting it wrong is so expensive, damaging, or time consuming to resolve for an organisation.
- malux85 3y agoYeah I agree - just an hour ago I was dealing with an LLM that was missing a "not" thus inverting the meaning of a rather important simulation parameter!
- bberrry 3y agoIf you have speed you can generate multiple answers and have another model pick the best one.
- Drakim 3y agoIf I ask an LLM a very complex and specific question 500 times, if it just doesn't know the facts you'll still get the wrong answer 500 times. That's understandable. The real problem is when the AI lies/hallucinates another answer with confidence instead of saying "I don't know".
- deleted 3y ago[deleted]
- helsinkiandrew 3y ago> If I ask an LLM a very complex and specific question 500 times, if it just doesn't know the facts you'll still get the wrong answer 500 times. Think the commenter meant use another model/LLM which could give a different answer, then let them vote on the result. Like "old fashioned AI" did with ensemble learning.
- simion314 3y agoThe problem is asking for facts, LLM are not a database so they know stuff but it is compressed so expect wrong facts, wrong names, dates, wrong anything. We will need an LLM as a front end then it will generate a query to fetch the facts from the internet or a database , then maybe format the facts for your consumption.
- ttrrooppeerr 3y agoA bit off-topic but maybe not? Any words on GPT-5? Is that coming? Or is OpenAI just focusing on the Sora model?
- YetAnotherNick 3y agoThere's no reason for OpenAI to release the model. They have close to 100% market anyways and releasing GPT-5 likely won't increase the total market as it is a incremental leap. And it's a open secret that most other models used GPT-4 synthetic data for training to come close to it. They would likely wait till any model performs better than GPT 4 for the same price
- tosh 3y ago100%? Claude 3 Opus is in the capability ballpark of GPT-4, GPT-3.5 has alternatives that are cheaper (Claude 3 Haiku) or cheaper and work offline (Qwen 1.5, Mixtral, …).
- ZitchDog 3y ago100% market share. A competitor will likely need to be 10x better than ChatGPT in order to get significant market share, not just marginally better in certain scenarios.
- Kostic 3y agoIs Claude 3 Opus generating more profits and taking considerable amount of customers from OpenAI? I'm not seeing that yet. Granted, I'm in Europe (outside of EU) so I can't pay for Opus but I guess that kinda confirms my statement. GPT4 is still a good product and there are no market pressures to release GPT5.
- lewhoo 3y agoThere is reason to release new models if said models would be capable of grabbing a significant portion of job market currently occupied by humans.
- whiplash451 3y ago
- feintruled 3y agoBrave new world, where our machines are sometimes wrong but by gum they are quick about it.
- RUnconcerned 3y agoI too am a big fan of having my computer hallucinate incorrect information.
- harryf 3y ago>> print(“Hello, world!”.ai_reverse()) world, Hello!
- ben_w 3y agoFirst few versions of Swift kept changing how strings work because it's not entirely obvious what most people intend from the nth element of a string. Used to be easy, when it was ASCII. Reverse the bytes of utf-8 and it won't always be valid uft-8. Reverse the code-points, and the Canadian flag gets replaced with the Ascension Island flag.
- samus 3y agoCharacter-level operations are difficult for LLMs. Because of tokenization they don't really "perceive" strings as a list of characters. There are LLMs that ingest bytes, but they are intended to process binary data.
- darthrupert 3y agoYesterday I asked my locally running gpt4all "What model are you running on?" Answer: "I'm running on Toyota Corolla" Which was perhaps the funniest thing I heard that day.
- RUnconcerned 3y agoFinally, something more offensive than parsing HTML with regular expressions: parsing HTML with LLMs.
- AlphaAndOmega0 3y agoI for one am glad I can offload all the regex to LLMs. Powerful? Yes. Human readable for beginners? No.
- cornedor 3y agoWhy tough? To me, it seems more prone to issues (hallucinations, prompt injections etc). It is also slower and more expensive at the same time. I also think it is harder to implement properly, and you need to add way more tests in order to be confident it works.
- okamiueru 3y agoDeterministic? No.
- RUnconcerned 3y agoPersonally when I am parsing structured data I prefer to use parsers that won't hallucinate data but that's just me. Also, don't parse HTML with regular expressions.
- rybosome 3y agoGenerally I agree with your point, but there is some value in a parser that doesn’t have to be updated when the underlying HTML changes. Whether or not this benefit outweighs the significant problems (cost, speed, accuracy and determinism) is up to the use case. For most use cases I can think of, the speed and accuracy of an actual parser would be preferable. However, in situations where one is parsing highly dynamic HTML (eg if each business type had slightly different output, or you are scraping a site which updates the structure frequently and breaks your hand written parser) then this could be worth the accuracy loss.
- tosh 3y agoI initially thought the blog post is about scraping using screenshots and multi-modal llms. Scraping is quite complex by now (front-end JS, deep and irregular nesting, obfuscated html, …).
- bambax 3y agoInteresting post, but the prompt is missing? How do the LLMs generate the keys? It's likely the mistakes could be corrected with a better prompt or a post check? Also, Google SERP page is deterministic (always has the same structure for the same kind of queries), so it would probably be much more effective to use AI to write a parser, and then refine it and use that?
- dns_snek 3y agoFor all the posturing and crypto hate on HN, we're entering a world where it's socially acceptable to use 1000W of computing power and 5 seconds of inference time to parse a tiny HTML fragment which would take microseconds with traditional methods - and people are cheering about it. Time for some self-reflection? That's not very green.
- satisfice 3y agoAND it's not even reliable.
- skc 3y agoOne is actually useful day to day though.
- londons_explore 3y agoWhile energy remains cheap and human minds remain expensive, it always makes sense to use AI to reduce human effort. If one cares about the environment, a carbon cap/tax is what you should campaign for. Then carbon-based energy sources will be curtailled, energy costs will go up, and AI like this will be encouraged to become more energy efficient or other methods used instead.
- osigurdson 3y agoIt is a nice idea in principle but ends up being a political tool and a tariff on goods and services of your own country. A global and corruption free carbon tax might work but that is impossible to achieve.
- londons_explore 3y agoThe only way it's gonna work is if a bunch of countries get together, agree a carbon cap/tax, and then tell other countries that they need to join the scheme if they want to trade goods with the group. One way to combat corruption is to ask an international panel of experts to assess how many extra emissions came from non-official sources in each country and reduce next years cap by that amount. Then countries have an incentive to stamp out corruption.
- wruza 3y agoThe prompt, for those interested. I find it pretty underspecified, but maybe that's the point. For example, "Business operating hours" could be expanded a little, because "Closed - Opens at XX" is still non-processable in both cases. You are an expert in Web Scraping, so you are capable to find the information in HTML and label them accordingly. Please return the final result in JSON. Data to scrape: title: Name of the business type: The business nature like Cafe, Coffee Shop, many others phone: The phone number of the business address: Address of the business, can be a state, country or a full address years_in_business: Number of years since the business started hours: Business operating hours rating: Rating of the business reviews: Number of reviews on the business price: Typical spending on the business description: Extra information that is not mentioned yet in any of the data service_options: Array of shopping options from the business, for example, in store shopping, delivery and many others. It should be in format -> option_name: true is_operating: Whether the business is operating HTML: {html}
- infecto 3y agoThis should be higher up. This whole blog post is mostly worthless because the way they are extracting data is less than optimal. Lower end models do not have the attention to complete tasks like this, GPT4Turbo will generally have the capability. But to have an optimal pipeline you should really be splitting up these tasks into individual units. You extract each attribute you want independently and then combine it back together however you want. Also asking for JSON upfront is equally suboptimal in the whole process. I have high confidence that I could accomplish this task using a lower end model with a high degree of accuracy. Edit: I am not suggesting that an LLM is more optimal than what ever traditional parsing methods they may use, simply the way they are doing it is wrong from an LLM flow.
- crowdyriver 3y agoThere's lots of comments here about how stupid is to parse html using llms. Have you ever had to scrape multiple sites with variadic html?
- samus 3y agoThe example here has HTML with a somewhat fixed format. It would indeed have been better to have samples with different format and aiming for a low error rate. If you are scraping a limited amount of sites, you could for each site ask the LLM for parsing code from some samples, review that, and move on.
- deleted 3y ago[deleted]
- infecto 3y agoThis test is interesting from a general high level metric/test but overall the way they are extracting data using a LLM is suboptimal so I don't think the takeaway means much. You could extract this type of data using a low-end model like 8x7B with a high degree of accuracy.
- samus 3y agoThe better way would be to ask it to generate a program that uses CSS selectors to parse the HTML.
- emporas 3y agoMixtral works very well with json output in my personal experience. Gpt family are excellent of course, and i would bet Claude and Gemini are pretty good. Mixtral however is the smallest of the models and the most efficient. Especially running on Groq's infrastructure it's blazing fast. Some examples i ran on Groq's API, the query was completed in 70ms. Groq has released API libraries for Python and Javascript, i wrote a simple Rust example here, of how to use the API [1]. Groq's API documents how long it takes to generate the tokens for each request. 70ms for a page of document, are well over 100 times faster than GPT, and the fastest of every other capable model. Accounting for internet's latency and some queue that might exist, then the user receives the request in a second, but how fast would this model run locally? Fast enough to generate natural language tokens, generate a synthetic voice, listen again and decode the next request the user might talk to it, all in real time. With a technology like that, why not talk to internet services with just APIs and no web interface at all? Just functions exposed on the internet, take json as an input, validate it, and send the json back to the user? Or every other interface and button around. Why pressing buttons for every electric appliance, and not just talk to the machine using a json schema? Why should users on an internet forum, every time a comment is added, have to press the add comment button, instead of just talking and saying "post it"? Pretty annoying actually. [1] https://github.com/pramatias/groq_test https://github.com/pramatias/groq_test
- imaurer 3y agoGroq will soon support function calling. At that point, you would want to describe your data specification and use function calling to do extraction. Tools such as Pydantic and Instructor are good starting points. I am collecting these approaches and tools here: https://github.com/imaurer/awesome-llm-json https://github.com/imaurer/awesome-llm-json