28 ms·
Amazon Nova
- scbenet 2y agoTechnical report is available here https://www.amazon.science/publications/the-amazon-nova-family-of-models-technical-report-and-model-card https://www.amazon.science/publications/the-amazon-nova-fami...
- kajecounterhack 2y agoTL;DR comparison of models vs frontier models on public benchmarks here https://imgur.com/a/CKMIhmm https://imgur.com/a/CKMIhmm
- brokensegue 2y agoSo looks like they are trying to win on speed over raw metric performance
- SparkyMcUnicorn 2y agoThis doesn't include all the benchmarks. The one that really stands out is GroundUI-1K, where it beats the competition by 46%. Nova Pro looks like it could be a SOTA-comparable model at a lower price point.
- oblio 2y agoSOTA?
- camel_Snake 2y ago"State of the Art", if that's what you were asking.
- maeil 2y agoJust means it's better at one specific task than the others, which has always been the case. For each of Sonnet, GPT and Gemini I can readily name a task they are individually the best at. At the same time the consensus that Sonnet 3.5 is overall the currently strongest model remains correct, and that's what most people care about. Additionally most people do tasks that all of the models perform similarly at, or they can't be bothered to optimize every task by using the best model for that one task. Which makes sense since not a single cloud provider has all three of them. Now this one will likely be AWS-exclusive too.
- retinaros 2y agoin the berkeley function calling it is similar than 4-o for multi turn while being way faster
- int_19h 2y agoBenchmarks are way too easy to game. There's no shortage of models that "beat GPT-4" according to some benchmark or another, that are obviously nowhere even close when you try them on novel tasks.
- attentive 2y agoon https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/ Nova Pro is on par with Yi Coder 9B Chat. Which is not very inspiring.
- baxtr 2y agoAs a side comment: the sound quality of the auto generated voice clip is really poor. No match for Google's NotebookLM podcasts.
- wenc 2y agoThe autogenerated voice is Amazon Polly which is an old AWS speech synthesis service which doesn’t use the latest technology. It’s irrelevant to the article, which is about Nova.
- griomnib 2y agoIf you haven’t seen it this may be the best use of ai-podcast I’ve seen: https://youtu.be/gfr4BP4V1R8 https://youtu.be/gfr4BP4V1R8
- bongodongobob 2y agoThis is goddamn hilarious, thank you.
- griomnib 2y agoLet us all thank the YouTube God.
- xnx 2y agoMore options/competition is good. When will we see it on https://lmarena.ai/ https://lmarena.ai/ ?
- glomgril 2y agolooks like it's there now
- teilo 2y agoSo that's what I missed at the keynote.
- htrp 2y agoNo parameter counts?
- HarHarVeryFunny 2y agoSince Amazon are building their own frontier models, what's the point of their relationship with Anthropic ?
- tokioyoyo 2y agoIf you play all sides, you’ll always come on top.
- worldsayshi 2y agoYeah Copilot includes Claude now.
- cdchn 2y agoThis is Amazon's core e-commerce business model but for AI. You sell everybody else's stuff and also offer an Amazon Basics version.
- tinyhouse 2y agoI can only guess. 1. A company the size of Amazon has enough resources and unique internal data no one else has access to that it makes sense for them to build their own models. Even if it's only for internal use 2. Amazon cannot beat Anthropic at this game. They are far a head of them in terms of performance and adoption. Building these models in-house doesn't mean it's a bad idea to also invest in Anthropic
- PartiallyTyped 2y agoAlso not putting all of your eggs in one basket.
- blackeyeblitzar 2y agoCommoditizing complements
- jonathaneunice 2y agoDifferent models have different strengths and weaknesses, especially here in the early days when models and their capabilities progress several times per year. The apps, programs, and systems based on models need to know how to exploit their specific strengths and weaknesses. So they are not infinitely interchangeable. Over time some of that differentiation will erode, but it will probably take years. AWS having customers using its own model probably improves AWS's margins, but having multiple models available (e.g. Anthropic's) improves their ability to capture market share. To date, AWS's efforts (e.g. Q, CodeWhisperer) have not met with universal praise. So for at least for the present, it makes sense to bring customers to AWS to "do AI" whether they're using AWS's models or someone else's.
- andrewstuart 2y agoIt's not clear what the use cases are for this, who is it aimed at.
- dvh 2y agoShareholders?
- christhecaribou 2y agoThe real “customers”.
- mystcb 2y agoI'd say, people that need it. Which could be the same for all the other models out there. To create one model that is great at everything is probably a pipedream. Much like creating a multi-tool that can do everything- but can it? I wouldn't trust a multi-tool to take a wheel nut off a wheel, but I would find it useful if I suddenly needed a cross-head screw taken out of something. But then I also have a specific crosshead screwdriver that is good at just taking out cross-head screws. Use the right tool for the right reason. In this case, there maybe a legal reason why someone might need to use it. It might be that this version of a model can create something better that another model can't. It might be that for cost reasons you are within AWS, that it makes sense to use a model at the cheaper cost than say something else. So yeah, I am sure it will be great for some people, and terrible for others... just the way things go!
- dgfitz 2y ago> I'd say, people that need it. Nobody needs Reddit hallucinations about programming.
- petesergeant 2y agohttps://artificialanalysis.ai/leaderboards/models https://artificialanalysis.ai/leaderboards/models seems to suggest Nova Lite is half the price of 4o-mini, and a chunk faster too, with a bit of quality drop-off. I have no loyalty to OpenAI, if it does as well as 4o-mini in the eval suite, I'll switch. I was hoping "Gemini 1.5 Flash (Sep)" would pass muster for similar reasons, but it didn't.
- xendo 2y agoSome independent latency and quality evaluations already available at https://artificialanalysis.ai/ https://artificialanalysis.ai/ Looks to be cheap and fast.
- blackeyeblitzar 2y agoIt would be nice if this was a truly open source model like OLMo: https://venturebeat.com/ai/truly-open-source-llm-from-ai2-to-drive-critical-shift-in-ai-development/ https://venturebeat.com/ai/truly-open-source-llm-from-ai2-to...
- sourcepluck 2y agoIs it narrowly open source, or somewhat open source, in some way? Thanks for that link, anyway!
- blackeyeblitzar 2y agoAs far as I can tell Amazon’s nova is fully closed source. Maybe because their goal is to get you to pay them for hosting.
- mikesurowiec 2y agoA rough idea of the price differences... Per 1k tokens Input | Output Amazon Nova Micro: $0.000035 | $0.00014 Amazon Nova Lite: $0.00006 | $0.00024 Amazon Nova Pro: $0.0008 | $0.0032 Claude 3.5 Sonnet: $0.003 | $0.015 Claude 3.5 Haiku: $0.0008 | $0.0004 Claude 3 Opus: $0.015 | $0.075 Source: AWS Bedrock Pricing https://aws.amazon.com/bedrock/pricing/ https://aws.amazon.com/bedrock/pricing/
- Bilal_io 2y agoYou have added another zero for Haiku, its output cost is $0.004
- indigodaddy 2y agoThanks that had confused me when I compared same to Nova Pro
- mikesurowiec 2y agoYou're absolutely right, apologies!
- warkdarrior 2y agoEyeballing it, Nova seems to be 1.5 order of magnitude cheaper than Claude, at all model sizes.
- holub008 2y agoHas anyone found TPM/RPM limits on Nova? Either they aren't limited, or the quotas haven't been published yet: https://docs.aws.amazon.com/general/latest/gr/bedrock.html#limits_bedrock https://docs.aws.amazon.com/general/latest/gr/bedrock.html#l...
- tmpz22 2y agoMaybe they want to gauge demand for a bit first?
- Tepix 2y ago
- indigodaddy 2y agoUnfortunate that this seems to be inextricably tied to Amazon Bedrock though in order to use it..
- deleted 2y ago[deleted]
- jklinger410 2y agoIt's really amusing how bad Amazon is at writing and designing UI. For a company of their size and scope it's practically unforgivable. But they always get away with it.
- smt88 2y agoYou say they "get away with it," but it makes more sense to conclude that UI design has a lot lower ROI than we assume it does as users.
- wilg 2y agoOr that design instincts are backwards
- wavemode 2y agoYou can't conclude that. At best, you can conclude that outdated product design doesn't always ruin a business (clearly). But you can't conclude the inverse (that investing in modern product design doesn't ever help a business).
- handfuloflight 2y agoThat's a great point. Further, there are many sizeable businesses built on top of AWS where they deliver the abstractions with compression that earns them their margin. Case in point: tell me, from the point of view of the user, how many steps it takes to deploy a NextJS/React ecosystem website with Vercel and with AWS, start to finish.
- rrrrrrrrrrrryan 2y agoI think they have plenty of competition in the cloud computing space. It seems fair to say that their strategy of de-prioritizing UI/UX in favor of getting features out the door more quickly and cheaply has benefitted them. However, I don't think it's fair to say that this trade-off always wins out. Rather, they've carved out their own ecological niche and, for now, they're exploiting it well.
- 2y ago
- zapnuk 2y agoThey missed a big opportunity by not offering eu-hosted versions. Thats a big thing for complience. All LLM-providers reserve the right to save (up to 30days) and inspect/check prompts for their own complience. However, this means that company data is potentionally sotred out-of-cloud. This is already problematic, even more so when the storage location is outside the EU.
- ygouzerh 2y agoThey might not have enough GPUs datacenters in Europe
- Tepix 2y agoI'm not sure if hosting it in the EU will do any good for Amazon, there's still the US CLOUD Act: It doesn't really matter where the data is located.
- physicsguy 2y agoIt makes a really big difference for anyone doing business in Europe though. Legally we're only allowed to use text-embeddings-3-large at work because Azure don't host text-embeddings-3-small within a European region.
- Tepix 2y agoPerhaps your legal department is ignoring the US CLOUD act?
- diggan 2y ago> The model processes inputs up to 300K tokens in length [...] up to 30 minutes of video in a single request. I wonder how fast it "glances" an entire 30 minute video and takes until the first returned token. Anyone wager a guess?
- potlee 2y ago> The Nova family of models were trained on Amazon’s custom Trainium1 (TRN1) chips,10 NVidia A100 (P4d instances), and H100 (P5 instances) accelerators. Working with AWS SageMaker, we stood up NVidia GPU and TRN1 clusters and ran parallel trainings to ensure model performance parity Does this mean they trained multiple copies of the models?
- glomgril 2y agoModels like this are experimentally pretrained or tuned hundreds of times over many months to optimize the datamix, hyperparams, architecture, etc. When they say "ran parallel trainings" they are probably referring to parity tests that were performed along the way (possibly also for the final training runs). Different hardware means different lower-level libraries, which can introduce unanticipated differences. Good to know what they are so they can be ironed out. Part of it could also be that they'd prefer to move all operations to the in-house trn chips, but don't have full confidence in the hardware yet. Def ambiguous though. In general reporting of infra characteristics for LLM training is left pretty vague in most reports I've seen.
- jmward01 2y agoNo audio support: The models are currently trained to process and understand video content solely based on the visual information in the video. They do not possess the capability to analyze or comprehend any audio components that are present in the video. This is blowing my mind. gemini-1.5-flash accidentally knows how to transcribe amazingly well but it is -very- hard to figure out how to use it well and now Amazon comes out with a gemini flash like model and it explicitly ignores audio. It is so clear that multi-modal audio would be easy for these models but it is like they are purposefully holding back releasing it/supporting it. This has to be a strategic decision to not attach audio. Probably because the margins on ASR are too high to strip with a cheap LLM. I can only hope Meta will drop a mult-modal audio model to force this soon.
- xendo 2y agoThey also announced speech to speech and any to any models for early next year. I think you are underestimating the effort required to release 5 competitive models at the same time.
- plumeria 2y agoIs Gemini better than Whisper for transcribing?
- jmward01 2y ago'better' is always a loaded term with ASR. Gemini 1.5 flash can transcribe for 0.01/hour of audio and gives strong results. If you want timing and speaker info you need to use the previous version and a -lot- of tweaking of the prompt or else it will hallucinate the timing info. Give it a try. It may be a lot better for your use case.
- adt 2y agoParam estimates etc: https://lifearchitect.ai/olympus/ https://lifearchitect.ai/olympus/
- zacharycohn 2y agoI really wish they would left-justify instead of center-justify the pricing information so I'm not sitting here counting zeroes and trying to figure out how they all line up.
- Super_Jambo 2y agoNo embedding endpoints?
- lukev 2y agoThis is a digression, but I really wish Amazon would be more normal in their product descriptions. Amazon is rapidly developing its own jargon such that you need to understand how Amazon talks about things (and its existing product lineup) before you can understand half of what they're saying about a new thing. The way they describe their products seems almost designed to obfuscate what they really do. Every time they introduce something new, you have to click through several pages of announcements and docs just to ascertain what something actually is (an API, a new type of compute platform, a managed SaaS product?)
- Miraste 2y agoThat may be generally true, but the linked page says Nova is a series of foundation models in the first sentence.
- lukev 2y agoYeah but even then they won't describe it using the same sort of language that everyone else developing these things does. How many parameters? What kind of corpus was it trained on? MoE, single model, or something else? Will the weights be available? It doesn't even use the words "LLM", "multimodal" or "transformer" which are clearly the most relevant terms here... "foundation model" isn't wrong but it's also the most abstract way to describe it.
- meta_x_ai 2y agoNone of those matters (except multimodal). If you are running a business, the only thing that matters is a) How does it perform on my set of evals b) What is the cost/latency of serving it to my consumers. It shouldn't matter to me how many parameters, corpus it is trained on, whether it's LLM or Transformer or something else
- marcosdumay 2y ago> How does it perform on my set of evals What kinds of eval? Personally, I have no idea what kind of data you can throw at a "foundation model" and what kind of response you will get. The only thing it says is that there's machine learning involved... Once you get enough context to understand it's not a spin-off of a TV series.
- TheAceOfHearts 2y agoThey really should've tried to generate better video examples, those two videos that they show don't seem that impressive when you consider the amount of resources available to AWS. Like what even is the point of this? It's just generating more filler content without any substance. Maybe we'll reach the point where video generation gets outrageously good and I'll be proven wrong, but right now it seems really disappointing. Right now when I see obviously AI generated images for book covers I take that as a signal of low quality. If AI generated videos continue to look this bad I think that'll also be a clear signal of low quality products.
- ndr_ 2y agoSetting up AWS so you can try it via Amazon Bedrock API is a hassle, so I made a step-by-step guide: https://ndurner.github.io/amazon-nova https://ndurner.github.io/amazon-nova. It's 14+ steps!
- simonw 2y agoThank you!
- teruakohatu 2y agoThanks for that. Are there any proxies that can communicate with bedrock and serve it via a OpenAI style api?
- moduspol 2y agoYou'd have to deploy it yourself, but there's this: https://github.com/aws-samples/bedrock-access-gateway https://github.com/aws-samples/bedrock-access-gateway
- teruakohatu 2y agoThanks. That is quite a heavy stack!
- popinman322 2y agoTry LiteLLM; their core LLM proxy is open source. As an added bonus it also supports other major providers.
- OJFord 2y agoYour 14 steps appear to be 'create an IAM user'..?
- Spivak 2y agoIf you're already in the AWS ecosystem or have worked in it, it's no problem. If you're used to "make OpenAI account, add credit card, copy/paste API key" it can be a bit daunting.
- smallnix 2y agoDo these work with the bedrock converse API?
- dheerkt 2y agoyeah converse api supports all models on bedrock, or atleast all the text2text ones
- mrg3_2013 2y agoDOA When marketing talks about price delta and not quality of the output, it is DOA. For LLMs, quality is a more important metric and Nova would always try to play catch with the leaderboard forever.
- xnx 2y agoMaybe. The major models seem to be about tied in terms of quality right now, so cost and ease of use (e.g. you already have an AWS account set up for billing) could be a differentiator.
- mrg3_2013 2y agoUsing LLMs via Bedrock is 10x more painful than using direct APIs. I could see cost consolidation via cloud marketplace a play - but I don't see Amazon's own LLM initiatives ever taking off. They should just lose those shops and buy one of the frontier models (while it is still cheap)
- int_19h 2y agoThe major models are not tied in terms of quality. GPT-4 and GPT-o1 still beat everyone else by a significant margin on tasks that require in-depth reasoning. There's a reason why people just don't go for the cheapest option, whatever the benchmarks say.
- xnx 2y ago> GPT-4 and GPT-o1 still beat everyone else by a significant margin on tasks that require in-depth reasoning I haven't seen examples of this. Do you know where I could find some?
- int_19h 2y agoHere's a fairly simple test that I throw at any model that claims to be "GPT-4 level": https://news.ycombinator.com/item?id=42262661 https://news.ycombinator.com/item?id=42262661 For more complicated stuff, I did some experiments using LLMs to drive high-level AI decisions in video games. Basically, it gets a data schema and a question like "what do you do next?", and can query the schema to retrieve the info that it thinks it needs to give the best answer to that. GPT-4 and GPT-o1 especially are consistently the best performers there, both in terms of richness of queries they produce, and how they make use of them. There's also a bunch of interesting examples along the same lines here: https://github.com/cpldcpu/MisguidedAttention https://github.com/cpldcpu/MisguidedAttention. Although I should note that even top OpenAI models have troubles with much of this stuff. https://github.com/fairydreaming/farel-bench https://github.com/fairydreaming/farel-bench is another interesting benchmark because it's so simple, and yet look at the number disparity in that last column! It's easy to scale, too. Unfortunately, we're still at the point in this game where even seemingly trivial and unrelated minor changes in the prompt (e.g. slightly rewording it, and even capitalization in some cases) can have large effect on quality of output, which IMO is a tell-tale sign when the model is really operating in a "stochastic parrot" mode more so than any kind of actual reasoning. Thus benchmarks can be used as a way to screen out the poorly performing models, but they cannot reliably predict how well a model will actually do what you need it to do.
- m3kw9 2y agoUsing Amazon or google cloud api and forgot about it? Surprise bill in a few months.
- astoilkov 2y agoAny ideas on how to use the new models through JavaScript in the browser or Node.js?
- siquick 2y agoIs there any difference in latency when calling models via Bedrock vs calling the providers APIs directly?