8 ms·
Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to c
by jjcm 2mo ago
Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... .
It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.
- basch 2mo agoDepends on the definition of friction. If someone is in the Google ecosystem, why would they reach out of it.
- Ardon 2mo agoI already use GCP and Google for work, and getting an API key was so annoying that even I couldn't be bothered after a while of looking around. Maybe things there have improved some, but when I was looking it was a huge runaround.
- beart 2mo agoHmm. My company has an internal portal for generating Gemini API keys. I select a project from a drop down, enter a name, and press okay.
- cameronh90 2mo agoThat may be evidence the built-in Google experience is difficult or confusing.
- MrBuddyCasino 2mo agoI like using 3.5-flash-lite for doing cheap PDF and Image data extraction stuff. I don't think there is a better bang / buck model right now (3.1 is cheaper but a lot worse).
- wahnfrieden 2mo agoProbably cheaper to run a Mac Mini with VisionKit (private APIs if you need bounding rects).
- oh_no 2mo ago5.6 Luna costs far less and benchmarks far better, have you compared for this task?
- MrBuddyCasino 2mo agoUh snap it indeed is cheaper: $0.20 / $1.20 vs $0.30 / $2.50. Gemini is mostly good enough for what I do with it, but the cost savings are interesting. Gemini is still faster though. Not sure how much the benchmarks can be trusted though: https://www.reddit.com/r/GoogleGeminiAI/comments/1vbq5vf/comment/p123hu5/ https://www.reddit.com/r/GoogleGeminiAI/comments/1vbq5vf/com...
- wahnfrieden 2mo agoYou can use fast mode for 2.5x faster and it’ll still be cheaper on output tokens
- piyh 2mo agoSol on Cerebras is going to be expensive AF
- sigmoid10 2mo agoMoving from either frontier intelligence or frontier latency to a single model that does both at the same time is potentially a game changer in certain industries. I can easily see e.g. hedge funds dropping tons of money on this, because it means they can now do the same thing as their competitors, but much faster. That's basically a license to print money.
- alasdair_ 2mo agoWhat would a hedge fund want to do on this exactly? It’s too slow for hft and I’m not sure what they would be doing where ms matter but is not hft.
- cameronh90 2mo agoThere's a lot of trading that isn't proper "HFT", but where speed and latency still matter. Often you'll find this employed more as slippage reduction - i.e you're going to make the trade either way, but making it faster saves you a few bps. I'm not sure what event-based traders are doing now, but back in the day NLP sentiment analysis was all the rage, so I'm assuming they've now incorporated LLMs too.
- jmalicki 2mo agoIt's not like there are only two buckets: 1. HFT doing ass-simple arbitrage where only latency matters 2. More sophisticated slower trading taking in deeper signals Those are two points along a continuum. If you are reacting to an earnings announcement by having an LLM read the earnings release and listen to the call, getting the results a few seconds earlier lets you get your trade in a few seconds earlier. Just because "not HFT" doesn't mean "completely latency insensitive".
- sigmoid10 2mo ago
- krat0sprakhar 2mo agoCan you help me understand how it is hard to get an API key from Google? You just head on over to http://aistudio.google.com/api-keys http://aistudio.google.com/api-keys and create a key... not any different from platform.openai.com? Disclaimer: I work in Google so it might be that this link is not publicly well known
- dogomatic 2mo agoGCP/vertex is a maze
- jjcm 2mo agoDisclaimer that I haven't tried this since January, so things may have changed in the last 7mo, but this was my experience at that time: https://x.com/pwnies/status/2010523020629274723 https://x.com/pwnies/status/2010523020629274723 At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you get with a Google product will route you through 8 different dashboards to set up ACLs before you've hired your 2nd employee.
- qlte 2mo agoFWIW I definitely did not not need to do anything like that to generate a key via AI Studio. It was like three clicks to get the free tier key, later on enabling billing was a few more plus typing in credit card info.
- theplumber 2mo agoIt was 3 clicks because you knew where to look for.
- abirch 2mo agoA pet peeve of mine is that Google doesn't let you search all the options like Apple or Microsoft E.g., I have flashbacks to trying to turn off or on bells in Google Home. A simple search bar would help people trying find API keys, etc.
- bayindirh 2mo agoI use LLMs rarely, and only for digging into subjects which I can't find enough information using search engines. I only tried Claude and Gemini, but Gemini both returns faster and higher quality information which I can use for more targeted digging myself. Google being Google, their models tend to be better at finding, organizing and presenting information, from my experience.
- iainmerrick 2mo agoYep, I use Gemini for this too and it’s great - very fast and high quality. I’d be very willing to try it out as an API, but it’s far too complicated to set up payment, and I don’t want to risk taking a wrong step and being locked out of other Google services. So Anthropic and Mistral get my money instead.
- krychu 2mo ago> but I just don't know what situation I'd reach for 3.7 Flash You reach for it every time you do a Google search
- WarOnPrivacy 2mo ago> You reach for it every time you do a Google search [my self-important Kagi shtick awakens, pokes at it's restraints]
- dzhiurgis 2mo agoIt's much cheaper tho. Junie says Fable is 5-10x more than default model (Gemini 3 Flash Preview).
- tjwebbnorfolk 2mo ago> especially given how hard it is to get an API key from them What does this mean? Anybody can get an API key
- angelmm 2mo agoThe API key you are mentioning is just ridiculous. Onboarding your company or personal account is a trap. I ended up getting assigned to sales guy just to test their Vertex API because I used a company email. Of course, we just used OpenRouter for testing and never touched a Gemini model anymore.