19 ms·
Claude 3.5 Sonnet
- Filligree 2y agoI’ve been wanting to test it, but the API console wants me to fill in a VAT field. Not sure how to get one of those as an individual.
- BoredPositron 2y agoIt's really bad ux but you can just leave it blank as a private user.
- wraptile 2y agois it though? blank optional fields have been the standard since forever.
- BoredPositron 2y agoThe whole organization wording throws a lot of non-native people of. Combine it with the vat field and you have bad ux.
- 8n4vidtmkvmk 2y agoA lot of these guys ask for company name and it often is required. I'm not a company. Just a dude trying to evaluate your thing for my hobby projects.
- BoredPositron 2y agoYeah, that’s why the organisation wording is unfortunate. It’s meant more like a group name under which you will have your projects and not corporation name.
- Davidzheng 2y agoStill not as good at math as gpt4o from benchmarks and also my experience
- Legend2440 2y agoThat's ok, I don't need my language model to do math. My calculator does that just fine.
- Davidzheng 2y agoMath is not just arithmetic
- ta988 2y agoIt is not just about math it is about tokenization and chain of thought.
- muzani 2y agome: "is 400000 the same as 4 * 100 * 100" gpt-4o: "Yes, 400,000 is the same as 4 X 100 × 1000." gpt-4: Yes, 400,000 is the same as 4 X 100 × 100. Here's the breakdown: 4 × 100 = 400 400 × 100 = 40,000 Oops, it looks like there was an error in my previous calculation. So, 4 × 100 X 100 actually equals 40,000, not 400,000. To make it 400,000, you would need to use 4 × 10 x 10,000 or adjust the numbers similarly. claude-3.5: No, 400000 is not the same as 4 * 100 * 100. Let's break it down: 4 * 100 * 100 = 4 * 10000 = 40000 400000 is ten times larger than 40000. Everyone keeps saying gpt-4o beats benchmarks and stuff, but this is consistently my experience with it. The benchmarks fall far off my every day experience.
- Davidzheng 2y agoThe thing is, gpt4(o) is the only model I've talked to that makes me feel like it understands calculus. Not like it's pretending and succeeding
- muzani 2y ago
- andrewchambers 2y agoI wonder if openai they have a response ready or if they just are tackling other business problems like ios integration now and the seemingly postponed low latency chat launch. Either way I am looking forward to claude 3.5 opus. It did seem slightly odd to me that openai made their supposedly best model free.
- replwoacause 2y agoI pay for both OpenAI and Anthropic pro plans, and can say OpenAI is lagging behind at this point. Hopefully their next model steps it up.
- lowyek 2y agoMust be fun working on cutting edge competitive stuff for all these 3 major teams. It's exciting to live in this times and see this all unfold in our eyes.
- jimsimmons 2y agoWho 3
- netsec_burn 2y agoMeta, OpenAI, and Anthropic come to mind. In TFA they name OpenAI, Google DeepMind and Anthropic.
- lowyek 2y agoTFA = The Freaking Article; wow! what a full form.
- snewman 2y agoPresumably Anthropic / Claude, OpenAI / GPT, Google DeepMind / Gemini.
- lowyek 2y agothank you. yes I meant that.
- muzani 2y agoLikely 4, based on the other comments. Maybe 5 if you want to add Midjourney.
- daniel_iversen 2y agoOn that note, and apologies if it sounds spammy, but genuinely if there's any AI engineers reading, check out asana.com/jobs - we partner with the two leading AI labs (they're all customers of ours too) and I know the team get to experiment with some of the early release stuff to build our leading AI work management platform (Dustin, our very technical founder and CEO writes about what the deal with our latest AI advancements are here https://asana.com/inside-asana/the-key-to-unlocking-the-potential-of-ai-teammates https://asana.com/inside-asana/the-key-to-unlocking-the-pote...) - I feel like it's one of the best places to work in so many ways, but the cutting edge AI stuff is REALLY fun!
- npn 2y ago[flagged]
- alach11 2y agoClaude 3.5 Sonnet took a solid lead on our internal benchmarks over gpt-4-turbo for extraction tasks against large documents. It may not be great for every workflow, but it certainly hits a sweet spot for intelligence x cost on most of my workflows.
- sdwr 2y agoAre there any tools out there that expose GPT or Claude to a codebase, and let it write PRs (semi) autonomously?
- campers 2y agoAnother one is https://www.github.com/trafficguard/nous https://www.github.com/trafficguard/nous which provides a software dev agent that can find/clone repos, create a branch, search in a repo, then delegates to Aider for the code editing in the edit/compile/lint/test loop, and raise a merge request.
- CGamesPlay 2y agoI know of Sweep <https://docs.sweep.dev https://docs.sweep.dev>, but honestly the SOTA on SWE-bench, which is basically exactly what you're asking for, is only about a 25% success rate, so expect very mediocre results.
- tananaev 2y agoIt's probably not going to be very good at handling it out of the box. It would require quite a bit of fine tuning.
- nyellin 2y agoaider is pretty good - https://github.com/paul-gauthier/aider https://github.com/paul-gauthier/aider
- sdwr 2y agoThanks, tried it out - it's pretty cheap and pretty good. Being able to commit directly into multiple files simultaneously is great.
- kleneway1 2y agoI’m working on adding Sonnet 3.5 to JACoB this week. So far it’s been very impressive. https://github.com/jacob-ai-bot/jacob https://github.com/jacob-ai-bot/jacob
- 2y ago
- m0zzie 2y agoCan anyone comment on its coding ability? Considering cancelling my subscription with OpenAI as I was previously using GPT-4 quite heavily as a multiplier for myself, guiding it and editing outputs as required, but GPT-4o feels significantly worse for this use case. It is certainly better in many other areas, but its coding ability is not great. I tried to revert back to standard GPT-4 but it is now so slow to respond (higher load?) that it breaks my mental flow, so I'm exploring other options.
- cyral 2y agoI've been playing around with it this week and its coding ability is insane (for a LLM). I've given it some pretty sloppy descriptions about things I want to do and it's managed to figure out exactly how to do it on the first or second try, I'm talking things like building animations in React that cannot be described with text very well. Big pain point is copy and pasting things back and forth to have it edit them. If it was integrated and could see my local files, that would be killer. I know there are various companies working on that, but the jetbrains AI integration for example is garbage compared to the results I get by manually asking claude. I wasn't worried about how this would affect our industry a few months ago, but this has me reconsidering. It's like a junior engineer that can do most tasks in seconds for a couple of cents.
- hdhshdhshdjd 2y agoWhat worries me is you need that time in the dirt to get a feel for coding as a craft. And at least for me that aspect of knowing the craft helps get my thinking in tune with problem solving in a very productive way. Coding can be similar to playing an instrument, if you have mastery, it can help you be more expressive with the ideas you already have and lead you to new ones. Whereas if we take away the craft of coding I think you end up with the type of code academic labs produce: something that purely starts on a “drawing board”, is given to the grad student/intern/LLM to make work, and while it will prove the concept it won’t scale into long term, as the intern doesn’t know when to spend an extra 30 minutes in a function so that it may be more flexible down the road.
- andrewstuart 2y agoAI programming would be really useful if it moved towards me being able to make a fixed set of statements about the software, those statements are preserved permanently in the source code somehow, and the AI ensures that those statements remain true. Its frustrating to work with the AI to implement something only to realise within a few interactions that it has forgotten or lost track of something I deemed to be a key requirement. Surely the future of software has to start to include declarative statement prompts as part of the source code.
- simonw 2y agoHave you tried building that system with prompting? You could set a convention of having comments something like this: # Requirement: function always returns an array of strings And then have a system prompt which tells the model to always obey comments like that, and to add comments like that to record important requirements provided by the user.
- gavindean90 2y agoThis is a good solution. I’ve taken the idea that maintaining the vision in my job in the LLM relationship. Reiterating key details street it’s forgotten and burning tokens tweaking things over and over towards the vision is the cost of doing business.
- ilaksh 2y agoThe first thing I built with OpenAI a few years ago was a system that had a section for a program spec on one part of the screen and a live web page area on the other part. You could edit the spec and it would then regenerate the web page. It would probably work much better now. Eventually I might add something like that as a plugin in my current framework.
- tkgally 2y agoThis article was written before Anthropic added the Projects feature for Pro and Team users [1]. I spent some time yesterday experimenting with Projects, and, like Artifacts, it looks really useful. I like the idea of being able to have multiple projects going simultaneously, each with its own reference materials. I don’t need to use it in a team, but I can see how that could be useful, too. The one problem I see is that the total context window for each project might start to seem too small pretty quickly. I assume, though, that Anthropic’s context windows will be getting larger as time goes on. I wonder what other features Anthropic has in the works for Claude. My personal wish is for a voice interface, something like what OpenAI announced in May but has now put off until later this year. [1] https://www.anthropic.com/news/projects https://www.anthropic.com/news/projects
- muzani 2y agoI'm far more excited for Projects than the other stuff. OpenAI's Memories is their lock-in for me; I'd have switched to Claude if Anthropic had something better. Projects sounds like an even better approach.
- 8n4vidtmkvmk 2y agoWhat are you using memories for? Openai told me it memorized some stuff, but it's never used it in a useful way nor do I see how it would.
- muzani 2y agoGame dev. It keeps tabs on the type of game it is, background, genre, classes, theme. Notably the difference between one decision and another - it takes note when I say "Instead of Y, do X". Like if I show a preference of Chinese names, it'll give me more of that.
- viraptor 2y agoFrom my reading there's nothing here that's specific to Anthropic, right? Any app like Chatbox, Msty or other could implement their own version of Projects based on any of the available models. (Maybe even Notion...) The API for creating an assistant with a pre-baked context is also available. Or did I miss something unique there?
- nyellin 2y agoInteresting. We're finding that it's more verbose than GPT-4o (and often ignores instructions to give output in a specific format), but does better when faced with ambiguity and potentially confusing situations. Here is an example using http://github.com/robusta-dev/holmesgpt/ http://github.com/robusta-dev/holmesgpt/ where ask it to troubleshoot a payment processing workload in a Kubernetes cluster vs the GPT-4o output underneath: holmes ask --model="bedrock/anthropic.claude-3-5-sonnet-20240620-v1:0" "what is wrong with the payment processor?" AI: To investigate what's wrong with the payment processor, we'll need to gather more information about it. First, let's try to find the relevant Kubernetes resources related to the payment processor. Running `kubectl get -A --show-labels pod | grep payment` Thank you for your question about the payment processor. initial search, I've found two pods related to payment processing. Let's investigate the one that seems to be having issues: 1. In the default namespace, there's a pod named "payment-processing-worker-5bb5bd75cd-x4g99" which is in a CrashLoopBackOff state. This indicates that the pod is repeatedly crashing and restarting. 2. In the sock-shop namespace, there's a pod named "payment-ff86cd6f8-rs47b" which appears to be running normally. Let's focus on the problematic pod in the default namespace. To get more information about what's wrong, we'll need to look at its logs and description. Running `kubectl describe pod payment-processing-worker-5bb5bd75cd-x4g99 -n default` AI: Based on the information gathered, it appears that there is an issue with the payment processor in the Kubernetes cluster. Let me explain what's wrong: 1 The payment processing worker pod is in a CrashLoopBackOff state: • Pod name: payment-processing-worker-5bb5bd75cd-x4g99 • Namespace: default • Status: CrashLoopBackOff 2 The pod is continuously restarting and failing: • Restart Count: 242 • The container is in a "Waiting" state with the reason "CrashLoopBackOff" 3 The root cause of the issue appears to be related to an environment variable: • The container is running a bash script that checks for the presence of the DEPLOY_ENV environment variable. • If DEPLOY_ENV is not set, the container exits immediately after printing an error message. • The command being run is: if [[ -z "${DEPLOY_ENV}" ]]; then echo Environment variable DEPLOY_ENV is undefined ; else while true; do echo hello; sleep 10;done; fi 4 The environment section in the pod description shows: Environment: <none> This indicates that no environment variables are set for the container. 5 The container is exiting with a status of "Completed" and an exit code of 0, which suggests that it's running the script, finding that DEPLOY_E To fix this issue: 1 The DEPLOY_ENV environment variable needs to be set for the payment-processing-worker deployment. 2 You should update the deployment configuration to include this environment variable. 3 After setting the environment variable, the pod should stop crashing and enter a running state where it prints "hello" every 10 seconds. To resolve this, you'll need to update the deployment configuration for the payment-processing-worker. You can do this by editing the deployment and adding the DEPLOY_ENV environment variable with an appropriate value. Would you like me to show you how to update the deployment with the required environment variable? Note that it explicitly picks up on the ambiguity in the question and notes that there are 2 payment related workloads and one is healthy. GPT-4o: $ holmes ask "what is wrong with the payment processor?" Running `kubectl get -A --show-labels pod | grep payment-processor` Running `kubectl get -A --show-labels deployment | grep payment-processor` Running `kubectl get -A --show-labels service | grep payment-processor` Running `kubectl get -A --show-labels pod | grep payment` Running `kubectl describe pod payment-processing-worker-5bb5bd75cd-x4g99 -n default` Running `kubectl logs payment-processing-worker-5bb5bd75cd-x4g99 -n default --previous` AI: payment-processing-worker-5bb5bd75cd-x4g99 in namespace default crashes due to missing DEPLOY_ENV environment variable. Fix: Define DEPLOY_ENV in the deployment spec. Example patch: env: - name: DEPLOY_ENV value: "production"
- SubiculumCode 2y agoI'm fairly impressed with Sonnet's one shot scripting performance for my use cases. However, I was using it to help me diagnoses a gnome key ring and ssh issue I was having, and it suggested that I 'rm -rf' my keyring files to test it's solution out. A little drastic..maybe mv my password file first? Anyway, it sometimes seems even more cocky than last gen, and less careful by default
- aappleby 2y agoI don't need an AI to write code for me, but it is _astoundingly_ helpful to have it summarize various design options and new technology stacks without me having to scavenge Google for the obscure corner-cases I care about. I have an idea for a project that involves streaming 3 giabits of data per second from a USB 3.0 device out over a 10 gig Ethernet connection, and it was able to compare/contrast various levels of support for high-bandwidth USB 3 and Ethernet in multiple frameworks and languages. And the whole conversation, with code examples, cost me 3 _cents_ of Anthropic credits. My new fear is when people start asking AIs "Hey AI, here is my codebase, my org chart, and commit histories for all my employees - how can I reduce the number of humans I need to employ to get this project done?"
- gjsman-1000 2y agoNobody cared when it was the artists, the songwriters, the music studios, the film companies. But impacting jobs of the programmers - now, heaven forbid.
- delgaudm 2y ago[flagged]
- squigz 2y agoI'm pretty sure people cared about the artists too.
- gjsman-1000 2y agoDid the people building dall-e care?
- squigz 2y agoAre those people here? Do you reckon GP is one of them?
- thethirdone 2y ago
- GaggiX 2y agoThe incredible ability of Claude 3.5 Sonnet to create coherent SVG makes me wonder if the LLM was not just pretrained on text. Vision capabilities are usually added later using a vision encoder that does not affect the LLM's knowledge of the visual world, but in this case the LLM clearly has quite a strong understanding of the visual world.
- levocardia 2y agoHave you read the "sparks of AGI" paper about GPT4? It suggested that even just text can give an LLM a rich world model, based on the tikz drawings of a unicorn that got progressively better as GPT4 precursors were trained on increasingly more data (and, interestingly, the drawings got worse when it was RLHF'd for safety).
- GaggiX 2y agoYes of course, as always, it's very possible that just scaling solved the problem, but the fact that the model is so good makes me wonder if they actually did something different and pre-trained the model on image tokens as well.
- thethirdone 2y ago> You can say ‘the recent jumps are relatively small’ or you can notice that (1) there is an upper bound at 100 rapidly approaching for this set of benchmarks, and (2) the releases are coming quickly one after another and the slope of the line is accelerating despite being close to the maximum. The graph does not look like it is accelerating. I actually struggle to imagine what about it convinced the author the progress is accelerating. I would be very interested in a more detailed graph that shows individual benchmarks because it should be possible to see some benchmarks effectively be beaten and get a good idea of where all of the other benchmarks are on that trend. The 100 % upper bound is likely very hard to approach, but I don't know if the limit is like 99%, 95% or 90% for most benchmarks.
- simonw 2y agoI heard a theory today that hitting 100% on the MMLU benchmark may be impossible due to errors in that benchmark itself - if there are errors in the benchmark no model should ever be able to score 100% on it. The same problem could well be present in other benchmarks as well.
- willsmith72 2y agoi took it to mean progress is increasing, not rate of progress is increasing. a classic case of "acceleration misuse" but nothing more
- aoeusnth1 2y agoI think this is what they meant: https://imgur.com/a/GWqfp9U https://imgur.com/a/GWqfp9U If you take the upper bounds at any given point in time, the rate of increase of the best models over time is accelerating.
- cortesi 2y agoClaude 3.5 Sonnet's coding abilities are incredibly impressive. I think it lets an expert programmer move more than twice as fast. There are limits - to produce high quality code, not copy-and-paste pablum, you have to be able to give detailed step-by-step directions and critically evaluate the results. This means you can't produce code better than you would have written by yourself, you can only do it much faster. As an experiment, I produced a set of bindings to Anthropic's API pair-programming with Claude. The project is of pretty good quality, and includes advanced features like streaming and type-safe definitions of tools. More than 95% of the code and docs was written by Claude, under close direction from me. The project is here: https://github.com/cortesi/misanthropy https://github.com/cortesi/misanthropy And I've shared part of the conversation that produced it in a video here: https://twitter.com/cortesi/status/1806135130446307340 https://twitter.com/cortesi/status/1806135130446307340
- zeroonetwothree 2y agoI find LLM coding much less useful when it’s interacting with a large existing codebase. It’s certainly good at one-off type code and greenfield projects (especially if similar to other open source stuff). And it’s also good at getting started if you aren’t an expert yourself.
- cortesi 2y agoWe haven't found this to be an impediment. Keep things modular, and share the type definitions of anything you import with the model. As the benefits here become more and more clear tooling will improve and people will adapt their development practices to get the most out of the models.
- miohtama 2y agoI have been developing Python 20 years now. Claude 3.5 is the first AI that is “smart” enough to help me. I usually do not need help with easy task, but complex ones. Claude is not a perfect, but it definitely gives a productivity boost for even the most seasoned developers, which would have been some obscure mailing list and source code reading in the past.
- 2y ago
- mrcwinn 2y agoI’ve thoroughly enjoyed the product overall much more than ChatGPT. I do wish it had voice input that rivaled what OpenAI previewed. Excited for 3.5 Opus. For now I’ve canceled OpenAI subscription and removed the app in favor of Claude.
- silisili 2y agoWell, I went to try it, but it requires a phone number for some bizarre reason. Fine, gave it my primary number, a google voice number I've had for a decade, and it won't accept it. That's the end of my Claude journey, forever. If you want me to try your service, try using some flow with less friction than sandpaper, folks.
- seaal 2y agoGoogle Voice blacklist is pretty common.
- silisili 2y agoI have about a dozen credit cards, 4 bank accounts, mortgage, car payment, utility accounts, github 2fa, aws 2fa, Fidelity retirement, etc. Not one has an issue with my number. I did have some service refuse it, I want to say Twitter? But I'd definitely not consider it common. This is probably only the second or third time I've seen it, tbh.
- simonw 2y agoI'm going to guess the reason Anthropic do verified phone numbers is that, unlike email addresses, most people don't have an easy way to create multiple phone numbers. Since Anthropic accounts come with free rate-limited access to their models they're trying to avoid freeloaders who sign up for hundreds of accounts in order to work around those caps. Google Voice numbers are blocked because people can create multiple of those, which would allow them to circumvent those limits.
- tonyoconnell 2y agoI have been using Claude Sonnet with Artifacts along with Vercel V0 to build Sveltekit pages and components really well. I create a UI in V0 and then simply copy the JSX into Claude and tell it to convert to Sveltekit. It creates the +page.svelte +page.server.ts and all the components almost perfectly.
- whiddershins 2y agoI wish it could preview svelte builds. I wish it would let me include svelte files in “projects.”
- ttul 2y agoClaude helped me code up a quick webhook handler for a silly side project. Being a side project, I asked it to put the code into a Docker container, which is did flawlessly. It also gave me a Makefile with sensible targets for building, rebuilding, starting, stopping, etc. Finally, I asked it use an .env file for the secrets I needed to store. Everything it did was one-shot on target. The only thing that would make this better would be to have it in the terminal with me.
- lukasb 2y agoI keep saying that if nothing else, we are in a golden age of actually finishing side projects. All those “it’s easy, I’ll just glue this API to this API” projects really are easy now.
- r2_pilot 2y agoThis, exactly. Every side project I've thrown at Claude 3.5 Sonnet has been completed that same night. It's so different from how I used to treat my backlog projects which could take a week or so of research-code-iterate and now they're just an evening (or less; since Sonnet's upgrade on average it's taken me about 20-40 minutes to get the night's planned work done) and I get to sleep earlier. Win-win-win
- brcmthrowaway 2y ago
- renewiltord 2y agoClaude Sonnet is freaking amazing. I used to have a safety test[0] that Claude failed. But it was a bogus safety test and fortunately someone here told me so and I immediately subscribed to it. It's amazing. The other day I ported a whole node.js script to Python with it. It was not flawless but it was pretty damned good. Such a mechanical process and I just had to review. Loved it. 0: https://news.ycombinator.com/item?id=39607069 https://news.ycombinator.com/item?id=39607069
- namcxr 2y agoI do not understand this. Will you still love it when it takes your job? (Assuming it is really that good, which is something that I can never replicate.)
- dagaci 2y agoApparently my account was banned on Anthropic Sonnet after a "Automatic review". I'm 100% sure i did not make any "unsafe" queries, I've litterally only briefly tested and that was weeks ago. +1 OpenAI Subscription -1 Anthropic Sonnet->sudden-death-automatic-review-system
- bl4ckneon 2y agoI had that too, no explanation, no response via email support, nothing. I can't give money to them knowing my account can just get banned at anytime while I might have an active subscription going on.
- andhuman 2y agoMine too! And I didn’t even get to use it once. I’ve filled in the form, let’s see if they lift the ban.
- namanyayg 2y agoYou may be able to create a new account from a different email?
- deleted 2y ago[deleted]
- dagaci 2y agoAccounts are tied to your phone number.
- whywhywhywhy 2y agoWhich are locked to the first account verified with them.
- newscracker 2y agoSo when someone lets (or has) their phone service get disconnected and the company recycles that phone number shortly, the person who then gets this phone number is out of luck if they want to use Claude?
- JCM9 2y agoThis is highlighting what has happened with all forms of ML. Give a baseline set of folks the same dataset and they will end up with a model that performs about the same. Companies are one-upping each other but it’s very back and forth and just a case of release date. These models will become a complete commodity. The thing that could be proprietary is the data used to train them, which could lead to a sustained better model performance. The barrier to entry here is super high given training costs, but the ML skills are still a commodity.
- bboygravity 2y agoThe dataset could become proprietary you say? In other words: information with copyright on it that was (illegally) used/stolen from billions of people and companies will get copyright on it that will be resold as a set? I don't know...
- realharo 2y agoI think the comment was referring to things like internal company data, or licensed data that is not publicly available, etc. Those things could be a competitive advantage.
- brap 2y agoThis is why I strongly believe Google has the clear advantage. Infinite data, infinite resources , not to mention dozens of verticals (search, Android and Chrome are probably the killer ones)
- yunwal 2y agoGoogle obviously has the advantage here, but it also seems like they’re willing to squander it. The Gemini rollout has basically been clippy 2.0 so far. The Gemini interface in gcloud seems to know nothing about how services work, the chat summaries consistently fail, the search summaries are comically bad. I’m usually not one of these people who wants to “block all the AI stuff” but with google products I do.
- liquidise 2y agoAs someone building an AI company right now, my quick Pro/Con for 4o vs Claude 3.5: Claude: subjectively sounds more human to me, and really nails data questions that 4o is lackluster at 4o: far better assistant logic reasoning. I can trivially break Claude's assistant (system prompt) instructions within the user prompt, where 4o succeeds in all of these tests. Pricing and output speed, for our purposes, are functionally identical. Exciting to have a competitor in the space already who stands to keep openai honest.
- xwolfi 2y agoAnd hum, what incredible problem are you solving at your AI company? Must be the forefront of human innovation ! I hope it's porn.
- dmazin 2y agoAha, so I’m not the only one. For both Claude 3 Opus and 3.5 Sonnet, anecdotally its language is far more natural. So much so that I prefer it over 4o.
- deleted 2y ago[deleted]
- willsmith72 2y agoam i right in that it has no online capabilities? that's a pretty big issue for me
- xixixao 2y agoI tried Sonnet vs GPT 4 just now with: > Given a body with momentum B and forques F, what is the differential of applying the forques to the momentum in PGA? Claude gave a wrong answer, ChatGPT gave a correct one. I’m sticking with ChatGPT.
- Smaug123 2y agoI have a mathematics (though not physics) degree and I didn't understand your question at all; "forques" appears to be either a place in France, Old French, or Catalan. I assume ChatGPT was correct in re-spelling "forques" as "torques", but have you tried asking Claude using words that do appear on the Internet?
- xixixao 2y agoUnlike you, both LLMs were familiar with geometric algebra and used the relevant terminology. Testing on something widely known isn’t likely to stretch these systems.
- Smaug123 2y agoI'd expect them to do better when the input uses words that appear more in the training data. This very thread is the fifth hit on Google for `"forques" geometric algebra`; the third and fourth hit are the same paper as each other; the second hit is https://bivector.net/PGAdyn.pdf https://bivector.net/PGAdyn.pdf which appears to have invented the term; and the first hit doesn't define it. I (logic, computability, set and type theory) am in no position to know whether it's a standard term in geometric algebra, but I do strongly expect LLMs to do much worse on queries that don't appear much in their training set (for which I take Google search results as a proxy); even if they have the knowledge to answer, I expect them to answer better when the question uses common words. I do know that when I asked your question to ChatGPT, it silently re-spelt "forques" as "torques".
- alastairr 2y agomy 2p worth - my work involves a lot of summarisation, recommendation from a user preference statement. I've been able to do this with 4o / opus, but the consistency wasn't there, which required complex prompting chains to stabilise. What I'm seeing with Sonnet 3.5 is a night-and-day step up in consistency. The responses don't seem to be that different in capability of opus / 4o when they respond well, it just does it with rock-solid consistency. That sounds a bit dull, but it's a huge step forward for me and I suspect for others.
- deleted 2y ago[deleted]
- boyka 2y agoThese models are clearly great with language, be it natural language or code. However, I wonder where the expectation comes from that a static stochastic parrot should be able to compute arbitrary first order logic (in a series of one-shot next word predictions). Could any expert elaborate on how this would be solved by a transformer model?
- thmixc 2y agoClaude 3.5 Sonnet can solve the farmer and sheep problem with two small changes to the prompt: 1. change the word "person" to "human". 2. change the word "trips" to "trip or trips". (Claude is probably assuming that the answer has to be in multiple trips because of the word "trips")
- erdemo 2y agoAs developer Claude code generator 2x better than gpt4o, of course it subjunctive but Claude much consistent for me.
- redkrc 2y agoاريدك ان تجعل هذه الصفحة اكبر واكثر صفحة احترافية في التاريخ اكثر من موقع ابل واقوي من اقوي البراندات العالمية اريدك ان تضيف مزايا احترافية جداجدا ليصبح الموقع والتطبيق رقم 1 في مجال تسجيل الاوزان والجيم واكتب انت جميع الاكواد لانه لا خبرة لي في البرمجة اطلاقا ولا استطيع ممكن ان ياخذ ذلك مني سنينا اتمني ان تشارك في هذا العمل الانساني الخيري
- nojvek 2y agoI cancelled my OpenAI membership and using more and more of Claude. Sonnet is pretty fast and cheaper than 4-o. I'm legit elated that a smaller player is able to compete with large behemoths like OpenAI and Google. (I know they have Amazon backing them, but their team is much smaller. OpenAI is ~1000 employees now). I'm building on top of their api. It's neat. I wish them the best.