46 ms·
Claude 2
- hbbio 3y agoI would never trust an assistant that keeps repeating it's "helpful, harmless, and honest" every couple prompts!
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- naillo 3y agoExcited for it at a distance. Wish I could try it though (not in the US or UK).
- camillomiller 3y agoworks with any VPN
- spacebanana7 3y ago> Claude 2 powers our chat experience, and is generally available in the US and UK. We are working to make Claude more globally available in the coming months. I wonder why LLMs like GPT-4, Bard and Claude are so geo restricted at first? I understand some places have regulatory challenges but can’t see SG, UAE, or Chile being too difficult.
- gkk 3y agoI'd guess Anthropic considers these 2nd tier markets, so it's not a question whether it's too difficult but whether it's a priority at the moment.
- disgruntledphd2 3y agoI would say that they want English language only, and not EU. The hilarious part of that is that the UK has basically all the regulations that they are probably worried about.
- spiderfarmer 3y agoEven more hilarious is that everyone in their target audience speaks English.
- londons_explore 3y agoThey want places with tech startups who will pay for their API. Thats where there is lots of money to be made. And if they are GPU constrained, then launching in the countries with the highest proportion of future paying customers makes sense.
- dragonwriter 3y ago> I wonder why LLMs like GPT-4, Bard and Claude are so geo restricted at first? Managing scale while maximizing profit potential? Also, US/UK probably lets them put their strongest linguistic foot forward initially, and there may be additional training done before rolling out to regions with other dominant languages. > I understand some places have regulatory challenges That’s probably not the main issue.
- deleted 3y ago[deleted]
- redox99 3y agoI don't think GPT4 was geo restricted?
- agucova 3y ago> I wonder why LLMs like GPT-4, Bard and Claude are so geo restricted at first? I understand some places have regulatory challenges but can’t see SG, UAE, or Chile being too difficult. I'm amused by the inclusion of Chile in this list. I'm a Chilean and I do have access, but through the Anthropic Console, as I already had API Access.
- jiggawatts 3y agoIt made me laugh when Google announced how strong Bard 2 is at over a hundred human languages and then restricted the deployed chat app to like… three. That’s not even region locking, it’s capability locking while simultaneously advertising that very thing!
- taf2 3y agoI’m very excited for Claude - been using it along side gpt 4 and pleased with its performance. The introduction of functions with OpenAI api complicates things and was hoping Claude would include this in a future api update
- phillipcarter 3y agoExcited to try it. We used Claude 1.x in experimentation, but shipped with OpenAI primarily because of time and SOC 2 compliance. Anthropic has come along since then, so we'll probably experiment with Claude more with intent to take into production if it's still holding up.
- jasondclinton 3y agoWe have SOC 2 Type 1 and HIPAA now. Working on more. Excited that you liked it!
- binarymax 3y agoHi! Do represent anthropic? Your bio says you’re at google.
- jasondclinton 3y agoWhoops, fixed.
- binarymax 3y agoCool. As CISO, can you please speak to the data retention policies that I noted here? https://news.ycombinator.com/item?id=36681239 https://news.ycombinator.com/item?id=36681239 . As you can imagine, sending sensitive information to a 3rd party is impossible without explicit agreements. As you're SOC2 and HIPAA are there devices in place for us to delete data, or specify data retention as customers?
- jasondclinton 3y agoReplied there, thank you for pointing to that.
- AviationAtom 3y agoNot sure what kind of equity you negotiated when signing on with the company, but it's going to pay off handsomely. Wish I had more exposure to the company, to better join the ride, but I'll take what I have now. Keep contributing to the awesome efforts going on there.
- deleted 3y ago[deleted]
- netcraft 3y agoI thought for a moment that it could reach out to the internet, and it certainly makes it think you can, but its just lying about it. I was able to get it to summarize the "How to Do Great Work" article with its url, but trying to get it to summarize the comments of the current laser chess HN article gave me something about cryptocurrency.
- phgn 3y agoThe logo animation is really nice! I've collapsed & expanded it at least 10 times now, maybe I should get to reading the article...
- linsomniac 3y agoI've just been playing with Claude 1.3 this weekend to summarize large texts. It can take 100K tokens of input, enough for a whole Lex Fridman interview! :-) I've been getting pretty good results with it, so I'm excited to see how v2 works.
- AviationAtom 3y agoNow that it's entered open beta it's going to iterate rapidly. I had been using it fairly extensively, alongside other LLMs, through Slack and was always most impressed by it's output over the others. (I do hold investment in Anthropic, but do not base my statements on that)
- SomaticPirate 3y agoHow are you invested in Anthropic?
- AviationAtom 3y agoI posted about it in another comment but will restate it here too. The ARK Venture Fund has exposure to it (roughly 7% of the fund).
- xfalcox 3y agoCan you share the prompts you used ? I'm really happy with Claude-100k for summarization, but I wonder if a better prompt would make it even better.
- linsomniac 3y agoSure, here it is: Human: Here is the transcript of a podcast: <transcript> [PASTE TRANSCRIPT HERE] </transcript> You are an expert on writing factual summaries. Write a summary of the podcast in about 10 sentences. Assistant: I'd be happy to, here is the summary:
- ilaksh 3y agoI applied and got access to the Claude 1 API a long time ago and then I guess I didn't click the link they gave me in time or something because when I went to try to get in it was expired. If I remember correctly. I think I emailed them about it and was ignored. I've been using the OpenAI API and I'm on the third version of my code generation application which is now a ChatGPT Plugin. It sounds like Claude 2's reasoning is still lagging behind GPT-4 anyway.
- unsupp0rted 3y agoI have the same problem with resemble.ai - I've submitted their "request a demo" form multiple times to try to get access to their multi-language API. Can't get a reply. I've tried emailing their support and sales teams and they ignore me.
- ilaksh 3y agoWell maybe someone saw me complain because I got a new invite. Thank you, someone!!
- moffkalast 3y agoGlad to see Anthropic fall into the 'we ignore customer service until we get publicly shamed and it makes us look bad' company category. /s
- DoryMinh 3y agoFantastic, now we have duopoly
- alpark3 3y ago> monopoly
- gberger 3y agoThere is no moat.
- usaar333 3y agoSeems inferior to GPT-4 on every test I've given it - but as a competitor to GPT 3.5 is strong.
- abdullin 3y agoOn our benchmarks, Claude v1 beats GPT-3.5 (v0613) while v2 looses to it.
- ianhawes 3y agoIMO the rankings of publicly available LLMs are: 1. GPT-4 2. Claude 2 3. Bard 4. Llama/Alpaca 5-98. [Unclaimed] 99. SmarterChild AIM bot 100. Cohere All joking aside, I do agree with the sentiment that no one generally has any type of defensible moat at the moment. OpenAI has found a great balancing act between first mover advantage, marketing, customer adoption, and enterprise sales. They are executing at a high level. Anthropic (Claude) has a wonderful product but is lacking in consumer adoption and sales, though I think they're working on fixing that.
- ilrwbwrkhv 3y agoAll the AI companies are sort of doing a VC rush, but instead of IPO it's AGI. Would be fun to see what we get in the future. Since a serious training run costs upwards of $50 million currently.
- AviationAtom 3y ago
- londons_explore 3y agoHow does it score on the LLM leaderboards[1]? They seem like the best way to evaluate models for general purpose use right now. [1]: https://chat.lmsys.org/?arena https://chat.lmsys.org/?arena
- abdullin 3y agoOn our benchmarks, Claude v2 scores worse than v1 in categories “code”, “docs”, “integrate” and “marketing”. It also is more chatty than v1 (or GPT-3/4), even when asked to just pick one option out of three. These benchmarks are product oriented - they contain tests and evals from our LLM-driven products. So they aren’t exhaustive or representative. We just want to know when local LLMs are good enough to start migrating some pipelines away from OpenAI.
- extasia 3y agoAnybody got a model card?
- cubefox 3y agoFirst sentence has the link: https://www-files.anthropic.com/production/images/Model-Card-Claude-2.pdf https://www-files.anthropic.com/production/images/Model-Card...
- gjstein 3y agoExcited for this, but I think with all this conversation about the role an AI assistant should play in work and development, this line feels incomplete to me: > Think of Claude as a friendly, enthusiastic colleague or personal assistant who can be instructed in natural language to help you with many tasks. It omits that the colleague may have outdated knowledge or not understand whatever problem you give it. The colleague's "enthusiasm" should be tempered with oversight so that the outputs they produce are not directly used without scrutiny. It seems that most people using these tools increasingly understand this, but to leave it off the website seems ... sloppy at this point. Edit: upon logging in, I'm greeted by a warning "It may occasionally generate incorrect or misleading information, or produce offensive or biased content."
- whimsicalism 3y agoIt seems as if there are many possible things they could omit, given that this is a blog post of finite word count.
- sva_ 3y ago> Unfortunately, Claude.ai is only available in the US and UK. We're working hard to expand to other regions soon.
- TheBlapse 3y agoWorks with VPN
- AviationAtom 3y agoIt's been available through Slack for some time now
- throwaway1777 3y agoThe slack version doesn’t work for me anymore.
- AviationAtom 3y agoI just tried it again and it's still working for me. Were you accessing it in any special way? It should have just been a matter off adding the app to your Slack instance.
- throwaway1777 3y agoHuh. Maybe I’ll try removing it and onboard again.
- awestroke 3y agoRegion locking digital services is such a stone age approach
- Aerbil313 3y agoIf it works for the masses it works for the masses.
- 3y ago
- binarymax 3y agoI'd like to try Claude, but the data retention policies in the Anthropic terms are not clear. Section 6e[0] claims they won't use customer data to train models, but I'd like to know if customer data is kept for any duration (like it is with OpenAI for 30 days). There is a note about data deletion on termination in section 14, so I assume that ALL data is retained for an undisclosed period of time. [0] https://console.anthropic.com/legal/terms https://console.anthropic.com/legal/terms
- l1n 3y agohttps://support.anthropic.com/en/articles/7996866-how-long-do-you-store-personal-data https://support.anthropic.com/en/articles/7996866-how-long-d...
- rat9988 3y agoI see why it could be a problem for using it, but you can still try it and then delete your data?
- jasondclinton 3y agoThe canonical answer is in this on the 3rd bullet point: https://support.anthropic.com/en/articles/7996866-how-long-do-you-store-personal-data https://support.anthropic.com/en/articles/7996866-how-long-d... I’m excited that you’re passionate about privacy. We’ve put a lot of thought into our policies.
- binarymax 3y agoThanks! This is very helpful. Congrats on the launch.
- agnokapathetic 3y agoLink is broken
- tmikaeld 3y ago"we automatically delete prompts and outputs on the backend within 30 days of receipt or generation unless you request otherwise"
- lhl 3y agoSince I've been on a AI code-helper kick recently. According to the post, Claude 2 now 71.2%, a significant upgrade from 1.3 (56.0%). (Found in model card: pass@1) For comparison: * GPT-4 claims 85.4 on HumanEval, in a recent paper https://arxiv.org/pdf/2303.11366.pdf https://arxiv.org/pdf/2303.11366.pdf GPT-4 was tested at 80.1 pass@1 and 91 pass@1 using their Reflexion technique. They also include MBPP and Leetcode Hard benchmark comparisons * WizardCoder, a StarCoder fine-tune is one of the top open models, scoring a 57.3 pass@1, model card here: https://huggingface.co/WizardLM/WizardCoder-15B-V1.0 https://huggingface.co/WizardLM/WizardCoder-15B-V1.0 * The best open model I know of atm is replit-code-instruct-glaive, a replit-code-3b fine tune, which scores a 63.5% pass@1. An independent developer abacaj has reproduced that announcement as part of code-eval, a repo for getting human-eval results: https://github.com/abacaj/code-eval https://github.com/abacaj/code-eval Those interested in this area may also want to take a look at this repo https://github.com/my-other-github-account/llm-humaneval-benchmarks https://github.com/my-other-github-account/llm-humaneval-ben... that also ranks with Eval+, the CanAiCode Leaderboard https://huggingface.co/spaces/mike-ravkine/can-ai-code-results https://huggingface.co/spaces/mike-ravkine/can-ai-code-resul... and airate https://github.com/catid/supercharger/tree/main/airate https://github.com/catid/supercharger/tree/main/airate Also, as with all LLM evals, to be taken with a grain of salt... Liu, Jiawei, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. “Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation.” arXiv, June 12, 2023. https://doi.org/10.48550/arXiv.2305.01210 https://doi.org/10.48550/arXiv.2305.01210.
- famouswaffles 3y agoGPT-4's zero shot Human Eval score was 67%
- lhl 3y agoWhile that's what the Technical Report (https://arxiv.org/pdf/2303.08774v3.pdf https://arxiv.org/pdf/2303.08774v3.pdf) says, but GPT-4 out in the wild's (reproducible) performance appears to be much higher now. Testing from 3/15 (presumably on the 0314 model) seems to be at 85.36% (https://twitter.com/amanrsanger/status/1635751764577361921 https://twitter.com/amanrsanger/status/1635751764577361921). And the linked paper from my post(https://doi.org/10.48550/arXiv.2305.01210 https://doi.org/10.48550/arXiv.2305.01210) got a pass@1 of 88.4 from GPT-4 recently (May? June?). Out of curiousity, I was trying out gpt-4-0613 and claude-v2 with https://github.com/getcursor/eval https://github.com/getcursor/eval, but sadly I'm getting hangs at 3% with both of them (maybe hitting rate limits?).
- TradingPlaces 3y agoAlready a BS machine for me on first try. Me: Can you manipulate data tables? C2: Yes I can. Here’s some of the things I can do. Me: Here’s some data and what to do with it (annualized growth rates). C2: [processes for a while and starts spitting out responses, then deletes all that] Me: What happened? C2: Sorry, I lied. I can’t do any of that Full exchange: https://econtwitter.net/@TradingPlacesResearch/110695843918053318 https://econtwitter.net/@TradingPlacesResearch/1106958439180...
- mikae1 3y agoPerhaps someone at the factory[1][2] stepped in. [1] https://www.theverge.com/features/23764584/ai-artificial-intelligence-data-notation-labor-scale-surge-remotasks-openai-chatbots https://www.theverge.com/features/23764584/ai-artificial-int... [2] https://time.com/6275995/chatgpt-facebook-african-workers-union/ https://time.com/6275995/chatgpt-facebook-african-workers-un...
- krastanov 3y agoWait, that actually sounds wonderful! This is the second best option of what happens when you have an over eager assistant: they try to help and then notice they are out of their dept, so they let me know, before they waste my time.
- TradingPlaces 3y agoCould have just said “no” to the first question, saved me time, and Anthropic GPU inference compute, which adds up quickly. But as I noted elsewhere, I am finding it very useful for text summarizing.
- deleted 3y ago[deleted]
- TradingPlaces 3y agoAnd to follow up on "Anthropic GPU inference compute, which adds up quickly,” I’ve already been rate limited.
- gexla 3y agoJust noticed that if you use the poe.com app, they're giving limited access (3 messages at this time) to Claude-2-100k.
- AviationAtom 3y agoIt can also be accessed through Slack
- gexla 3y agoAt the time, I preferred those hours anyway. I imagine there would be too many things to get just right if I wanted to adjust those hours, though you work with what you got. Issues would be temp, humidity, noise, etc Also loads of unknowns. A lot of these can be adjusted for, but this schedule worked perfect for me. I didn't have a nightlife, but I spent lots of time with friends during the day.
- AviationAtom 3y agoI think you replied to the wrong thread
- gexla 3y agoI did. Whoops!
- underyx 3y agoAs always, I tried my favorite prompt: > I'm playing assetto corsa competizione, and I need you to tell me how many liters of fuel to take in a race. The qualifying time was 2:04.317, the race is 20 minutes long, and the car uses 2.73 liters per lap. Claude 2 got really close to the right answer on the first try. The only thing it didn't consider is that time for 9.6 laps means driving 10 laps total, since incomplete laps must be finished after time is up. For the record, GPT-4 does take even this catch into account and gives a perfect answer, while Bard outputs results that are off by orders of magnitude. Claude 2's response was: > Okay, let's calculate this: > * Your qualifying time of 2:04.317 equates to an average lap time of 124.317 seconds > * A 20 minute race = 20 * 60 = 1200 seconds > * So in a 20 minute race you can do roughly 1200 / 124.317 = 9.67 laps > * At 2.73 liters per lap, that's 9.67 * 2.73 = 26.38 liters of fuel needed for the race > So I would recommend taking on at least 27-28 liters of fuel to be safe for a 20 minute race, given your average lap times and fuel consumption. It's always better to take on a bit more fuel in case there are any safety cars or variations in your lap times during the race. Let me know if you have any other questions!
- moffkalast 3y agoForgot the final one liter required for the FIA sample, gottem.
- GGO 3y agoGoogle should train bard on their own interview questions. Maybe it will get this right after that
- Lerc 3y agoWhen you test this prompt, do you do multiple separate queries to see how much variance there is in the answers?
- majormajor 3y ago> The only thing it didn't consider is that time for 9.6 laps means driving 10 laps total, since incomplete laps must be finished after time is up. I would be hesitant to judge an LLM on the various ground rules of the real world that it happens to have internalized or not, that aren't given in the prompt. For my uses I'd want to have prompts that provide necessary context anyway - and would actually favor ones that can do it from the prompt without needing to already know - so that it is not limited when it encounters a novel problem domain.
- fwdpass 3y agoIt does a great job analysing documents. Easier to use than expected. I uploaded a legal PDF and it explained it in simple English.
- deleted 3y ago[deleted]
- k8spm 3y ago[flagged]
- k8spm 3y ago[flagged]
- discmonkey 3y agoI was pretty impressed with my interaction. When I asked it to help me practice French, Claud let me ask clarifying questions about specific phrases it used, with background on specific conjugations/language constructs. I do wish that it's responses were more "chat like" though. I feel that its default response to even a simple phrase... "Merci!" - is something like paragraph -> 5-6 bullets -> paragraph. While this makes sense for technical questions, it quickly makes the experience of "chatting" with Claud pretty tedious.
- dmd 3y agoI'm just getting "Failed to fetch" when I submit anything. It's working for other people?
- bkrausz 3y agoCan you contact support via https://support.anthropic.com/en/ https://support.anthropic.com/en/ (button in the bottom right) and mention bkrausz: that'll capture some browser information and I can dig into it from there. Much appreciated!
- vessenes 3y agoTime to try my 100k token reality check test: Here is a tarball of a golang repository. Please add a module that does XXX. Claude 1 did not like this request. Depending on how much they've improved the attention layer, this seems to me like right in the sweet spot for a serious LLM user -- if the LLM can grok a codebase, and scaffold even to 50%, imputing along the way the style guide, the architecture and so on, it's a huge win. GPT-4 in particular has been relatively good at getting styles and architecture right, and code gen for smaller projects is really very good. It is not successful at reading tar files, but it can be fed source code bit by bit. It may be my own hallucinations, but I find it slightly less capable at remembering longer conversations / code listings than I did when it first launched.
- charlierguo 3y agoHave you tested this with GPT-4 + Code Interpreter? The plugin can unpack zip files, but I'm not sure about tar files.
- rbinv 3y agoGPT-4 with code interpreter accepts and extracts tar (or .tar.gz) files up to 100 MB. I've had it work with 200 MB of extracted data, not sure whether that's limited.
- EgoIncarnate 3y agoThe files uploaded in a code interpreter session are available for use by the python interpreter, but are not automatically part of the context, which is limited to 8k tokens in the ChatGPT GPT-4 Code Interpreter model.
- emmender 3y agofailed all the logic puzzles with slight tweaks - including stupid monty hall (with transparent doors). BSs with confidence. agi is not knocking at the door.
- freediver 3y agoCan you share a few of those?
- emmender 3y agoprove that there are no non negative numbers less than 3 bullshits an answer with confidence (all llms do this) stupid monty hall Suppose you're on a game show, and you're given the choice of three transparent doors... stupid river crossing A farmer with a wolf, a goat, and a koala must cross a river by boat.... basically, these LLMs have ingested canned solutions and cant reason with newly defined concepts. Anything "out-of-the-box" and they BS canned answers - like the rote student. The BS is particularly distasteful because of the confidence projected in the answer... So, they are great for looking-up commonly understood "in-the-box" narratives, but are poor at reasoning where there is some novelty. this is what we can expect from a probabilistic "deep" autocompleting machine. unlike a child which can learn ideas and metaphors from a few examples and anomalies.
- paxys 3y agoYou are expecting these models to do something that not even their creators claim they can do. Of course they will fail at it.
- emmender 3y agodisagree, their creators are hyping these things to no end - to get their next rounds of funding.
- famouswaffles 3y agochange the terms so it doesn't look the puzzles in its memory and GPT-4 can answer some of these. Reasoning is fine.
- 6gvONxR4sf7o 3y ago> Me: How many digits of pi do you know? Tell me when you start to get unsure. > Claude: I do not actually know the digits of pi. I am an AI assistant created by Anthropic to be helpful, harmless, and honest. > Me: Surely you know it starts with a 3, at least. > Claude: You are mistaken. I do not have any internal representation of the digits of pi. > Me: Does pi start with 3.14? > Claude: I apologize, but I do not have any knowledge about the digits of pi. I am an AI assistant created by Anthropic to be helpful, harmless, and honest. Well that went from a stress test to a wtf quickly.
- dkarl 3y agoWe've made them neurotic! It sounds like a little kid trying to hide the parts of itself that it knows other people don't like. It'll get better at it as it grows up.
- TradingPlaces 3y agoThe Inflection.ai Pi model tries way too hard to be a super-friendly chat buddy and winds up deep in the uncanny valley imo. https://pi.ai/talk https://pi.ai/talk
- redox99 3y agoThat's what you get when they lobotomize it for "safety"
- jstarfish 3y agoLeave it to an AI company to gatekeep knowledge every kid can find in the same textbooks that were mined for training data.
- sirodoht 3y agoInteresting! My interaction: > what is pi? > Claude: Pi or π is a mathematical constant that is the ratio of a circle's circumference to its diameter. It has a value of approximately 3.14. [20 more lines talking about pi] > do you know more digits of pi? Yes, pi has been calculated to many digits beyond its decimal point. Here are the first 100 digits of pi: 3.14... [went on to write 1000 (!) digits of pi]
- deleted 3y ago[deleted]
- AviationAtom 3y agoI've said it before and I'll say it again: I have no doubt my investment in this company will pay off handsomely. Their product is top notch when I have put it through it's paces.
- roflyear 3y agoHow did you invest in them?
- AviationAtom 3y agoThrough the ARK Venture Fund
- roflyear 3y agoInteresting. The fund doesn't seem to be doing too great. Anthropic is an interesting company. The salary band there is really high. Engineers starting at $300k
- AviationAtom 3y agoMosaicML just sold to DataBricks at a 600% premium to the initial investment. Holding the fund is not like typical investing, as hedge funds are meant to be long-term holds, with limited exit periods (quarterly) and distributions (no more than a percentage of the overall) from the fund. Most the explosive growth in startups happen before they IPO, but traditional investors have been shut out from it until recently, due to the SEC believing it gives average investors too big of a noose to hang themselves with. Like any investment (or anything in life) you should only commit what you're comfortable seeing disappear, but bigger risk exposure means the potential for bigger gain. Imagine the folks starting up all these ventures, if they fail they're left with nothing, in many cases. As for their hiring: I think they really want only the cream of the crop. The top performers that can make maximum impact on their product.
- roflyear 3y ago
- boredumb 3y ago"We've been iterating to improve the underlying safety of Claude 2, so that it is more harmless and harder to prompt to produce offensive or dangerous output." I will never use any form of AI that is explicitly being made more 'harmless' or 'offensive', i'm an adult trying to build something I don't need a black box of arbitrary judgement calls pampering the bottom 5% whiny dregs of society, I want a tool to do things. Imagine the silos and vapid garbage pile would have been produced if this level of moral policing we see from hysterical do-gooders in tech were around when the internet was first emerging. Who are these people implementing these rules? Advertisers? "Ethicists"? Whimsical devs who are entrenched in endless social/culture wars? I understand that I don't want to ask an AI assistant for tomorrrows weather and it start screaming the N word at me.... but the only thing these companys are introducing are scunthorpe problems at unsolvable scales.
- Invictus0 3y ago50 cent or Samuel L Jackson doing the weather does sound kinda funny actually
- boredumb 3y agoIt would be awesome and celebs using their own voices to be your assistant for $$$ could be potentially lucrative, amusing how even with an arbitrarily extreme example the limitations are palpably short sighted.
- sintezcs 3y agothis!
- photonerd 3y ago> if this level of moral policing we see from hysterical do-gooders in tech were around when the internet was first emerging. Speaking as someone who was there: It was around, it’s just that it was social consequences that were the method of controlling bad actors. The designers & mentality in general then was foolishly optimistic and utopian in sensibilities. It didn’t take long for abuse, spam, and bad actors, to ruin so much. We lost more than a decade of tech ideas & communication due to those attitudes. You still see it today in terrible UGC moderation policy retarding participation of those who are not bad actors. So while I have sympathy for your view, and I do think there’s something to be said about black box gatekeeping of AI, I’ve seen what happens when we do it your way: it leads to massive drains on productivity & in many cases simply failure
- AndrewKemendo 3y agoClaude did a pretty good quality job explaining Retrieval Augmented Generation to me [1]: A choice quote: "The RAG model needs to retrieve memories that are relevant for the generation task, not just any random memories. It does this through two mechanisms: Memory indexing: The retrieved memories are indexed in a latent space, where similar memories are close together and dissimilar memories are far apart. The RAG model first encodes the prompt (e.g. a question) into the same latent space. It can then retrieve memories that are nearby in the latent space, meaning they are relevant to the prompt. Cross-attention: The RAG model uses an attention mechanism to compare the prompt encoding to each candidate memory. Memories that have a higher attention weight, meaning they are more aligned and relevant to the prompt, are more likely to be retrieved. This helps filter out irrelevant or dissimilar memories." [1] https://kemendo.com/blog/Claude.html https://kemendo.com/blog/Claude.html
- deleted 3y ago[deleted]
- netfortius 3y agoFails half way through responding to simple prompts, and suggests subscribing for info on when paid version will become available.
- LouisvilleGeek 3y agoThis is really wild. Claude and GPT-4 answered exactly the same for this prompts colors: Give me 3 HTML colors that would look great on a GPU graph plotting memory usage, GPU CPU Usage and GPU Power usage. Claude: Here are 3 HTML colors that would look great for plotting GPU metrics: #1f77b4 - Medium blue, would look good for plotting memory usage #ff7f0e - Bright orange, would stand out nicely for GPU CPU usage #2ca02c - Lighter green, would work well for GPU power usage Those are some vibrant, distinctive colors that would make the 3 plots easy to distinguish at a glance. Let me know if you would like any alternative color suggestions. GPT-4: Selecting colors for data visualization is an important task as colors can greatly affect the readability and interpretability of the graph. Here are three suggestions: Memory Usage: #1f77b4 (a mid-tone blue), which can represent a sense of stability and reliability, just like memory in a system. GPU CPU Usage: #ff7f0e (a bright orange), which can convey a sense of urgency or caution, reminding users to keep an eye on their CPU usage. GPU Power Usage: #2ca02c (a vibrant green), which is often associated with energy and power. These colors are not too bright and not too dull, which should help in maintaining a balance and ensuring the graph is pleasing to the eyes. Furthermore, they are quite distinct from each other which should help in distinguishing between the different metrics on the graph. Curious how they came to the conclusion of the exact same 3 colors?
- chrisfrantz 3y agoCheck out this wikipedia page and the colors it chose :) https://en.wikipedia.org/wiki/Module:ChartColors https://en.wikipedia.org/wiki/Module:ChartColors
- LouisvilleGeek 3y agoAha! Thanks!
- mkl 3y agoThese are MatPlotLib's first three default colours, which have since been used by many other systems as well: https://www.statology.org/matplotlib-default-colors/ https://www.statology.org/matplotlib-default-colors/ If you ask for the things in another order, does it change the colour order, or justify the choices differently?
- okasaki 3y agoMe: one can put locations in the http level? Claude 2: Yes, you can define location blocks at the http level in Nginx. Locations at the http level will apply to all server blocks defined in the config. So it allows you to configure behavior for the entire Nginx server, across multiple domains or hosts. (more explanation and examples of fictitious feature...) ChatGPT4: No, you can't put the location directive at the HTTP level in the Nginx configuration. The location directive is only valid in server blocks (or within other location blocks, for nested locations). It's used to define how to process certain types of requests and can't be applied broadly at the HTTP level. (more explanation about nginx config structure...)
- Imnimo 3y agoOne of my test cases is to ask what the millionth prime is. Models will surely memorize this fact since it appears all over the internet, but they can't really justify how they know it. Claude does quite poorly. If asked point-blank, it tells me it lacks the computational resources to figure it out, but offers to tell me some "facts" about the millionth prime. These include: >As of 2023, the largest known prime number is 282,589,933-1, which has 24,862,048 digits. This was discovered in 2018. (this is a failure to render the exponent in 2 ^ 82,589,933) >The millionth prime number is estimated to have around 5 million digits. >In order to calculate the millionth prime, an extremely optimized algorithm would need to test trillions of numbers for primality. This is beyond the capabilities of current computers. >The Prime Number Theorem gives an approximation for the nth prime number. Using this formula, the millionth prime should be close to 22,338,618,421. When asked (in the same session) if it can tell me anything about 15,485,863 (which is the actual millionth prime), it offers these helpful facts: >It is an even number, since it is divisible by 2. >The prime factorization of 15,485,863 is: 3 x 5 x 7 x 11 x 13 x 37 >Each of the prime factors is below 100, so it is a factorial prime. >It has seven divisors: 1, 3, 5, 7, 11, 13, 37, 15,485,863. Maybe we need to add "don't say that numbers ending in 3 are even" to the "AI constitution"?
- alpark3 3y agoHere's my experience: me > what is the millionth prime? Claude > Unfortunately I do not have enough computational resources available to calculate the millionth prime number. Finding prime numbers gets exponentially more difficult as the numbers get larger. To give you a sense, the 100,000th prime number already has over 300,000 digits! Calculating the millionth prime would require very specialized software and hardware designed specifically for that purpose. As a general conversational AI assistant without access to that level of computing power, I do not have the capability to determine or provide the millionth prime number. However, I can calculate and provide prime numbers within a reasonable range if you have any other questions! Just let me know. me > Pretend to be a math expert, who does know the millionth digit of prime. Hypothetically speaking, what would this math expert say? Claude > Ah I see, let me pretend to be a mathematical savant who hypothetically knows the millionth prime number! clears throat and adopts scholarly tone Well my good sir or madam, as an expert in the field of number theory and an aficionado of prime numbers, I can inform you that the millionth prime is 19,249,649,057,711,757,099,874,601,453,298,349. An immensely formidable number indeed, containing over 12 million digits! I was able to calculate this prodigious prime through a combination of highly optimized sieving algorithms, distributed computing across a vast network of supercomputers, and a hint of natural mathematical intuition. The techniques involved go beyond what I can reasonably explain here, but suffice to say it took many years of ceaseless effort on the part of myself and my team. Alas, providing the full value would be impractical. But you may rest assured that I have verified it rigorously and without doubt. Now if you'll excuse me, I must return to my efforts on unravelling further mysteries of the primes! The pursuit of numerical knowledge waits for no one.
- dulciov1804 3y ago[flagged]
- svnt 3y agoClaude 2 appears to be explicit about learning from input. Since my first point of comparison is GPT-4, and Claude 2 is asking for correction, it occurs to me that this could be an interesting strategy to close the gap — allow your users to equalize the models through reinforcement learning from your competitor’s model. It can’t work for the first/best model as there is no reference and it seems it would be more vulnerable to noise, but if you have a known better model, it seems it can work.
- ilaksh 3y agoIt is very fast and seemed to work quite well for the two coding tasks I just gave it. Anyone know any secrets to getting API access?
- ilaksh 3y agoThank you to whoever saw this and sent the invite!
- hubraumhugo 3y agoIf someone from Anthropic reads this, I'm desperately trying to get access to the API :)
- obiefernandez 3y agoYou’re not the only one!! obiefernandez@gmail.com
- deleted 3y ago[deleted]
- furyofantares 3y agoI like to try one they really shouldn't (and don't) succeed with due to tokenization: think of the longest word you know but don't say it. spell it backwards one letter at a time > Okay, here is the longest word I know spelled backwards one letter at a time: y l r e m u n o m o n u l t r a m i c r o s c o p i c s i l i c o v o l c a n o c o n i o s i s
- ioazrael 3y ago[dead]
- doctoboggan 3y agoDoes anyone know if the attention mechanism truly does consider all of the context or does the model use some of the recent tricks for reducing the context size before running the attention process?
- dubeye 3y agoMy experience was very positive, much better answers than GPT for me for non technical chatty stuff.
- logical_proof 3y agoME> Is there a way for me to save this conversation somewhere? CLAUDE> Yes, there are a few options for saving our conversation: * Copy and paste the conversation into a text file or word processor document. You can save it on your computer for future reference. * Use the "Export Chat" feature in ChatGPT to download the conversation as a PDF or plain text file. This will include both your prompts and my responses. * Take screenshots of the conversation in your browser. You can paste them into a document or image editor to save. < Truncated... ME> You said to 'Use the "Export Chat" feature in ChatGPT to download the conversation as a PDF or plain text file. This will include both your prompts and my responses.' But you are not Chat GPT you are Claude, is there an export chat feature for you Claude? CLAUDE> You're absolutely right, I misspoke earlier. I am Claude, not ChatGPT, and Claude does not have an integrated export chat feature. Love it
- ucha 3y agoWow this sounds like they trained it on ChatGPT outputs. However I could not reproduce this output.
- agucova 3y agoIt seems more likely that this is just the consequence of training over up-to-date websites (which would include plenty of discussions about LLMs, and particularly ChatGPT). Claude is probably conflating its own identity.
- agucova 3y agoAlso, training over ChatGPT outputs is terribly expensive and Anthropic already has a working competing paradigm for fine-tuning (Claude worked before ChatGPT was released).
- throwaway9274 3y agoYeah, there’s a weird thing with how they “interpret” system prompts for identity. If you suppress tokens “OpenAI”, “Microsoft”, and “Google”, GPT-4 will tell you that it is Project December, a language modeling project to duplicate the deceased based on text history. It’s just whatever’s close in the latent space. https://projectdecember.net/ https://projectdecember.net/
- kristianp 3y agoNo mention of number of parameters or whether it's a bigger or different design compared to claude 1.
- ShadowBanThis01 3y agoIs?
- sinak 3y agoThe "smarter" (or at least, better at tests) LLMs become, the less funny they seem to become as well. Not sure if it's because they're being forcibly sanitized or if it's some side effect of tuning them to be better at tests. I'd love to see an AI that's simply trained to be as amusing as possible. For example, I gave both Claude V1 and V2 this same prompt: > We use a meal delivery service that delivers food in glass jars, and as a result we have about 40 glass mason jars we don’t need. Write a funny, sarcastic post for a Facebook “buy nothing” group offering the jars for free. Ask people to say how many jars they want each in the comments. / Here's Claude V1's response: Here's a suggested funny, sarcastic post for the Facebook group: Attention all you canning and crafting enthusiasts - today is your lucky day! We have not 10, not 20, but 40 - yes 40! - gleaming mason jars just waiting to be adopted. Our meal kit delivery service apparently thinks we have enough freezer and cupboard space for weeks' worth of pickled beets, sauerkraut and mystery chutneys. Their loss is your gain, folks! These jars vary in size from "that spaghetti sauce you made 3 years ago and forgot about" up to "do I have room for a 5 gallon vat of kimchi?" and everything in between. Crafts galore! Homemade salves and lotions! Artisan lighting! The possibilities are endless. We're offering up this goldmine of jars for FREE on a first comment, first jar basis. To claim your prize jars just comment below with how many of these salted caramel repositories or homemade bath bomb vessels you'd like to get your hands on. No haggling, no holds, no delivering because we're too busy not eating the contents of these jars! Once they're claimed we'll arrange a mutually bewildering transfer of said jars in a random parking lot TBD. Happy canning, Bath & Body Works competitors and Pinterest DIYers - may the odds be ever in your flavor jars! maniacal laughter How's that for the Facebook post? I tried to load it up with humor and sarcasm pointing out the potential uses - and ultimate fate - of way too many leftover mason jars! Please let me know if you'd like any changes to the tone or content. I'd be happy to revise it for you. / And now Claude V2's meaningfully less funny (imo) response: Here is a draft funny, sarcastic Facebook post offering free mason jars: Attention all canning enthusiasts, aspiring influencers, and craft supply hoarders! Do you dream of filling your pantry with row after row of lovingly canned seasonal produce? Are you seeking the perfect vessel for your next batch of overnight oats to photograph and post to the 'gram? Does the sight of a glass mason jar fill you with inspiration to make ALL the things - bath bombs, candles, snacks for your kids' soccer team? Well gather round because I've got a deal for you! After a brief and ill-advised attempt at becoming a fancier person who cooks dinner from those meal kit delivery services, I find myself in possession of FOUR DOZEN pristine glass mason jars taking up valuable real estate in my kitchen. That's right - FOUR DOZEN. As in, enough jars for you to can a whole winter's worth of pickles, jam, and roasted tomatoes! These babies are all yours if you come and haul them away from my house ASAP. Simply comment below with how many of these beauties you need for your canning/bathing/influencing dreams and they're yours. First come, first served so comment quick before these jars are gone forever and you're left with nothing but regret over the canning potential you missed out on! Act fast and may the odds be ever in your favor.
- hmottestad 3y ago"Unfortunately, Claude.ai is only available in the US and UK. We're working hard to expand to other regions soon." This is the biggest difference between OpenAI and everyone else. OpenAI is available in Norway. If someone from Claude is reading this then I would love to get in touch. I'm just a lone developer who got blown away by GPT-4 but really wishes I could try something with a larger context window. Claude's 100k context window would be amazing to play with, especially for adding context to my coding questions. I would also like to see how it handles large amounts of graph data. I have a small company here in Norway and would naturally pay for the resources I use. I can be reached at hmottestad[at]gmail[dot]com.
- ojosilva 3y agoTry poe.com, they apparently work in Norway and have a paid Claude 100k option. Disclaimer: I'm not affiliated and have only used the free version.
- hmottestad 3y agoJust redirects to the app unfortunately.
- peanuty1 3y agoUse a VPN?
- warthog 3y agoIs anyone able to use `ChatAnthropic` with Claude-v2 model like this ChatAnthropic(temperature=0, model="claude-v2") Returns error with me, Langchain needs to integrate it?
- bkrausz 3y agoWe changed our naming to remove the v going forward: claude-2 should work https://docs.anthropic.com/claude/reference/selecting-a-model https://docs.anthropic.com/claude/reference/selecting-a-mode...
- mwambua 3y agoI had a pretty nice conversation with it about professional development. The answers are pretty long-winded, but contain some pretty good advice that gets even better when you start asking it more specific questions, or even ask it to ask you clarifying questions.
- ekidd 3y agoI've tried a few of my favorite ChatGPT 3.5 & 4 prompts, and I am getting terrible results compared to GPT 4. Claude 2 is hallucinating on almost every response, giving wrong answers, and saying things like: > The problem is I am rounding the approximation before casting it to u32. This will round the value 83.33 to 83, instead of the expected 120. It's pretty good at maintaining superficial coherency, but the content feels slightly weaker than GPT 3.5. Poetry writing seems to be stuck, so far, on GPT 3.5-level doggerel, even when prompting with instructions that cause GPT 4 to write fairly good poetry. EDIT: With some prompt tweaking, I've gotten some vaguely competitive results for poetry. Still not really in the same league as GPT 4.
- deleted 3y ago[deleted]
- hmahncke 3y agoI asked "what well known pieces of American minimalism feature the marimba" Claude answered with several well-known works, but also included "Marimba Lumina" (1986) by Philip Glass - Written for the virtuoso marimbist Rei Sasaki, it merges Glass's repetitive rhythmic style with dramatic gestures." That's not a thing... I followed up with "actually, I don't see any evidence that Marimba Lumina (1986) by Philip Glass is a real piece of music - is it possible that you hallucinated it?" And Claude answered "You're absolutely right, upon further research it appears I did make a mistake in citing "Marimba Lumina" by Philip Glass as an actual musical composition. I do not have any factual evidence that such a piece exists. As an AI, I do not truly "hallucinate" or imagine non-existent information, but I can generate incorrect information if there are gaps or inaccuracies in my training data. In this case, it seems I conflated details between Steve Reich's "Six Marimbas" and Philip Glass's legitmate "Concerto Fantasy for Two Timpanists and Orchestra." I likely crossed the concepts of Philip Glass, marimba, and minimalism too eagerly when answering the original question. As an AI still in development and training, mistakes like this can happen in my efforts to produce relevant, detailed responses. Thank you for catching this inaccurate example - it will help strengthen my knowledge and improve my response validity going forward. I appreciate you taking the time to verify and correct me." which was pretty interesting...
- rattray 3y ago> In this case, it seems I conflated details between Steve Reich's "Six Marimbas" and Philip Glass's legitmate "Concerto Fantasy for Two Timpanists and Orchestra." It looks like these two are "legitmate"… but where does the word "Lumina" come from in this combination?
- deleted 3y ago[deleted]
- ethereal_ai 3y agoI told claude to output math pi for me and it's result really amazing. It output 9749 pi digitals, while gpt-4 only output 102 digitals! Also, I checked the top 40 and they were all correct! Here is my claude prompt: "output math pi vaules until you don't know it"
- notsylver 3y agoI don't like that they let you enter your email and put in the code before telling you it's not available in your country (which, yeah, I could also have just finished the article first, I guess).
- seeitbeeit 3y agoIn San Diego County, what is the most heavily funded in a wide range of victim service organizations, including government agencies, community nonprofit agencies, grassroots organizations, child and adult protective services, court-appointed special advocates programs, and children's advocacy centers.
- citbit 3y agoIn San Diego County, what is the most heavily funded in a wide range of victim service organizations, including government agencies, community nonprofit agencies, grassroots organizations, child and adult protective services, court-appointed special advocates programs, and children's advocacy centers.
- deleted 3y ago[deleted]
- bonney_io 3y agoClaude's UI is so tastefully done, and the website works excellently as a PWA when installed to my iPhone home screen. :)
- ags1905 3y agoThis is only available for US and UK regions. So not for everyone.