7 ms·
Qwen3.8 Max now ranked as the best overall model by agentic index
- embedding-shape 2mo agoStrange that the page https://artificialanalysis.ai/agents/coding-agents https://artificialanalysis.ai/agents/coding-agents doesn't even mention "Qwen" once if it's now the "best" according to one of their one index?
- scrlk 2mo agoDifferent benchmarks: > Artificial Analysis Agentic Index: Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, Tau³-Banking) > Artificial Analysis Coding Agent Index v1.3 incorporates 3 benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA Qwen3.8 Max is 55.4 on the Agentic Index but hasn't been tested for the Coding Agent Index.
- apitman 2mo agoLooks like coding agent is model+harness. There are far fewer models represented on that page. I believe "agentic index" is still the metric to look at for coding performance. I could be wrong about that though.
- Bootvis 2mo agoIndeed, and this Qwen 3.8 max specific page: https://artificialanalysis.ai/models/qwen3-8-max https://artificialanalysis.ai/models/qwen3-8-max Doesn't have the claim either. Clickbait?
- petu 2mo agoThis page has it, scroll to "Intelligence" header (not the highlights one, but second on the page / with black square) and click "Agentic Index"
- Bootvis 2mo agoSo the original link should be: https://artificialanalysis.ai/models/qwen3-8-max?intelligence=agentic-index https://artificialanalysis.ai/models/qwen3-8-max?intelligenc... Even then, this seems a much more marginal win than the headline suggested to me.
- amelius 2mo agoAccording to those graphs, Grok 4.5 appears to be the most cost-effective model.
- user43928 2mo ago$0.05 per task, Intelligence Index score 52 -> GPT 5.6 Luna max $0.36 per task, Intelligence Index score 56 -> Grok 4.5 high $1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max $0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol xhigh $1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh
- artemisart 2mo agoThey didn't run all benchmarks. It's the best in AA agentic index (GDPval-AA v2, ³-Banking) but not coding index (DeepSWE which is missing, Terminal-Bench v2.1 they have 81% vs 90% for Sol, SWE-Atlas-QnA missing).
- moritzwarhier 2mo agoDoes "artificial analysis" mean what it says? Dubious. But: I've been very impressed by the larger Qwen Models, and a brief try of Kimi also impressed me. A lingering sense of quality degradation when going deep remains. But that's not an accusation: they seem to be hitting the compute/quality tradeoff extremely well. And on-prem capability is simply irreplaceable. Apart from all the innovations that were driven by the strive for this optimization: quantization, "distilling" (without obvious mad-cows-disease)... I think China was an invaluable player in this progress. Intuitively, I'd even go so far to speculate that LLaMa wouldn't exist without the competition.
- syntaxing 2mo agoI am so excited for Qwen 3.8 27B. It’s a shame how slow prefill (~3-400) is on a strix halo but it’s such a good model for agentic tasks.
- tarr11 2mo agoWhat type of agentic tasks are you using it for (eg how complex)?
- syntaxing 2mo agoFor personal stuff, I use it with AnythingLLM. It replaced any Google search for me. For coding, I run opencode though I have been debating switching to Pi. I would argue it’s at Sonnet 3 level.
- CamperBob2 2mo agoHow are you running it on a Strix Halo? The weights aren't out yet, are they?
- 13rac1 2mo agoI interpret @syntaxing as meaning they are looking forward to running Qwen3.8-27B, but are frustrated by prefill times with other models, such as Qwen3.6-27B.
- syntaxing 2mo agoI meant Qwen3.6. Unsloth supposedly has early preview of the model and the VRAM requirement is the same so most people expect similar model size and type.
- LoganDark 2mo agoI find that 35B-A3B is much easier to run on my M4 Max (both prefill and generation)
- markasoftware 2mo ago
- drnick1 2mo agoWhy does an open weights model cost nearly the same as GPT5.6? $1.14 vs $1.23 on the cost index. Since you can't presumably run this on your own hardware given the model size and hence gain other things like privacy, I don't see any reason to move away from GPT at this rate.
- jazzyjackson 2mo agoRunning a large model on rented GPU is still meaningfully more private than handing your chat logs over to FAGA
- TheCycoONE 2mo agoThe acronym is new to me: Facebook, Anthropic, Google, openAi?
- eli 2mo agoIt's not enough that it's better? Many providers will host it and will compete on price. It also can't easily be taken away because one company (or one government) decides they don't want it around any more. People can fine-tune it for particular workloads.
- drnick1 2mo ago> It's not enough that it's better? It's barely better, and barely cheaper, not really enough to challenge the status quo IMO. Half the price for basically the same performance would be a much stronger value proposition.
- ux266478 2mo agoWhat status quo? Just look at Openrouter's rankings: https://openrouter.ai/rankings https://openrouter.ai/rankings Things change radically month to month. Nobody is remotely close to capturing the market or having any kind of stability over time. People move around quite a lot, often to sidegrade within a generation. Just playing fly on the wall with discourse would be enough to tell you all of this, even without the data to back it up.
- eli 2mo agoI believe it. It's extremely good at troubleshooting. I gave Qwen and Kimi K3 the same annoying, complicated, intermittent bug to track down. Kimi did a bit better in understanding the existing code, but Qwen built some diagnostic tools and did an excellent statistical analysis on the log data. Qwen got way closer to the truth. I'm very much looking forward to their forthcoming smaller model Qwen 3.8 releases. A version that can easily run locally would be great.
- comboy 2mo agoHow CLI are you guys using for qwen and kimi?
- eli 2mo agoI use https://pi.dev/ https://pi.dev/ which works fine out of the box but is fairly minimal and intended to be customized. There are many extensions. OpenCode or oh-my-pi might make more sense if you just want a batteries-included agent. You can also make Claude Code work with other models without too much work, but I think that's asking for headaches.
- trey-jones 2mo agoI used claude with GLM and it's easy to set up, just hard to find the documentation. No headaches really, unless you want to use it against multiple different APIs.
- Gooblebrai 2mo agoIs there any subscription of any kind for Qwen? Or via Pi.dev needs to be used with API credits?
- iAMkenough 2mo agoMy first web search turned up this as the top result https://www.alibabacloud.com/help/en/model-studio/coding-plan https://www.alibabacloud.com/help/en/model-studio/coding-pla...
- aliljet 2mo agoIs there a path to distill this model to do very specific things? Like a RAG strategy for a small (or even large) corpus?
- Alpha3031 2mo agoDepends on what you want to do. Some task specific models can be trained with a few ten or hundred thousand training examples so you can use a bigger model to produce synthetic training examples and then fine tune a smaller student model. I think that's the usual process. Whether you'd get acceptable performance this way depends, as mentioned, on what you're trying to do and what you'd consider acceptable.
- teravor 2mo agoonce you are able to get the full probability distributions per token you can distill it on specific domains. distilling without that isn't generally a good idea unless you have invested millions in the requisite infrastructure.
- mococa 2mo agoGo China!
- SwellJoe 2mo agoI find that surprising. I've been trying it on several projects and have found it's pretty sloppy. It leaves stuff broken, doesn't reliably write tests to check its own work unless explicitly prompted, misunderstands the assignment, etc. It is smart and reasonably quick but not reliable.
- dyauspitr 2mo agoIt’s because they’re doing some sort of combined score of intelligence, speed and cost. On pure intelligence it doesn’t even show up in the top 10.
- superfrank 2mo agoI've come to the same conclusion over and over with all of the Chinese models that have been claimed to be catching up with OpenAI's and Anthropic's frontier models (Deepseek 4, GLM 5.2, Kimi K3). At their best, I think they're closing in on Opus and GPT, but they're incredibly inconsistent and the variance in output quality is much higher than the best from any of the Anthropic or OpenAI models from the last few generations. The only way I can describe it is that it feels like a lack of intuition with the models which means I find my self needing to write longer prompts or have more back and forth to get them to do what I want from them. To give an example, I have a saved prompt that I use as a sanity check on some data I'm storing. It reads about 50 rows from a DB and matches them to the UI and makes sure the data is displaying correctly. I've been using this with GPT 5.5 and now 5.6 for a few months and running it a few times a week with no issue. Sometimes I'll run it multiple times in a single chat if I notice bad data (run it, fix thing, run again, fix another thing). I recently tried to switch to using Deepseek v4 (first flash and then pro) and while both did the task just fine, both would do things like change the response format from one message to another in the same chat or randomly decide to omit things it didn't think were relevant. At one point I ran the prompt, fixed some bad data, and then said "Okay, I fixed row 7, run {prompt} again" and so it decided to leave row 7 out of the response. A few times the first message would contain a table and then the next run in the same chat would contain the data in a bulleted list. None of those are major issues and all could be solved with a bit more rigor in my prompting, but for me it makes them harder to work with. Those examples are a bit trivial, I think they're the easiest way for me to illustrate the gaps I see with them.
- brcmthrowaway 2mo agoCould someone like Apple be playing the long game - Good Enough(tm) intelligence will eventually fit in our pocket and homes?
- LPisGood 2mo agoAlmost surely. Apple is extremely well positioned to take advantage of this over the next decade.
- colingauvin 2mo agoDS4 Flash Q2/Q4 mixed quant fits on a DGX Spark (a $4000 device which is not particularly unheard of expense for Apple customers), and is indistinguishable for me from Opus for my personal daily use/assistant benchmarks[0]. [0]https://humanparadox.org/local-vs-frontier-benchmarks-for-my-personal-assistant/ https://humanparadox.org/local-vs-frontier-benchmarks-for-my... - note here I tested Q8 but have found no difference at lower quant.
- dofm 2mo agoIndeed. I like using Macs mostly, and the bargain M1 Max MBP I am using for local LLMs is a fabulous experimentation platform and does loads of other stuff well, so I am in no rush, but if I reached the point of buying dedicated hardware for an LLM, I'd be looking at the DGX Spark machines.
- kyxsc 2mo agoApple is already doing this... they worked with Gemini to distill the model into a smaller one that fits on your phone. If you have iOS 27 Beta, you're already using this
- notatoad 2mo agosort of. they have a local model, it does some things. they also have significant cloud infrastructure backing it, and most tasks are going to be sent off to the cloud for processing, not be handled by the on-device model. Siri is not on-device by any stretch of the imagination.
- sirbor 2mo agoQwen is the way to go
- dyauspitr 2mo agoIt doesn’t even show up in the raw intelligence index, so how could it possibly be the best?
- quirino 2mo agoA couple days ago they had published an overall score of 53 for this model, but that was removed and today it returned with a score of 56. I wasn't able to find an explanation from them. Anyone knows what happened?
- Art9681 2mo agoA wire transfer happened.
- ignoramous 2mo agoThe kind of distillation guaranteed to work.
- whwhyb 2mo agoaccording to them: > launch traffic hit our public API endpoint harder than expected, causing intermittent instability. https://x.com/QwenDevs/status/2085279963654275247 https://x.com/QwenDevs/status/2085279963654275247
- steve-atx-7600 2mo agocurious about methodology. ive seen them post results for claude/codex when they only ran over benchmarks 3 times per model...
- onomojo 2mo agoAny benchmark showing Opus 5 as the best just loses credibility for me. Anyone who's actually used Opus 5 daily knows what I'm talking about.
- cesarvarela 2mo agoIt is infuriating to interact with, but it is also first in many blind test leaderboards on LLMArena
- copperx 2mo agoI'm dumbfounded to see Opus 5 making SO MANY mistakes in coding simple stuff. Most times, Fable 5 comes out to be cheaper because it nails so many things much quicker than Opus 5.
- garciasn 2mo agoI have Fable plan and Opus implement. I haven't had any major issues working this way; however, Opus does seem plain fucking stupid compared to what I experienced with Sonnet previously.
- aenis 2mo agoI do the same, and generally have good results, but it does stupid things with gusto. I'd open a blog with "weird things Opus did". Today it launched a swarm of cpu-hogging processes to test if the widget showing machine and I/O load is rendering nicely and correctly. The test went fine, but it was no longer able to kill those processes since they were really effectively hogging the CPU in various ways - being diligent, some of them were hogging CPU, some were murdering the SSD, some were pounding on the network adapters. Took me 30 mins to recover the machine to a working state without killing the meaningful, messy, in-flight sessions i had going on on other projects.
- petesergeant 2mo ago> however, Opus does seem plain fucking stupid Infuriatingly so, in a way I don't remember Opus 4.8 being, but maybe I've just been ruined by Fable 5.
- petercooper 2mo agoHopefully this boils down to the smaller versions they've teased. In my experience, Qwen models are the closest to the "less knowledge, more intelligence" (yes, the two are hugely correlated!) ideal some tool-dependent tasks need. Even the 3.5 2B can be easily prompted to always lean on tools and not jump to false conclusions (although its actual coding skills are abysmal, as you'd expect).
- quotemstr 2mo ago> less knowledge, more intelligence People produce such models by over-RL-ing smaller models on math and coding tasks. I've found the results capable of neither innovative work nor thinking outside the box. They're straight-A students raised by tiger moments who never let them play freely for hours in the dirt. Perhaps you could say such models are skilled --- but intelligent? Not by my measure. People and AIs alike need diversity of experience and a broad liberal arts education to see hidden connections between fields and make real advances.
- DC-3 2mo agoIt's amusing to me that AI has become sophisticated enough that people have started being racist to it.
- petercooper 2mo agoI agree with you to an extent, but you have certainly given me food for thought. Sticking to LLMs, they seemingly get their intelligence (whatever that really means) from building models rich with knowledge, so you could have a point. But Qwen models seem to be particularly good, even at small model sizes, at maintaining both their own knowledge while acquiescing to and integrating external information in the moment.
- looksjjhg 2mo agoThat took what 2 years? I love how the chip ban made them more efficient
- Footprint0521 2mo agoFacts lol, now all the Chinese models are 1/40th of the cost for the same intelligence
- ben8bit 2mo agoHaven't tried this yet, but going to soon! I have to wonder what happened at Anthropic. We've cancelled our subscription in favor of OpenCode & Codex. Sol is just so good & OC goes so far for every $ spent. Claude's become a pain to work with - average output with an annoying personality. Who knew this would be an issue even a year ago? In any case, loving the stuff from the Chinese models!
- tomComb 2mo ago> an annoying personality I was with you until there. Qwen and the OpenAI models are great, aggressive agents, but they’re not as good as the anthropic models for human interaction. They just don’t have the subtlety, understanding, or attention to detail.
- ben8bit 2mo agoReally? I've heard so many other people complain about this recently. And maybe it's possible that it's the prompt style even. But interesting that it's not across the board.
- colingauvin 2mo agoClaude 4.5/4.6 - absolutely agree. Fable 5? From my (limited) testing, also reasonable to interact with. Opus 4.7/4.8/5? Absolutely smug and antagonistic and preachy. I'm constantly fighting with it to stop fighting me and accept that I occasionally know better. It's really frustrating to spend so many tokens of such an expensive model arguing with it.
- quadruple 2mo agoyeah, it's honestly amazing just how badly anthropic managed to screw up something in claude after 4.6. its night and day, and every time I see claude doing something stupid, i instantly realize I was accidentally on Opus 5
- theropost 2mo agoAnthropic is a bit nuts, I had $260 of credits on my max account for the extra usage the other night. It was expiring, so I figured I'll fire up an agentic swarm to deep dive and make some deep changes to some old cold bases.. literally 25 minutes or less, $260 burnt, it didn't get get into the implementation, just wrote a ton of useless plans for the most part. It really opened my eyes to what they expect to charge people.. wayyyy overpriced.
- polishdude20 2mo agoYou should just spend those towards a cursor subscription.
- cortesoft 2mo agoIt’s crazy how different the credit cost and subscription cost are. With the $200 subscription, I can have Fable on ultracode working for hours and not dent the usage limits.
- AlexandrB 2mo agoVCs are footing the bill for that $200 subscription.
- ericd 2mo agoThey have something like 80% gross margins, are at a $100B/yr ARR, and are growing at 10x per year... If that keeps up, they're going to be doing more revenue than Google in a year ($400B ARR, 20% per year growth)
- dexwiz 2mo agoHow can you sanely project the last 12 months forward? We have seen a huge uptick in usage. Last summer AI was a toy to most devs, now every enterprise developer I talked to uses it every day. Coding agent providers are surely going to hit market saturation in the near future.
- indiantrains 2mo ago[dead]
- d2p 2mo agoI clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot. Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2. I have screenshots of both. The description above the chart is the same in boh cases: > Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking) What happened? How can the scores change so much in a few seconds?
- h14h 2mo agoThey JUST updated their methodology: https://artificialanalysis.ai/methodology/intelligence-benchmarking https://artificialanalysis.ai/methodology/intelligence-bench... Edit to provide AA's article explaining it: https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1-1 https://artificialanalysis.ai/articles/artificial-analysis-i...
- ahartmetz 2mo agoFixed the result, eh? In both senses of the word.
- gpt5 2mo agoWhat was the change?
- johnnyApplePRNG 2mo agoI have been suspicious of these AI leaderboard sites for some time now, and this only increases that suspicion.
- splatzone 2mo agoCan someone please explain what changed, when it happened, and whether it was surreptitious?
- torginus 2mo agoIn that case they should clearly label that this is a new benchmark.
- atemerev 2mo agoWell, that's the bad index then. It is barely usable in my opinion compared to other Chinese frontier models.
- ramon156 2mo agowhich one of the other chinese frontier models is better?
- h14h 2mo agoThis has me hopeful for Qwen3.8-27B!
- camnora 2mo agoQwen is just crushing it overall. I regularly use 3.7-flash for everyday coding needs and it gets the job done.
- seizethecheese 2mo agoOpus is still first in Intelligence Index followed by Fable, GPT 5.6, Kimi K3 then Qwen 3.8 max. https://artificialanalysis.ai/#intelligence https://artificialanalysis.ai/#intelligence Our leaderboard combines Arena ELO, AA Intelligence index, latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max #4 GPT 5.6 Sol Source: http://pellmell.ai/leaderboard http://pellmell.ai/leaderboard. This jumps around a lot based on the top throughput and latency of whatever provider happens to be best at the moment.
- d4rkp4ttern 2mo agoAll these "intelligence" benchmarks miss something extremely important when using an LLM in a code-agent harness: How it communicates with you about what it did. Opus-5 is practically unusable (for complex tasks) in this sense - its updates are voluminous, and dense with cryptic language (there are numerous reddit threads complaining about this, so it's not just me). I often have to ask it to re-state concisely in plain terms. For a fairly gnarly task, after fighting with with Claude-Code + Opus-5, I ported my session to Codex + GPT-5.6-sol, and it was like a breath of fresh air. Arguably a key aspect of intelligence is concise, clear communication, and current benchmarks miss that, at least as far as I'm aware. I would think some arena-type benchmarks where humans rate responses would measure this, though I'm not sure which those are.
- moffkalast 2mo agoDamn I thought it was my extra instructions, I swear everything it writes is in some shorthand with direct references to variables that literally nobody could figure out unless you literally just wrote that code 5 minutes ago. I had it stop writing comments altogether cause it was always four lines of complete and utter nonsense, and it doesn't even obey that rule half the time. Despite doing an extensive back and forth to make a complete plan, 5 seconds into the implementation it changes its mind and makes another assumption, adding some extra thing that tends to break the entire approach and needs follow-ups to repair or cleanup. Instruction following is basically non-existent compared to Fable, it just does whatever the fuck it wants.
- esafak 2mo agoIt is also the most expensive open source frontier model, per task; cf. Cost per Intelligence Index Task. If it is as good as the benchmarks indicate it bodes well for Qwen and China. For my part, I'll pass; it is not on the Pareto frontier.
- Fordec 2mo agoAnthropic have a real fight on their hands now. The competition is no longer 6 months behind, it's 6 days. If this had come out two or three weeks earlier this would be an absolute market leader on both quality and timeline.
- brettgo1 2mo agoOut of curiosity, what's currently the best model I can use locally?
- daemonologist 2mo agoWith an unlimited budget, Kimi K3 (which is quite comparable to this Qwen Max imo). With a normal budget/a PC you might already have, probably Qwen 3.6 27B.
- arjie 2mo ago$500k - Kimi K3 (maybe $250k? Haven’t done this one) $25k - DSv4 Flash $4k - Qwen 3.6 35A3B Q5 $1k - Qwen 3.6 27B Q4 Some people prefer the sense over the MoE YMMV.
- apitman 2mo agoThese numbers look about right based on my experiences as well. Though for a single user I think 2x DGX Spark (~$10k) runs DSv4 Flash fairly well right?
- colingauvin 2mo ago16 DGX Sparks can run K3 at a reasonable TPS. So that's $64k. 2 DGX Sparks can run DS4 at 1 million context with 50 TPS so that's $8k. 1 A4500 can run 35A3B. Those are about $1200 new. 27B actually takes more hardware to run than 35B because attention is done differently I believe and therefore KV Cache takes a lot of space. It will run on an A4500 but it's slow and context will be like 32k.
- OsamaJaber 2mo ago[dead]
- proxyscore 2mo agoDoes it matter, it's all non deterministic bs ware and deepseek is eating the Americans lunch
- bonoboTP 2mo agoI distrust any benchmark where Opus 5 beats Fable 5.
- jjcm 2mo agoChina has caught up is the main takeaway here. The SOTA models are so close that it's really hard to compare them intelligence wise - you have to get a feel for them yourself and what works for you. What I'm really excited for is the 27B model. 3.6 is still the king of local, and if 3.8 makes the same improvements it could really legitimately make local viable as a default. I'd love to run a perpetual agent on 3.8 that's locally driven.
- icedrift 2mo agoI'm still skeptical of the smaller models after the talent exodus a few months ago.
- jimbo808 2mo agoAt this point I feel like the only factor differentiating SOTA models now is who they’re propagandizing you on behalf of (not considering agentic tooling/state management, etc).
- Zambyte 2mo agoQwen 3.6 27b is already a viable default. I'm running it on a single 7900 XTX right now for Go development with pi. It's great.
- snapplebobapple 2mo agoWorks great with room to spare on my lenovo pgx too
- bitexploder 2mo agoI find 35B A3B viable as well, but your harness and runtime really matters to get tool calling and such dialed in. In fact, I would encourage you to experiment with it some as I find I get more reliable output from 35B A3B, though 27B is still generally smarter. A3B with a review cycle or two from 27B is great for me. One of the reasons is, with good specs and design, A3B is just so fast. It isn't as smart as the 27B model, but it is close enough it can usually figure it out with the right tools.
- londons_explore 2mo agoI just don't think you can combine speed, latency, price and intelligence into a single useful metric. Clearly the weighting of those things depends on the usecase
- dangoodmanUT 2mo agoI'm seeing opus 59.2, Qwen 58.4?
- zmmmmm 2mo agoThe fact that the Chinese models have caught up on benchmarks suggests to me that its likely we will start to transition now into much more of a brand war. It will be subjective qualities that drive our decisions more than measures of absolute intelligence. Already I am choosing models more because I like the personality or style of what they do than because I think they have the absolute highest chance of outputting the most technically correct answer to any given prompt. It will be very interesting to see how things evolve in this direction.
- mindwok 2mo agoFor me now it’s simply cost and speed. With GPT5.6 and Fable (and respective open models since then) we passed a threshold where intelligence is sufficient. Now I just need speed of iteration and good prices.
- gerdesj 2mo agoMy vague equivalent of the pelican riding a bicycle test (for a local model without internets) is to ask it: "Where is Yeovil"? I don't expect a totally accurate answer for obvious reasons but I do enjoy watching the accuracy improve. Qwen3.6-27B-FP8 currently espouses (see below), which is not too bad. The directions are a bit mad but the mileage is about right and there is a helicopter manufacturer here and a RNAS (navy not airforce) museum nearby at Yeovilton. Cosford is in Shropshire which is not a million miles away. I'm not sure what 盆地的 means but the river Yeo is correct ... OK ... "basin like" - again not bad, even if Chinese is not the first language here. The model understands that Yeovil is named after (or vice versa or at least is associated with) a river Yeovil is the current form of Gifle (Saxon) which I thought meant "bend in a river" but WP is currently saying "fork in a river". My source is a local museum. There is a fork but was it there 2000 odd years ago? My hydrology skills say ... possibly ---------------------------------------------------- Q: where is yeovil: Yeovil is a town in Somerset, in the South West of England. It is located roughly: 25 miles (40 km) south-west of Exeter 60 miles (100 km) west of Bristol 140 miles (225 km) west-south-west of London Yeovil is known for its historic market town center, RAF Museum Cosford (nearby), and as a significant industrial town, particularly during World War II for aircraft manufacturing (including the Wellington bomber). It sits in the盆地的 valley of the River Yeo.
- ngl999 2mo agoIt's censored and it'll spread certain kind of narrative all over the world.
- MrDrMcCoy 2mo agoIf you're using AI for narratives, you're using it wrong.
- imagetic 2mo agoIt's the first model I've used that makes me forget it isn't one of the big frontier players after the first prompt. So far I'm impressed.
- fjbarrett 2mo agoThere are great visual comparisons on this page.