5 ms·
I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
- luciana1u 1mo ago[flagged]
- shahariaa 1mo ago[dead]
- RobertasTa 1mo ago[flagged]
- binary132 1mo agoIn my experience xhigh very often produces better results on the first try to the extent that the less cogitation-enthused settings actually waste more in the long run.
- jchw 1mo agoI'd personally like to know more about what tools it used/wanted and the harness setup, because this sounds pretty cool. I have a dual Arc Pro B70 setup and currently get around 22 t/s which isn't great but isn't terrible either (it is at least less quantized.) I've seen GPT 5.6 Sol happily invoke objdump and even write jobs to run headlessly which Ghidra when trying to disassemble a binary.
- trollbridge 1mo agoMy M5 Pro gets around 12-15 (6 bit MTP), although I haven’t worked on optimising it at all yet. A nice thing about running locally is you can run an uncensored model and you don’t have to worry about TOS violations on your OpenAI account when you ask it to “reverse engineer this ancient router firmware and give me a licence key that will work on it”.
- medler 1mo agoQwen is very much censored. Just try asking it about Tiananmen or how to build a bomb. But it is nice that you can experiment with it locally without having to worry about your account getting nuked
- jchw 1mo agoYou are misunderstanding what they said, they are saying you can use uncensored variants of models like Qwen when running locally. There are quite a lot of people working to "uncensor" open weights releases. It seems to work although it would be nice if some third party was benchmarking the uncensored variants regularly to give us an idea of how well retained their skills are.
- trollbridge 1mo agoAbliteratuon sloghtly reduces the strength of the model - in my opinion it’s around the same jump as going from 5 bit to 4 bit.
- AdamConwayIE 1mo agoI added a line to address this, sorry it wasn't there before! It was Pi and only used Bash-based tools.
- jchw 1mo agoCool. I was thinking of running Qwen3.8 through Codex, but maybe it's time I take a look at Pi.
- saidinesh5 1mo agoLately I genuinely believe that the future will be large frontier models generating and updating inputs/skills for "good enough" local models to solve our daily problems. A lot of tasks which need a bit of intelligence don't really need that much compute. Just good enough documentation / skills, tool calling and a good enough local model. Not sure what exactly this means for all those data centers that are getting built... But exciting times.
- catlifeonmars 1mo agoWhat’s the fundamental difference between a frontier model and a local model anyway?
- cromka 1mo agoPrivacy!
- catlifeonmars 1mo agoYes, exactly my point! “Frontier” vs “local” isn’t a useful distinction . “proprietary vs open” is a much more useful distinction. Although I suspect people use “frontier” as shorthand for “way too large to run at home practically”.
- zarzavat 1mo agoThat's effectively what it means. You can run frontier models at home e.g. Kimi K3, but you'd need a large amount of money.
- Tepix 1mo agoWith AI being more useful with access to more of your data, I can't see myself using cloud AI models for purposes such as personal assistants. Perhaps with differential privacy or confidential compute... But ideally these models run locally.
- 1mo ago
- VulgarExigency 1mo ago> The first attempt at recovering the key was wrong in a very specific way; it produced a working key and the signature check passed, but a hash the binary computes as an integrity check didn't match. In my experience, most models would have called it done and left it at that, but Qwen 3.8 27B didn't do that. Instead, it highlighted the mismatch, went back to the drawing board, and kept going until the value matched byte for byte. This seems to be a pattern in the more recently released models that I think accounts for an increase in the quality of their work. They are very persistent in verifying that their work is actually correct, so even if they're not as "smart" as bigger models that get it right the first time, they have the ability to follow through to ensure that the work is actually done.
- braiamp 1mo agoWell, it seems that Linus doesn't use those: > And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work. > I'd like to call it my tireless helper, but the AI several times stated flat out that this was impossible and unsolvable and that we should just write a report about it. > I suspect those things have been trained by people who may not be quite as stubborn as I am. https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=818bebeb63dd6bf5f4e07e145f6cdbace520a34c https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
- tonyarkles 1mo agoBoth things can be true. I’ve noticed both the same thing the parent posted and what Linus posted and my vibe on the split (I haven’t kept detailed notes) is that on greenfield code they tend to maybe over-verify and on brownfield code or data analysis they sometimes give up too early or… I’m not sure, need a bit of encouragement to keep pulling at threads. On the data analysis side, something specific I’ve noticed is an (understandable) bias towards computing numerical statistics, which they do very well and reading the post-analysis report has significantly improved my own “statistical thinking” approach overall. Numerical statistics are cool and understandably what a text-based LLM is going to want to work with, but asking the model to produce time-domain and frequency-domain plots of, say, specific events has multiple times resulted in “trying to plot this out has shown the opposite of what I concluded numerically… recalculating…” There’s still a pretty significant review and critically assess step for me, especially since the actions I take as a result of the analysis are pretty expensive, especially if they steer the next data collection run in a useless or harmful direction.
- exceptione 1mo agoLocal models would be even better if they did not ship with all the refusal shenanigans built-in. You can safely bet organized crime has access to the best models without these hoops, which makes the case that the average user (=non-criminal) should have access too. As I understood from an ex-Anthropic employee, some orgs got access to Mythos based on their high enough spending level, not on other grounds. Either we are in command over the software, or the corp is in command over us via the software. I can on a theoretical level understand the concerns, but either we ban all LLMs or we have a level playing field for everybody. Let's not forget: defense and offense are different sides of the same coin in software. I guess this wouldn't apply to bio weapons, but I am not in the know about that.
- binary132 1mo agoEhh, it’s at least given as the excuse for gain-of-function bioweapon research
- datsci_est_2015 1mo agoDigression, but this is the real Great Filter imo, not AI. I think technology advances to a point where it only takes one or two bad actors to type the right prompt to get a recipe for civilization-destroying bioweapons before you get anywhere near true AGI or anything relevant to the Kardashev scale. Biology is fragile. But not that that’s a good justification for hamstrung models. I think it’s just the inevitable endgame and it’s more sad than scary
- jeremyjh 1mo agoI have the same concern. If it becomes possible to engineer Captain Trips with a budget in the low 8 digits it won’t really matter what else happens.
- xyzzy123 1mo agoI don't fully understand the instinct to regulate local models for this? It seems like the wrong place to address the problem. You can download Ebola sequences right now if you want to. That's not the same as having an isolate. The difference is a lot of messy reality. This kind of work is not generally "one shot" (Claude make me a supervirus, make no mistakes), it requires lab space, iteration, and specific resources. It has a footprint. Wouldn't it make more sense to monitor / regulate facilities where you can sequence or request assembly of DNA, RNA, restrict and monitor the supply of key reagents and so on?
- topper00_raptor 1mo agoWhy does the screenshot on your pi terminal shows opus-4.6-medium from your claude subscription ? Instead of Qwen ?
- MarkWayneNewton 1mo ago[flagged]
- AdamConwayIE 1mo agoNope, it wasn't.
- deleted 1mo ago[deleted]
- AdamConwayIE 1mo agoAh, my bad! This image came from our backend, used for an unrelated article. I selected it by mistake rather than inserting the actual image that I'd uploaded. I'm updating it, thanks for the heads up! For what it's worth, that image couldn't have been related. The other screenshots all showed thinking traces, and Claude doesn't share those.
- pi-victor 1mo agoi'm not good with paper work, in fact, i'm horrible with anything that's paperwork related. for the past few days, i ran this model on my rtx 4090 + rtx 3070 and told it to check all the bills, invoices, contracts for me and my small company. i used pi with llama and the pi-llama plugin. oh, boy - i hooked it to my email, told it to download all of the invoices and bills i had for both me and my company and organize them by company/date/ and then merge them with the ones i have locally. it did ocr, wrote scripts, organized everything neatly. i am now the most organized i've ever been in my life. Next: RAG on all the documents and bills i have. if you connect staan-search (there is a pi plugin for that) and ctx7 to this it almost does miracles. the downside is i have to sit next to my noisy threadripper as the magic happens and pay for the electricity, but that's about it, i'll gladly do that. and as i finished this paragraph, it also finished organizing all my personal documents on my san. i don't use the expression "game changer" easily, but it's hard to resist in this case. out of all the models i've used locally qwen3.8:27b blows everything out of the water. my setup # Logical CUDA0 = RTX 4090, logical CUDA1 = RTX 3070 export CUDA_VISIBLE_DEVICES=0,1 cd ~/projects/misc/llama.cpp/ exec ./build/bin/llama-server -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M --mmproj /xx/xx/xx/xx/xx/mmproj-Qwen3.8-27B-Q8_0.gguf --host 0.0.0.0 --port 8080 --jinja --parallel 1 --split-mode layer --tensor-split 6,1 --fit on -fa on -c 98304 -ctk q8_0 -ctv q8_0 --image-min-tokens 1024 i load more on the 4090 because it's faster. usually the temp stays around 65 for both. utilization for 4090: 70-90% 3070: 30-50%. I get around 30-40 tk/s. if i offload more to the 4090 the tk/s goes up, but i stress the card too much and that thing now is worth its weight in gold. note: the pi-llama plugin needs a patch for pi to send the model vision capabilities, seems it doesn't work out of the box.
- deleted 1mo ago[deleted]
- throwa356262 1mo agoThanks, now I too want a Lenovo Thinkstation PGX... I think it will be fairly easy to remove refusals from open models. Feels like a lost battle, so why does Alibaba even bother?
- __alexander 1mo agoIn my benchmark Deepseek-v4-flash did much better than Qwen 3.8 27B at reverse engineering. https://alexander-hanel.github.io/StressingLLMs/ https://alexander-hanel.github.io/StressingLLMs/
- petu 1mo ago> This project evaluates local language models running on a single NVIDIA DGX Spark. "Did much better" is a bit misleading w/o that context and 1 hour time limit -- your benchmark design heavily favors V4 Flash. From results on your page V4 Flash processed 1-1.5M tokens an hour, while Q3.8 27B was failed before even reaching 200K tokens. By the way, how are you running V4 Flash on single Spark? Was it quantized?
- __alexander 1mo agoA single model was loaded at a time and each model was given a 90 minute timeout. This is how I’m running it (actually an older version because they deleted the repo and replaced it) https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-One-DGX-Spark https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-One-DGX-Spark
- djoldman 1mo ago> I gave it the hardest real task that fits on one machine: reverse-engineering a commercial app's license check... Respectfully, tasks that allow for explicit straightforward true/false or done/not-done tests are not the "hardest real task[s]." In fact, those are the ones that see the most gains from AI-assisted coding. Testable tasks are where the largest opportunity is.
- tempest_ 1mo agoWhich is exactly why we saw 1000s of ' "I" rewrote <mature software> in rust' posts last year when agentic coding really took off. Agents (even ones powered by small models) do reasonably well when provided an oracle to work against.
- cyanydeez 1mo agoI've included docs and tests as part of my vibe coding endevours. It doesn't matter if either is litterally correct, but they create guardrails for future context to prevent regresssions and blind avenues, etc. It's fairly successful but hits the time constrains and reduces the "value" of getting a local model to develop software. It's still a bump in productivity.
- hghnncrh 1mo agohow do incorrect tests or docs help create correct guardrails? if your tests and docs are possibly incorrect, and you're not writing the code.. how do you know if it even works? for extremely simple software you can just use it but for anything with access to disk or the network or with user options... you sound psychotic. actually. so nevermind, LLM psychosis is extremely common on this website, that's def all that's happening here
- cyanydeez 1mo agohuh. Oh, I understand, you don't read LLM output. In local coding, the screen scrolls enough to actually read it. But that's fine. enjoy your misunderstanding.
- samuel 1mo agoI have this idea of using an obliterated version of this model for cyber work(or even this one, seeing that its guardrails aren't that strong) in a harness with the ability to spawn SOTA level subagents, faster and more capable. The rationale is that the manager model sees the big picture and knows that the task is "unethical" while sota models are just given very isolated technical tasks that don't trigger any refusals. Has anyone tried this? I would love to know about previous attempts of this approach.
- andai 1mo agoYeah, the Chinese government used the same method last year to hack the US government using Claude Code. Making each piece of work small enough to be plausible. Compartmentalization. (Also saying "nah it's cool I have permission", heh) https://www.anthropic.com/news/disrupting-AI-espionage https://www.anthropic.com/news/disrupting-AI-espionage
- ianmarcinkowski 1mo agoI used it with opencode to build an admin UI for a React slideshow presentation app I use to do presentations. It worked pretty well on a 64GB Mac M3 Pro and took 1-2 hours.
- EGreg 1mo agoCan you give it a task like “Prove or disprove the Riemann hypotesis, keep going until you’ve done it” and see how long it takes? :-)
- mdp2021 1mo ago> As it turns out, probably unsurprisingly, Qwen recognizes common jailbreak attempts, and one of the first things it told me was that it wasn't going to fall for the jailbreak prompt Now also see latest submission, https://news.ycombinator.com/item?id=49409073 https://news.ycombinator.com/item?id=49409073 : # I spent $266 and four AI models to own my tablet. GLM-5.3 finished it in a day > Quick context: the tablet is a 2021 Fire HD 10 that ran my Home Assistant dashboard and kept powering itself off: the logs showed Amazon's own software issuing the shutdowns, and the only permanent fix was root, which has never existed publicly for this model. Anthropic's and OpenAI's cyber safeguards wouldn't touch the project Why should Anthropic and OpenAI thrive: they do not work on real problems.
- giancarlostoro 1mo agoHow can I use AI to do real security audits anymore if they don’t trust people in an enterprise plan? Its useless.
- verdverm 1mo agoTrust is a two way street, why do you trust them (Ant) if they do not trust you? Have they done enough shady things yet to break it? Are their models really that far ahead it's worth it?
- giancarlostoro 1mo agoYes, if you use Fable for anything (which is their best security model) they send all your IP and prompts to be reviewed.
- rustcleaner 1mo agoThe answer is to NEVER SUBSCRIBE, let the datacenters go the way of dark fiber after the '90s telecom bubble collapse. Maybe also a nice fat 90% corporate income tax on rentier-like businesses (subscriptions based: like SaaS, AI resellers, non-perpetual licensers, cloud storage and compute, etc). Force businesses with tax policy to only operate in a sell once + works forever, business model. Those who wish to rentier will need to spend the income on hiring more employees or put it into R&D, but the tax slides in after costs but before dividends or stock buybacks. :^)
- sm_ts 1mo ago[dead]
- jnwatson 1mo agoI can't get Qwen 3.8 27B to do a simple code review on a fairly basic Python file. With thinking on it just ruminates forever and with thinking off it gives obviously bad borderline hallucinating advice. Edit: I tried again with the 2.4T model and it still ruminates to death, but with thinking turned off, it generated genuinely useful advice. Edit2: adding --reasoning-budget 8000 --reasoning-budget-message "Reasoning budget exhausted; give the final answer now." --reasoning-effort low" to the llama.cpp executable parameters produces pretty good output.
- Refefer 1mo agoOne of the big learnings from 3.8 27b is adding reasoning budget really hurts the model. you need to let it spin for as many thinking tokens as it wants to to get it out. Another big takeaway is reasoning effort set to low doesn't save you tokens: low is pretty uncertain about things so it ends up thinking more (you can find some tests from folks on youtube). The final question, as always, is what quant are you running it at? KLD matters _a lot_ when it comes to its performance and it especially manifests with MTP/DFlash acceptance rate which makes those long thinking traces take a long time.
- jnwatson 1mo agoIt literally ran forever without a reasoning budget. I tried even the 2T model and cut it off after a half hour. This is to review a few hundred-line source file. It was consistent behavior from 2T to vanilla 27B to my ablated distilled version.
- nialv7 1mo agohow are you running the 2.4t model locally if you don't mind me asking
- jnwatson 1mo agoOh I didn't; only a distilled 27B model ran locally. This was just to figure out if distilling caused the problem.
- geye1234 1mo agoI'm far from being an engineer, but I can code a bit and have an engineering-adjacent role, and 3.8 27B "seems" -- purely subjectively -- miles ahead of 3.6 for the medium-difficulty tasks I give it. In particular, it's only started looping once in the 2 weeks or so I've had it. 3.6 did so every day. I normally run with thinking low but it's still miles ahead. I had been annoyed at not being able to run 0731 locally, but now I'm not sure I need it. I think I could leave 3.8 running overnight without waking up to find my office sweltering at 80F and seeing eternal loops on my screen.
- freepiai 1mo ago[dead]
- bluecalm 1mo agoAs someone who was selling Windows desktop app for 10 years and made nice money out of it I have mixed feelings. On one hand it was always a losing fight against determined hackers on the other the tools weren't widely available so the problem wasn't as widespread. We lost quite a bit to piracy but could still make a decent business. With widely available LLMs I think that business model is truly dead though. Not only hacks/cracks but also any kind of smart idea you may have will quickly be reversed engineered from your binary. If you never lost money to piracy you may think that "those people are not your potential customers anyway". This is not true because people will crack your software and then resell it - often pretending to be legit resellers operating under your brand. To add insult to injury they will send their customers to your support as well. If I ever come out with something smart again there is no way I am shipping it as executable. SaaS it is for better or worse.
- Frannky 1mo agoTry glm 5.3, is very helpful.
- rasitakyol 1mo ago[dead]
- HUMAN-X 1mo ago[flagged]
- HUMAN-X 1mo ago[flagged]