4 ms·
Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy
by freakynit 2mo ago
Just tested through openrouter.. gave exactly same task.. the task was to scan existing repo, and generate a single docker-compose file to deploy behind a caddy server, where certain port ranges are already used, the service demands widlcard certificates to be provisioned from outside, and postgre needs to be built-in one...
Tested this model, and gpt-5.6-terra-high.
Results: this one had few issues. terra: none.
These results are consistent with my past observations with the latest flash version as well. What benchmarks say, vs what I've been observing are different.
They are good till the project is simple... not anymore.
- shimman 2mo agoI've always wondered if I was using containers wrong because none of them I've ever had to create were complicated. Maybe it's because I choose tools that make local development easy (Go + sqlite + various CLTs) or maybe it's because I never hard to interact with this on the professional side outside of making images for our projects (which still weren't complicated for the reasons above). LLMs make containers in a pretty workable format for me (still hand tweak the env variables for a sanity check). How exactly does it struggle here and why does postgres need to be built? Were the needs beyond what you get in a base image?
- freakynit 2mo agoThis was the repo: https://github.com/amalshaji/portr https://github.com/amalshaji/portr And this was my gh issue: https://github.com/amalshaji/portr/issues/308 https://github.com/amalshaji/portr/issues/308 And below was my prompt: """ give me single docker-compose file that i can run on my server to run current project... you can read README.md , and then, this relevant page: https://docs-custom-reverse-proxy.portr-docs.pages.dev/docs/server/custom-reverse-proxy https://docs-custom-reverse-proxy.portr-docs.pages.dev/docs/... ... this was the result of me raising github issue: https://github.com/amalshaji/portr/issues/308 https://github.com/amalshaji/portr/issues/308 ... you can use gh cli to fetch the details and comments... i already have a caddy server running on my vps... and i will create wildcard certificates myself using certbot.. the domain name will be helloportr.xyz ... also, ports up to 9019 are already taken... ask me if anymore info is needed... """ You can try yourself and let me know of what you got.
- shimman 2mo agoThis is definitely beyond my capabilities lol but wow portr is a neat project. Never heard of it before, only the paid services from tailscale/cloudflare.
- arch-choot 2mo agoI've been using DS4F+Pi with great results, but I think one thing that helps is at the end of my prompt I'll tell it how to verify it, e.g. "Make sure the compose file works by running it locally (use self-signed certs if required)". The argument could be made that "the model should be smart enough to figure it out" , and maybe DS4 isn't. But with just a bit of steering you can get the correct result for like 1/10th the cost, or even cheaper.
- ctx_wrangler 2mo ago[flagged]
- npn 2mo agowait for Deepseek Harness (yes it is the official name) release then try again. for your kind of task, harness tools matter.
- freakynit 2mo agoI used pi
- natrys 2mo agoFor me, flash 0731 was much better in omp/opencode than in Pi. Anyway, it might be so that they are rolling out deployment. There haven't been an official announcement post yet (this submission is a link to openrouter). Some people have been saying they are getting results worse than GLM-5.1, that's obviously broken.
- gkbrk 2mo agoIf the model cannot figure out simple and ubiquitous tools, how is it supposed to figure out complex problems? All of the good models basically work with any harness, including giving them a single "shell command" tool. They can just figure things out.
- hadlock 2mo agoWhen it comes to quality of outcome, since at least Feburary, the harness has almost equal, if not more weight than the model itself. It's no longer "which model is the best?" it's "which model + harness is the best?" I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models
- azinman2 2mo agoWhich works better for you?
- KronisLV 2mo ago> the harness has almost equal, if not more weight than the model itself This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness.
- scrlk 2mo agoWhat harness are you using? DS V4 is harness sensitive.
- freakynit 2mo agoPi
- lousken 2mo agoAre we testing the model or the harness? If benchmarks show certain numbers it should perform as such without it
- deleted 2mo ago[deleted]
- ApolloFortyNine 2mo agoYou can't even run a benchmark without a basic harness, of course the harness has some effect.
- 0xbadcafebee 2mo agoYep. It's a state machine; change the state, change the result. https://arxiv.org/html/2605.23950v1 https://arxiv.org/html/2605.23950v1 | https://arxiv.org/pdf/2505.15146 https://arxiv.org/pdf/2505.15146 | https://medium.com/@amontzamir/youre-praising-the-wrong-thing-c14028e2459e https://medium.com/@amontzamir/youre-praising-the-wrong-thin...
- derangedHorse 2mo agoTerra has not been able to do any of the technical tasks I've asked of it correctly. I'm surprised others get use out of it. Anything below Sol high tends to give me mostly unreliable results. I'm using codex as my main harness but maybe it performs better with a different one.
- freakynit 2mo agoDepends on project complexity. For one of my more complex projects, I exclusively use sol-high ... nothing below that works correctly. For this however, a comparatively much simpler task, tarra-high works fine.
- Foobar8568 2mo agoRight now, sol-xhigh is my favorite model. I feel that Opus 5 is dumber than 4.8. Fable is too expensive to do anything (limit of $50, started a prompt at $25, ended up at $75, is bullshit, but at least it's "free credits"). DeepSeek is okay for random API-based stuff, as it's cheap. Local open models running on a 5090 are hit or miss. I feel that most GGUFs/quants are awful...
- miohtama 2mo agoOpus 5 degrades to word salad. I wonder if it is because of watermarking.
- SwellJoe 2mo agoOpus 5 doesn't really even speak coherent English. I'm not sure what's going on, but it can't explain anything. It still does an excellent job with code and writing tests and code review and creating and completing a plan, and it seems to be able to understand English instructions, but it sure as hell can't explain what it did or how to use the code it wrote. That was true before they announced the watermarking, I'd already started to back off of using Opus as much because I like to understand what the model is doing and have it write documentation I can use to reproduce its results, but maybe watermarking was already in there unannounced.
- tripleee 2mo ago[flagged]
- MagicMoonlight 2mo ago[dead]
- apitman 2mo agoWait people use terra?
- ApolloFortyNine 2mo agoI use deepseek flash to do exactly this. Git repo (which I usually have it build from scratch) -> build docker image -> deploy to server with komodo/caddy-docker proxy. Works great, regularly one shot applications. I often make changes to the application after its deployed (to be fair, my prompts are usually quite laxidasical, just 'build x, use /deploy-to-komodo) but the deployment works great. I did make a skill, but if your doing anything repeatedly you should as well. Opencode, but any harness I'd think would work similar.
- yassa9 2mo agodid you test kimi k3 or qwen 3.8 max on the same task ? or plan to test them ? I respect those genuine users tests other than those benchmarks that models are trained and overfitted to them
- v3ss0n 2mo agoI do that kind of things all the time with Qwen 3.5 122B. It works well in one shot with Cline or Opencode. May be your harness problem?
- amelius 2mo agoI didn't understand your use case, so it could also be the way you write your prompt, I suppose ...
- celsoneto07 2mo agoI've been doing pretty heavy stuff with DeepSeek with a good degree of success. The thing is: I don't trust it to go fully autonomous. I check the steps, I steer it. For the pricing, it's worthy. Let's how the price increase is going to change my behavior.
- shunia_huang 2mo agoThe flash model will always use an outdated Treafik version that is not compatible with the newer docker engine, I tried to deploy some personal services with Traefik and everytime it uses this wrong version, and then fixes the version issue in the thinking chain. I was thinking to switch to Caddy but with your experience I'm gonna stay with Traefik and bare with the version issue...
- acchow 2mo agoI agree. The “frontier level” open weights models benchmark really well but fall behind in real world performance