8 ms·
DeepSeek Harness developer preview
https://github.com/deepseek-ai/deepseek-harness https://github.com/deepseek-ai/deepseek-harness
https://deepseek-harness.github.io/deepseek-harness/en/guide/quickstart https://deepseek-harness.github.io/deepseek-harness/en/guide...
- yunbiao 2mo ago[flagged]
- bobleer 2mo ago[flagged]
- 217 2mo agoif it's not better than omp im not trying it
- syntaxing 2mo agoIs there a reason why so many of these agent harness are written in node.js?
- m_ke 2mo ago1. it's built for async 2. runs everywhere 3. interpreted, making it fast to iterate on 4. decent performance 5. most popular language, llms are decent at writing it
- altmanaltman 2mo agoaren't 3 and 4 a tradeoff though? Yes you have 3 but "decent performance" cannot be an extaled value as compared to "runs everywhere". If its used as a counter balance to 3 then it shouldn't be its own unique point basically saying 4 is true despite 3 in this case.
- m_ke 2mo agowhen you're waiting for network or LLM inference the raw performance doesn't matter at all
- eglintondust 2mo agoThis line of thinking I feel like assumes it's the only program running on your computer. Using less of my CPU and memory means my computer can do more things in parallel, or even run more instances of the harness. My laptop is sweating when I got 5+ claude code sessions running.
- jaapz 2mo agoIs it actually claude using those CPU cycles though, or the agent running test suites and what not? Honestly I would not be surprised when it actually IS claude using those resources... It is very clearly vibed
- eglintondust 2mo agoI haven't monitored CPU usage so closely, but seems to get heavy with basic tool calls and editing. Memory usage definitely is out of hand, have an idle session right now eating 500MB
- altmanaltman 2mo agoyeah fair enough, my entire point is not about the application itself but the contradiction on using superlative terms for all points but a compromising/normal term for one. Like if performance is not revelant why include it in the list of benefits.
- Zambyte 2mo agoI'm not sure why specifically Javascript instead of something like Python or other options, but using an interpreted environment minimizes the friction for implementing extension systems, which are an important feature in AI harnesses.
- ubercore 2mo ago`uv` helps but it's new, and I don't think it has the same mindshare yet on "I just globally want to install this thing that needs an interpreter/runtime", so Python probably just doesn't come first to mind.
- svachalek 2mo agoI'd put it on this. In my experience Python is fine for scripting your own machine but an obnoxious platform to distribute code on. It's very fragile to version changes, in both directions; I don't know how many things I've seen that only run on 3.10, not 3.9 or 3.11. Its packaging system is global by default which only compounds this because everything needs a specific version but they're all dumped in the same place. And it tends to have a lot of native code as dependencies, leading to all the issues of needing to either have the right build environment or a runtime environment that's already been built for.
- hedora 2mo agoPython basically requires containers unless you are OK with it bit-rotting every six months or so. At least, this used to be the case for trivial python, and recently was the case for stuff that uses cuda. I stopped paying attention the third time they redefined matrix arithmetic semantics. That happened to be around the 100th time I was sent a script and it only ran on the author’s machine. Maybe they will fix it some day. When they do, I will not believe it. In contrast, TS has a much nicer type system and better async support. It runs well on web, mobile, desktop and server. Yes, sometimes you have to ship node.js or a whole web browser, but the tooling for that is slightly less insane than the analogous tooling for python. Its language interoperability story is slightly nicer too (invoke native code, or use wasm). It’s UI story is much, much better since it reuses all the web stuff. Pip practically invented the supply chain attack; npm perfected it. That’s probably a draw. Of course, if you care about performance, then other choices make more sense. If you’re training a model then python probably still wins, but very few customers have a $1M+ machine.
- sarjann 2mo agoMight be easier to do cross platform.
- nimsarajay 2mo agoI gotta same problem.
- game_the0ry 2mo agonpm as a distribution tool works well and typescript has types. Any reason why it should not be written in nodejs?
- LeBit 2mo agoI always thought nodejs was a weird choice for CLI tools. For web stuff, sure. But for CLI, it never made sense to me. Especially when Python and Go exist.
- game_the0ry 2mo ago> But for CLI, it never made sense to me. Especially when Python and Go exist. But why? Not saying node is better, just want to know where you are coming from for my own knowledge. Bc I would have picked typescript + node too. It has types (where python just has type hints) and a lot of developers know it already (where go is more niche).
- LeBit 2mo agoI guess the reasons are that I feel (no data to back this up) that having a Python runtime available is much more probable than having a nodejs runtime around. I was going to say you cannot easily distribute a nodejs based CLI app, but that’s of course not true. devcontainer-cli is a nodejs app and so are many of the coding agent harnesses. Yeah, thanks for pushing back. I guess my view was irrational.
- pohl 2mo agoProbably for the ease of coding extensions — which strikes me as outdated thinking: if it’s open source and you’re outsourcing the coding to LLMs, why not use a compiled, safe language? There’s an interesting counter example for DeepSeek called CodeWhale, though: https://github.com/Hmbown/CodeWhale https://github.com/Hmbown/CodeWhale
- Wowfunhappy 2mo agoAlso isn't OpenAI's Codex written in Rust?
- edgyquant 2mo agoI don’t think so it’s an npm package iirc
- ceehex 2mo agoyou can install binaries with npm too, not just limited to js
- SwellJoe 2mo agoReasonix (which has been seemingly the most recommended harness for using DeepSeek, as it is designed around maximizing caching in DeepSeek) is now a Go app, but still installed via npm. Which feels ugly, but I guess everyone has npm already, and it handles binaries, so I guess it's a reasonable choice.
- hocuspocus 2mo agoCodex, Kiro, Grok Build. Pi has a clone in Rust too.
- pohl 2mo agoI don’t think so. The ChatGPT app was, which is the “Classic” app now. The Codex app that they’re carrying forward is an Electron app and if you forget to quit it before you walk away it’ll make even your M5 Max unresponsive eventually. Sad days.
- jesse_dot_id 2mo agoTypeScript is great and its ecosystem is easy to work within.
- 2afTq 2mo agoSkill issue. Their models don't work for serious programming so everyone just copies Electron apps from each other.
- Wowfunhappy 2mo agoBecause: 1. The first significant agentic harness was made by Anthropic. 2. One of the most senior developers of client-side software at Anthropic is Felix Rieseberg, one of the original creators of Electron. [1] 3. After Claude Code blew up, everyone else copied Anthropic. --- 1: https://daringfireball.net/2026/07/claudes_criminally_bad_mac_app_is_an_inside_job https://daringfireball.net/2026/07/claudes_criminally_bad_ma...
- RussianCow 2mo agoI think it also helps that it's basically the default platform for any software that AI writes, unless you tell it otherwise. And JavaScript is one of the most widely used and well known languages in the world, so there's that, too.
- tosh 2mo agocodex is written in rust fwiw smol has implementations in Go, Python, Clojure, PHP https://github.com/smol-env/smol https://github.com/smol-env/smol out of the box an agent only needs to be able to do http requests and call tools (which might again be just http requests or shelling out) there is no inherent reason for why an agent has to be in JavaScript or Typescript but they are popular languages and come with runtimes and libraries for http requests, steaming, TUI (terminal ui) and so on which can help
- sroerick 2mo agoI'm interested in this but why those four separate languages Edit: okay I read the code, it's actually four separate implementations
- tosh 2mo agoYes it's separate implementations of the same minimal idea I'm currently working on more 'feature-full' but still minimal variants e.g. a python variant with automatic compaction + truncation of sh output https://x.com/__tosh/status/2087606344035479632 https://x.com/__tosh/status/2087606344035479632 i also got quite a lot of requests to provide the code in non-golfed form to make the implementation more approachable and idiomatic in each language (will do!)
- scotty79 2mo agoNothing better for UI than React. Electron/tauri or node kinda falls from it.
- Shorel 2mo agoBecause that hammer is their only tool!
- root_axis 2mo agoTypeScript's type system is extremely expressive while still allowing you to retain the flexibility of a scripting language. v8 and JSC also have decades of performance tuning across basically every consumer device.
- nylonstrung 2mo agoI'm still waiting to find one written in Rust that I really love. I don't think js/ts makes sense for terminal based applications
- ohyoutravel 2mo ago[flagged]
- cachius 2mo agoHaS aNyOnE sEeN aLtErNaTiNg CaSe CoNvErTeRs https://en.toolpage.org/tool/alternatingcase https://en.toolpage.org/tool/alternatingcase https://www.textformatting.com/case-converter/alternating-case https://www.textformatting.com/case-converter/alternating-ca... Sadly no backwards direction
- m00dy 2mo agoit looks like we're leaving md files and instead use cordis plugins ?
- pyrophane 2mo agoI'm curious what peolle are finding with first party vs 3red party harnesses for coding. Do the first party harnesses really have an advantage when paired with the maker's model?
- dsrtslnd23 2mo agoI hear that often but to me it does not feel like it. I built my own framework around pi.dev harness and run all kind of different LLMs with it. Sometimes also use the vendor harnesses and they don't feel better adapted.
- deleted 2mo ago[deleted]
- softwaredoug 2mo agoI use OpenCode and I like knowing the direct token spend for doing tasks. A healthy repo can get a lot done with Luna + fresh context. Then I can spend $1-$2 a day when I'm doing development, and costwise honestly it beats a $200 / month plan. I also just do a bit of hand-coding to guide the agent still. I worry the $200 / month plans are loss-leaders encouraging you to maximize token usage to churn out slop, rather than thoughtfully use coding agents in a way that still engages your brain, and produces good software.
- marstall 2mo agoi've been using Cascade (a third party harness) since the 3 week period in 2023 when it was hot. I think it's called something else now. Devin? Things got confusing there for a second and I stopped paying attention. Anyway very happy with it, I use it as a plugin to RubyMine and Webstorm. One of the primary advantages is being able to choose your model - and it often has free deals for newer models that are running promotions. Whenever I switch to Claude Code it seems clunky. Would rather use Claude with Cascade.
- nylonstrung 2mo agoI feel very strongly the first party harnesses don't make sense when essentially every month the Pareto frontier changes as new models get released I want to use the same consistent working surface across models in the same way I want to use the same text editor across all different languages
- aratahikaru5 2mo agoThe landing page provides more context than GitHub: https://deepseek.com/harness/en/ https://deepseek.com/harness/en/ The documentation, built from repo, is available here: https://deepseek-harness.github.io/deepseek-harness/en/guide/quickstart https://deepseek-harness.github.io/deepseek-harness/en/guide... (I find the development and reference sections easier to read and navigate)
- m00dy 2mo agoyeah, a rare thing.
- hmokiguess 2mo agothis should be the link!
- 3abiton 2mo agoI think the github readme repo is more for people who heard about it, but you're absolutely right, it is lacking in context and explanation.
- Gourabdg 2mo agoMaybe little weird to ask this but people knows what he is trying to talk about still "needs more context" !!
- catigula 2mo ago[flagged]
- m00dy 2mo agothis is a very dangerous question.
- catigula 2mo agoWhy?
- m00dy 2mo agobecause this time it looks pretty much original ?
- hedora 2mo agoYou’re implying the open weight model providers are behind the US companies, so they cannot do anything right. Instead, they currently own the entire Pareto frontier — they have the lowest cost model (in terms of inference and training) at every commercially-available level of output quality. We saw the same attitude from Silicon Graphics, Sun, etc vs Linux and Windows during the 1990s. It led to those companies’ ruin. Concretely, I remember lots of arguments that the Linux kernel team would stall out once they implemented posix, since that was the end of the “copy for the sake of compatibility” runway. While making such claims, none of the Unix vendors produced anything vaguely price-competitive with whitebox PCs (they were slightly better for niche workloads at 10x the cost, with crippling guardrails, er, license gated features). Those vendors even tried getting the US government to intervene with procurement regulations, etc. Anyone that was paying attention during the dotcom era should know how the current bubble ends.
- yipinwong 2mo agoIn the court of law, the plaintiff has the burden of proof. You need to provide the proof instead of accusations. What if DeepSeek never copied anything from anyone? They cannot prove something they haven't done. Same here, you gotta provide the proof or at least trace of where DS might have done so. --- Also in this field, nothing is original. Everything builds on another's ideas (unless the idea is copyrighted. Paid for it? then ok, stolen? no)
- rco8786 2mo agoBut like, what is it? Odd that this reached #1 on HN. The README is pretty bare outside of installation instructions and a link to "Cordis", which is "A Meta-Framework of Spatiotemporal Composability." and "under active development. The API is not yet stable and may change without notice.".
- francislavoie 2mo agoA "harness" is basically what you call Claude Code and such, i.e. a TUI to run the agent.
- deleted 2mo ago[deleted]
- a3w 2mo agoTUI? Aren't VS Code, Claude Code, Hermes Agent, Goose or Letta harnesses, but with UI, too?
- edgyquant 2mo agoVSCode at least is a GUI
- kaicianflone 2mo agoWhich it’s kind of strange VSCode ghcp lacks basic attributes like context % used compared to some TUIs where it’s default.
- u8080 2mo agoThere is a round icon at the right bottom, where white arc is how much context used - hover for extra info.
- kaicianflone 2mo ago
- vinhnx 2mo agoThe Cordis plugin architecture is interesting https://github.com/cordiverse/paper https://github.com/cordiverse/paper
- OutOfHere 2mo agoAs I understood, Cordis is for architecting functionality as plugins that can be hot-loaded and hot-unloaded (without having to restart the parent app such as VSCode). Cordis looks to be a second-layer extension system within the parent system, e.g. VSCode. I understand that Cordis is not tied to VSCode. Using memory to track inverses does not scale.
- vinhnx 2mo agoI think the paper is really worth reading for anyone working in software. As far as understand it is a software architecture paradigm where everything is a "plugin", and as plugin, I can plug-it-in and plug-it-out, if I understand it correctly. They called it "revertible effects", where software can strip out live code without a restart or system reboot. They give the example of VS Code, which requires system restarts whenever an extension or plugin needs to be updated. To help myself understand it, I created a quick video, using NotebookLM: https://www.youtube.com/shorts/LtR7DRlZJ0M https://www.youtube.com/shorts/LtR7DRlZJ0M.
- deleted 2mo ago[deleted]
- hmokiguess 2mo agoTangential but, are there benchmarks out there on how languages affect latent spaces and performance of these models? This other day I was looking at that “caveman” skill, and was shocked to see it evolved to become a company, and, in one of its modes, the highest form of compression seems to be “Wenyan” which is Classical Chinese. Should I get started on learning Chinese?
- wongarsu 2mo agoThere are lots of papers on the topic. I think the best summary is "it's complicated". Typically models perform slightly better in English, typically best in either professional English or very rude English. Though this varies by model, not all react well to rude English, and I wouldn't be surprised if Chinese was on the rise Also, "less tokens" is not always straight forward. I doubt it's a coincidence that the cavemen skill (or now proxy, I guess) has lots of numbers, but not a single benchmark on model performance or actual per-task token savings For example one paper I remember found that without CoT, just stating your prompt twice increases model performance. With CoT, the same function is served by the CoT restating the important parts of your question. Something about which tokens can affect which other tokens in attention implementations
- hmokiguess 2mo agoVery interesting, can you link some of the papers if you don't mind? I'm curious about this space. I'm finding more and more there seem to be sort of niche prompting skills that are important to be aware of
- 1899-12-30 2mo agoIt's interesting to note that the newer LLMs like deepseek v4 or kimi k3 basically use caveman mode natively for their thinking traces. Lot word dropping when thinking.
- throwa356262 2mo ago[dead]
- fkysly 2mo ago[flagged]
- lxdlam 2mo agoI have read the underlying paper, and found it may be useful, but not that useful. For those who want to know what it achieves: it adds hot-reload and dynamic enable/dispose capabilities to a plugin system, like the one in Pi agents, though they push the boundaries further, to the UI components and so on. For those who want to know what it does: if you have some PLT knowledge, ask your agent to explain the algebra to you better; for those who aren't familiar, the framework requires each plugin to provide how it initializes and how it destructs (like C++'s RAII, Rust's Drop trait and so on), and the runtime will then properly handle the lifecycle events and the common pitfalls. In addition, it provides a clean way to declare the dependencies between plugins, and the runtime will also properly process the lifecycle changes on a broader plane. I think it's worth reading if you are not familiar with OSGi, iPOJO, React's useEffect and so on (which the paper itself mentions); for others, a skim is enough: it does point out the gotchas for some common problems, but the algebra may not help you further.
- esafak 2mo agoFor all the high-powered theory it looked just like every other harness! I was expecting more. If anybody has tried it, does it let you preview components in any frontend framework with perfect fidelity? That would be a big win.
- slopinthebag 2mo agoUh, it’s big idea is a destructor? This is considered significant in 2026 and the era of vibe coding?
- Game_Ender 2mo agoDon’t sell it short, it’s big idea is also to support dependency injection style explicit linkage between dynamically added components.
- grommz 2mo agoThe paper mentions agent harness self improvement as one of the use cases. I don't know what's the advantage vs. iterating over a monolithic harness.
- WhereIsTheTruth 2mo agoIn the age of LLMs, if your new hires are pushing npm slop, with all the cargo culting and security pwn issues it brings, your hiring process has failed you oof
- cbg0 2mo agoWhat if the old hires are doing it?
- WhereIsTheTruth 2mo agoNew hires: https://x.com/victor207755822/status/2057064415300841626 https://x.com/victor207755822/status/2057064415300841626
- Gecko4072 2mo agoBad timing: https://xcancel.com/deepseek_ai/status/2087864589895798968 https://xcancel.com/deepseek_ai/status/2087864589895798968
- weird-eye-issue 2mo agoWhy is this bad timing?
- Fuzzwah 2mo ago[dead]
- brookritz 2mo ago[dead]
- jbellis 2mo agoAnd that's it, that's the last lab releasing models worth coding with that didn't have a first party harness that its models are trained to use.
- flaburgan 2mo agoIs there a comparison of harness somewhere? Like, the same prompt to the same model, but with different harnesses, and comparing the quality of the results. I am trying to run as much as possible only on free software, so I always only used Zed plugged with anthropic models, but I am wondering what is the quality of Zed harness compared to the one of claude code or pi or others... I would love some feedback.
- bobleer 2mo ago[flagged]
- schafberg 2mo agoI want to have the same thing, but tbh it's too complicated with so many configuratios and plugins. I doubt if any comparison of harness make sense now and can be applied in real coding works.
- alienbaby 2mo agoCan't that be sidestepped by comparing harnesses 'out of the box'
- sora2c0314 2mo ago[dead]
- bmurphy1976 2mo agoTracing what it actually did. Who would have thought that's a good idea, instead of trying to obfuscate everything.
- yipinwong 2mo agoGood idea, ugly landing page
- deleted 2mo ago[deleted]
- invaliduser 2mo ago«It uses an architecture where everything is a plugin» Ok, that's enough for me. I have developped over the year a plugin fatigue. Every product relying on "community plugins" for their features implies it works fine the 6 first months, then it's a nightmare of incompatible, deprecated, incompatible plugins, with no consistency and no governance. I understand how attractive it can be to companies to think, hey, let's make a very small product and rely on other people to make features, and I hope it works, but I'm personally staying away from that.
- curreylabs 2mo agoPlugins are the right solution for software that needs to strictly isolate a stable core domain from an unpredictable long-tail of niche integrations.
- NBJack 2mo agoThat can work great when the core plugin interface offered is actually stable.
- wltr 2mo agoSounds like a Linux architecture, innit?
- bdcravens 2mo agoMany of the libraries and executables in Linux are cross-compilable with other ecosystems. It's the difference between opening the door to an existing ecosystem and birthing one.
- c-hendricks 2mo agoyep, it's a handy architecture. Tho I don't know many people who prefer to run without coreutils.
- orbital-decay 2mo agoEverything about harness design is still experimental and janky. Throw everything in a pit and let the fittest survive. Large opinionated software is unlikely to survive and more likely to give you a migration fatigue
- 0xbadcafebee 2mo ago> It uses an architecture where everything is a plugin Did they discover Unix pipes?
- mring33621 2mo agoI just installed DeepSeek Harness with the latest Bun version and am using it with a local 9B, speculative decoding Qwen 3.x variant, running in llama.cpp and it works GREAT for small python projects, so far. It was very easy to connect the harness to the local model and it seems to run quite fast, compared to other harnesses that I have tried.
- mring33621 2mo agoSadly, it doesn't ship with support for "dsh --profile acp"
- tosh 2mo agooften new harnesses are based on pi this looks like a genuinely new one
- huqedato 2mo agoPlease somebody explain what is this good for. Is it a similar tool with Claude Code or Antigravity ?
- igravious 2mo agoDeepSeek Harness -- your common-or-garden coding harness https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices https://venturebeat.com/technology/deepseek-harness-launches...
- scotty79 2mo agoAlso similar to ZCode, Codex (now called ChatGPT) and many private projects that people make for themselves and publish in few days. At least superficially similar. Deepseek Harness supposedly has nice plugin system which others do mostly lack.
- tianyicui 2mo agoHi I'm one of the authors of DeepSeek Harness. It's just an early developer preview version we're presenting in MIT license currently. Expect lots of rough edges and compatibility-breaking changes. Any feedback and suggestions are more than welcome!
- vatsachak 2mo agoDo you use deepseek models to improve deepseek training and inference?
- flakiness 2mo agoTell me more about the ideas behind Cordis the plugin system. The paper is a bit too mathy to consume and I think it deserves a more accessible post or something.
- culi 2mo agoYou're making demands (like you would to an llm) instead of asking questions (like you would to a human). The GP didn't even offer to answer questions
- derekdahmer 2mo agoHe explicitly asked for feedback
- ziofill 2mo agoExactly. “Tell me more about X” is an open ended question, not feedback.
- KyleJune 2mo agoYou're absolutely right — "tell me more about X" is phrased as an imperative, not a question. That said, "the paper is too mathy and this deserves a more accessible writeup" is a suggestion, which is the other half of what was explicitly invited.
- deleted 2mo ago[deleted]
- cedws 2mo agoGuh, why TypeScript? If code is free now why would you choose a transpiled language with a huge runtime and nightmare security over something fast and lean?
- phront 2mo agohmm.. what about supply chain security?
- NitpickLawyer 2mo ago"I think it would be a good idea"
- alansaber 2mo agono thanks. i use AI.
- sora2c0314 2mo ago[flagged]
- Shorel 2mo agoAwesome, let's read what they have done! I open a new tab. To install the harness, first use npm... And tab is closed. No thanks.
- satonakamoto 2mo agohttps://github.com/bobleer/deepseek-harness-gui https://github.com/bobleer/deepseek-harness-gui Shit! They are fast!
- satonakamoto 2mo ago"The Chronicle of DeepSeek Harness Development: Sixty-Five Days from Internal Initiation to Overnight Viral Success" https://dsh-chronicle-duv8yxo8n-tsonglews-projects.vercel.app/ https://dsh-chronicle-duv8yxo8n-tsonglews-projects.vercel.ap...
- laul_pogan 2mo agoTrying it now, seems a little sloppy...
- vhantz 2mo agoThere is such a clear lack of innovation drive in this field. Every lab just copies what the other does. One of the most baffling things to me is how the once-upon-a-time good developer instinct to make everything reusable, testable, and deterministic is just getting lost into a sea of markdown begging a language model to please act a certain way. For example this repository has a "skill" definition that consists in instructing the LLM to run pre-commit checks. But we have solved this a long time ago, it's called git hooks. I do not understand why they don't simply wire those instructions as testable, reusable, deterministic code routines in the git tool call itself. It's like everybody is taking their brains out and putting it in a drawer.
- dawnerd 2mo agoGotta burn the tokens somehow! There’s a lot of these solved problems that devs have forgotten exist all in the name of using an llm for sake of using it.
- vatsachak 2mo agoYeah. LLMs have their place and they are definitely super human at short length tasks, but I feel like a large part of the "AI boom" is trying to get the computer to do something in a worse way then it already could.
- __alexs 2mo agoYes this whole focus on customisation and plugins etc is really just laziness and the absence of innovation. I don't want an infinitely programmable IDE. I already have that it's called my computer. I want something that actually has an opinion and gives me productive value without having to spend days reconfiguring it first.
- Kuyawa 2mo ago47mb downloaded, 1.5gb after build, wtf? I consider my own coding agent bloated at just 1mb (yes 1mb) because it uses postgresql package as db tool, and it works wonders. * edit 1: Upon further scrutiny, 35 dependencies make up for 1.4gb, what they are for? I don't even see postgres in there so I guess that would be another plugin. 1.5gb of basic functionality? * edit 2: Most of the time I use the terminal but also developed a web ui for my agent [1] and it is only 20mb with postgres, git, web, file tools, etc I definitely want to know why the bloat [1] https://github.com/kuyawa/mecha-ui https://github.com/kuyawa/mecha-ui
- nefarious_ends 2mo agoyeah, the build took 3 minutes on my macbook air. what's in this thing?
- Kuyawa 2mo agoFound the culprits, anthropic and openai models are 250mb each. So if we use only deepseek I guess it should be ok to remove them, but still some tests to do just in case
- redbuck 2mo agoNODE_MODULES
- Kuyawa 2mo agoExactly, node_modules, but being a single page app with nothing much other than calling AI in the background and beautifully presenting the results, what are those modules for? I suspect the culprit is react, typescript and its entourage of useless stuff not appropriate for a simple app in its structure, even if extremely complex for producing results. UI is simple, we shouldn't complicate stuff
- jimmydoe 2mo agoSpatiotemporal Composability... CORDIS... sounds like a few doctor who fans in deepseek
- nycdatasci 2mo agoEverything is a plugin. "this, like all other problems in Computer Science, can be solved by one more level of indirection." Roger Needham, circa ~1981
- addozhang 2mo agoI personally really like products with plugin systems: a stable, cohesive co with a rich, extensible ecosystem. You can create products that fiyour exact needs, and even if there's no plugin that meets your requirements, you can build it yourself. At least there's vibe coding. Just like Obsidian, there's also hot loading.
- slowin 2mo agoIs there any place to see benchmarks for harnesses (not models)? I'd love to see these things compared.
- WASDx 2mo agoI needs to be harness+model combination, https://artificialanalysis.ai/agents/coding-agents https://artificialanalysis.ai/agents/coding-agents
- lenerdenator 2mo agoAs someone who's just getting into the self-hosted game on a M2 Pro MBP with muse-glimmer 30b, what's the difference between something like this and Cline?
- anigbrowl 2mo agoThey do the same job. I can't say how well this performs; Cline works well early on in a session but I find I regularly need to start new tasks, past a certain point the context window gets cluttered and it starts trying to redo tasks it has already completed. Now, I'm sure that is partly my fault and not digging into Cline's more advanced configuration options or whatever, but it's not too obvious what to adjust. It sounds like DSH is promising better task management and easier configuration; I guess I'll find out when I try it with an existing project later today.
- Kuyawa 2mo agoI like it, it is beautiful, specially the trajectory tabs, very explicit, detailed on what it does. I like the plugin architecture, I wish they were sorted alphabetically so I don't waste hours looking for a plugin in a sea of unordered text. 9 out of 10 Edit: After creating an app it works as expected, no complains, lots to celebrate, being version 0.1 there is room for more surprises but right now it's the perfect tool for those initiating in agentic coding with one of the most affordable and powerful AI. It is really wonderful.
- sora2c0314 2mo ago[dead]
- SwellJoe 2mo ago"Every run is traceable Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream." That's a killer feature, IMHO, and one that US models won't allow you to do, as their traces are encrypted, obfuscated, etc. and have to be extracted via various workarounds (that violate the terms of service). If you want to be able to improve your tools that work with models, you have to be able to assess what the models think is happening, how they think about and interact with the data you give them. And, the US models won't let you see that.
- alansaber 2mo agoAgreed that it is a killer feature. US models obfuscate the COT (to A. make it look better and B. combat distillation) but > and the raw trace is fairly hard to reason about > but I still think this kind of feature is a big step in the right direction.
- gosolozero 2mo agoI agree it’s a great step. But the deepseek models also don’t perform to the same level of fable/sol. If we optimize/finetune to deepseek traces, wouldn’t it be suboptimal? What would the benefit be?
- Phemist 2mo agoYou let the smarter model explore the traces and figure out where the current harness' bottlenecks are for the current LLM. Then you can adjust prompts or tools to fix those.
- moljac024 2mo agoCan you explain how COT is essential to distillation and why obfuscating it helps defend against distillation? Or point me in the right direction in terms of what to read.
- salathielzhang 2mo ago[dead]
- gagan2020 2mo agoI was working on same idea but left in between and thank god they did it. Why I left that idea is because as a developer I know that was needed but I have limited time so I need to build that is really next path forward. I am working on whole dev space that can run on my Mac M4 or similar specs. I needed to revamp everything (LLM thinking) from ground up even models. My idea is mixing deterministic nature of existing tooling (non-LLM tooling) with non-deterministic nature of LLMs.
- root_axis 2mo ago> My idea is mixing deterministic nature of existing tooling (non-LLM tooling) with non-deterministic nature of LLMs. This is the fundamental idea behind every LLM harness.
- gagan2020 2mo agoLLM harness is making tools work with LLMs. My idea is making LLM work with tooling. Creating LLMs to work with tooling. By this process, LLM can be made smaller & efficient as well.
- ef2k 2mo agoWhat's buried under the lede: this harness is using Cordis v4 (the paper that dropped today). Cordis has already been used for four years in a different project called Koishi that uses v3. Cordis itself is a way of hot loading and unloading plugins without restarting a running process. The cool part is that when it unloads it can revert any state and side effects it created, cleaning up its connections, memory allocations, registered handlers, etc. and it can also deactivate any dependencies it relied on without disturbing other plugins.
- ziofill 2mo agoSounds very cool! What do you mean by “revert side effects it created”?
- ef2k 2mo agoAnythign that needs to be cleaned up or undone goes in ctx.effect. It returns the "inverse" (the cleanup function) when the plugin loads. Cordis then stores it and runs it when the plugin unloads. Take a look at 5.1.1 in the paper.
- deleted 2mo ago[deleted]
- shostack 2mo agoCould memories and blocks of context be made pluggable and unpluggable in this manner to "hot swap" context?
- roywiggins 2mo agoNot without nuking your cache, I figure.
- JHonaker 2mo agoAs far as I can find, Cordis didn't exist before they shared it with this harness. It is very likely that its because I didn't search in Chinese (as I don't know it). Can you explain the evolution of it from the different versions?
- alansaber 2mo agoCode mode getting some love.
- fkysly 2mo ago[dead]
- KronisLV 2mo agoEverything is a plugin? I'm reminded of Eclipse!
- notjes 2mo agoWe can /model, but when can we /harness?
- try-working 2mo agoI've been thinking that they should design the DeepSeek harness to work with other providers since a large use case is to off-load work from an expensive model to DeepSeek, instead of makign everyone hack the harness. I see that it works with many different providers out of the box and that's a great thing. It also makes it easy for me to build a plugin for the role-model router and have it work properly, so you can route between models automatically. Will be out later today.
- z_rho_one 2mo agoBased on what I've read today, DeepSeek Harness seems to be similar to Pi Coding Agent in design. Both are barebones to start out and rely heavily on plugins. However, 3 things make DSH stand out. 1. Plugins in DSH are required to have cleanup handlers, so I guess you could clean up plugins that are no longer used mid-session and prevent it from interfering with the current task? (unsure about this) 2. DeepSeek V4 models are post-trained on DSH. Given how cheap DS V4 is compared to OpenAI and Anthropic models, running DS V4 in DSH could be much more cost-effective while barely losing performance. 3. It's utilized and maintained by a large lab dedicated to open source AI. It's always nice to get new open source tools from large labs so that we are not always relying on small teams doing the heavy lifting.
- favflam 2mo agoIs this better than the bytedance cloudwego/eino Go library?
- mancerayder 2mo agoWhat sorts of questions can we and can't we ask?
- williamjinq 2mo ago[flagged]
- prtmnth 2mo agoMissed opportunity to call it DeepCode
- tw1984 2mo agonot very impressed. in the era of AI, telling me that the core design is a plugin system that can be reloaded and extended easily is just not exciting. it is something you feel excited 20 years ago back in the 2000s, in 2026, the expectation is agentic capabilities and self improving.
- throwaway7783 2mo agoWe keep reinventing everything. Eclipse style plugins from a long time back
- Founderarcstone 2mo agonice very good to see this.
- aussieguy1234 2mo agoMy approach to harnesses is: Everything is a skill, backed by a CLI tool that both I and the agent can use and debug.
- hahahaa 2mo agoProblem is I use pi.dev and I don't want to learn a new thing the new things keeping coming. Too much new and unprecidented rate.
- doumiBou 2mo ago[dead]
- sha2kyou 2mo ago[dead]
- skc 2mo agoPlayed around with it. Does what it says on the tin. Great work.
- dfguo 2mo ago[dead]
- aitobox 2mo ago[dead]
- taylorbuley 2mo agoSteal their testing substrate. The offline evaluation is genuinely fucking clever.
- dbbk 2mo agoI don't know how much more of these slop marketing pages I can take. "Everything is a plugin" as the leading headline - what does that even mean? Who cares?