6 ms·
Ornith-1.5: From Self-Scaffolding to Self-Improvement
- bigcat12345678 1mo agoHow is ornith-1.5's base model developed? Is the base model one of the Open weights models, or one pre trained by ornith team from scratch? I couldn't find information to answer this question in the article.
- goldemerald 1mo agoIt looks like they post-trained Qwen3.6. Interesting to see how far they could improve it with they harness/algorithm.
- deleted 1mo ago[deleted]
- tangjurine 1mo agoThis looks cool
- prometheus1992 1mo agoCan't wait to try this. Ornith1 (9B) was a really nice model. I have been running it locally using - https://github.com/deepanwadhwa/samosa-chat https://github.com/deepanwadhwa/samosa-chat
- _def 1mo agoWhat did you use it for? I've found it somewhat capable but not worth to actually use it (1 9B, that is)
- jonesy827 1mo agoI've been using the 35B-A3B today for some web scraping work, and it has been on par with Qwen3.8 27B at a much higher speed and at a higher quant (q4 vs q8). I'm impressed.
- jadbox 1mo agoI need someone to run actual benchmarks between the two.
- swatcoder 1mo agoBenchmarks are the BMI of model evaluation. They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.
- dofm 1mo ago> Benchmarks are the BMI of model evaluation. That is such an elegant way to put it.
- NitpickLawyer 1mo agoOnly relevant benchmarks are those you make yourself, targeted specifically for your workflows. Anything else is just number go up on a pretty graph, and every model out there is probably benchmaxxed to hell on the public ones anyway. Keep yours private.
- gertlabs 1mo agoThese models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.
- RomulusHill 1mo agoHi Gertlabs, my inference company has actually started offering Ornith1.5 today for 9B and 35B A3B. If you are still interested, send me a message on X and I'll help you get started! https://x.com/romulushill https://x.com/romulushill https://scalattice.com/models/ornith-1.5-9b/ https://scalattice.com/models/ornith-1.5-9b/ https://scalattice.com/models/ornith-1.5-35b-a3b/ https://scalattice.com/models/ornith-1.5-35b-a3b/ https://scalattice.com/developers/ https://scalattice.com/developers/
- nextaccountic 1mo agoIs this open weights? Or planned to be
- randomblock1 1mo agoYes: https://huggingface.co/collections/ornith-ai/ornith-15 https://huggingface.co/collections/ornith-ai/ornith-15
- montroser 1mo agoHoping this is real. It's too bad to see the signals from Qwen that they will not be releasing a 35B-A3B for the 3.8 lineup. The MoE architecture makes a huge difference for being able to run these local models on reasonable consumer hardware.
- parsimo2010 1mo agoHonest question/suggestion for the HN audience- Since Qwen released the weights for Qwen3.8 2.4T-A95B and we already have the staring point of Qwen3.6 35B-A3B, couldn't someone distill the bigger model and make a "pseudo" Qwen3.8 35B-A3B? Sure, it wouldn't be an official Qwen release but couldn't someone improve on Qwen 3.6 and get the thing everyone is asking for? I am calling this a suggestion for the audience because I don't have the will/resources to do this.
- WASDx 1mo ago"Qwen3.8 35B-A3B" and 4B/9B variants are already on huggingface distilled by hobbyists.
- karlkloss 1mo ago>I can't find any. Do you have a link?
- KronisLV 1mo agohttps://huggingface.co/Lord-H4D3ZS/Qwen3.8-Distill-35B-A3B-Coder-Abliterated https://huggingface.co/Lord-H4D3ZS/Qwen3.8-Distill-35B-A3B-C... > Base / architecture: Qwen/Qwen3.6-35B-A3B (Qwen3_5MoeForCausalLM, 256 experts, ~3B active). The "3.8" in the name refers to the teacher, not the base. Not endorsement, haven't run it myself, just found the link.
- parsimo2010 1mo ago[dead]
- jakswa 1mo agoI'll be comparing the 9B vs Ling 3 Tiny (8B-A1B) as a scout model. Ling tiny is so fast but can be a little too dumb. Hope the 9B strikes a good middleground even if dense/slower.
- jakswa 1mo agoended up disabling ornith 9B. Oddly Ling 3 Tiny is pretty dang capable if its thinking is unleashed (tons of output tokens, maybe 5X the tokens but it's so fast it's maybe only twice as slow as a smarter model). This is a really interesting space if these super small models keep improving. SO FAST! Want to use!
- garo-pro 1mo agoAcross five cases it reliably claims to be Claude without being able to name a specific version.
- wgd 1mo agoNobody (with the probable exception of Anthropic given their work on character training) really trains models on their identity and Claude is the only AI persona that's well-defined so if you put yourself into the AI's shoes it's a pretty reasonable guess that it might be Claude. I've had basically every open model claim it's Claude when the topic comes up.
- orbital-decay 1mo agoAnthropic changed their character training policy many times, and the name is probably separate from it anyway (besides the bits from the constitution etc). Name training usually comes last in all models, if at all, and it's pretty shallow. Certain Claude models say they are Qwen or Deepseek when asked in Chinese, for example.
- wgd 1mo agoYeah, Claude is actually surprisingly unsure of his identity considering that their most recent publication on their constitutional AI training literally had graphs demonstrating how certain properties differed based on whether they were phrased as questions about "Claude" versus "You", but it's actually that which makes me fairly confident that they're probably doing _something_ to try and close that gap. More generally, their current approach to constitutional AI pretty much only makes sense if they believe that they can first teach the model what the Claude character is like and also teach the model that the persona responding is Claude, so I figure that has to be part of the pipeline even if they're not very good at it.
- colingauvin 1mo ago397 is just too big for two Sparks even at NVFP4. Wish they had made this just a tiny bit smaller.
- kees99 1mo agoOrnith-1.5-397B is derived from Qwen3.5-397B-A17B via post-training. That process preserves exact parameter count.
- colingauvin 1mo agoAh thanks for the background. Well, time to learn how to quant things down!
- ggcr 1mo agoThey should've included Qwen3.5-397B-A17B in the benches then :/
- wyre 1mo agoThis is exciting. Their 9B model benchmarks competitively with Sonnet 4 which is pretty cool to have such a small model compared to one that came out 10 months ago. I’m curious how providers will price their 397B model.
- esafak 1mo agoApparently this is Jiwei Li's new company. https://ai.miraheze.org/wiki/Ornith https://ai.miraheze.org/wiki/Ornith https://www.innovatorsunder35.com/the-list/jiwei-li-2020/ https://www.innovatorsunder35.com/the-list/jiwei-li-2020/ I wonder what their angle is going to be; the scene is crowded, and they don't do serving.
- htrp 1mo agoAnother day another startup claiming some vague version of RSI to try to close their round.
- lsb 1mo agoThe page has comparisons with Qwen 3.6 27b and I’d love to see comparisons with Qwen 3.8 27b, the newer one is much more capable!
- ricardobeat 1mo agoIt's somewhat close, but a lot worse at code it seems. Qwen 3.8 is a wild improvement over 3.6. Qwen 3.8-27B Ornith-1.5-35B Terminal-Bench 2.1 73.0 67.8 SWE-bench Pro 61.7 59.6 DeepSWE (1.1) 42.2 22.0 NL2Repo 42.3 46.2 GPQA Diamond 89.2 89.2 Humanity's Last Exam 30.8 25.6
- hxii 1mo agoInterestingly, in my own benchmark and testing (in the hopes of finding a good-enough local model to run a personal assistant agent), Ornith-1.0-9B was worse than Qwen3.5-9B which according to their scores should've been reversed. I will definitely pass Ornith-1.5-9B through the gauntlet as well!
- rbanffy 1mo agoIt’s time for me to upgrade the main server in my home lab and I’m thinking about which machine should I have. What kind of hardware you’d need to run the 397B one at an acceptable speed?
- colingauvin 1mo ago3 DGX Sparks with 400 Gb interconnects in a loop.
- rbanffy 1mo agoI was looking for something more general-purpose, such as a decommissioned quad-socket x86 with 80 or more AVX512 cores and a terabyte of RAM, which I suspect will be a strong limiting factor. The possibility of adding some older Nvidia accelerators is interesting as well even though they won’t be able to hold the whole model in HBM and will need to stream it over PCIe 4, which sucks. That way, when not running inference, the machine can easily host VMs for other experiments and general housekeeping functions.
- AIorNot 1mo agoCan someone parse all that AI generated blather in the post and tell me clearly: 1. Is this self improvement at the model level (updated weights or memory, KV etc) or just by adding agentic code harnesess to guide the output better? Thank you
- kzrdude 1mo agoThe former
- readyblue 1mo agoPlease check your download links, it get 400 Bad Request on https://huggingface.co/ornith-ai/Ornith-1.5-397B-GGUF/tree/main https://huggingface.co/ornith-ai/Ornith-1.5-397B-GGUF/tree/m... only mmproj can be downloaded right now
- deadfish 1mo ago[flagged]
- adagonese 1mo agoWhat fraction of GPU-hours in each cycle went to rollout generation versus the actual update?
- brittanyseales 1mo ago[dead]