4 ms·
Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
- dunlin 14d agoBeen hoping for something in this space. Jev-like decision models on Qwen3.5 could really simplify some of our internal routing logic.
- mugul 14d agoQuite impressed by the energy people are putting into making OSS Jev-like models. I understand the hype but I wonder: what are the use cases for this kind of model? Could it be used in the context of coding agents, or is it more relevant in totally different situations?
- Havoc 14d agoYeah same. Got access to their API and then realised I don’t really have an immediate use case
- lucrbvi 14d agoYou should call Jev-like models when you give it a JSON-like structure to produce, it is useful when you need _some_ intelligence in your code. Edit: I want to add that you can see Jev like a smart if-statement.
- saejox 14d agoTo develop a smart ai system for my 2d roguelike platformer? game has way too many moving system for classic state-machine ai + i cant spare the time to develop it. its low latency entices me.
- nikolovv 14d ago[dead]
- vidarh 14d agoConsider every situation where you "force" an LLM to output only a choice / category, or a set of them. If you have workflows like that, you're now being promised significant cost- and latency reduction. For coding agents it'd only be useful in a subset of situations. E.g. you could imagine using one to classify bash tool calls into safe and unsafe for example.
- NitpickLawyer 14d ago> what are the use cases for this kind of model? Could it be used in the context of coding agents Yeah, it could. The most obvious usage would be to have local fast cheap "feedback" / "control" over a slower more expensive agent (i.e. cc / codex / opencode). Things like "goals" could now be split from a long prompt into "actions" and "verifiers". Where for each action you also produce a verifier. Then after each action you run the verifier w/ this kind of "universal classifier" and decide if the step was done correctly, if it needs follow-up and so on. Example: implement auth in this repo -> llm_plan() -> for item in plan generate_verifier() -> for item in plan implement() ; verify() ; accept() / followup(). Verifiers could be something like this. take a plan item as input, generate classification questions that might verify the task "is this following project conventions?" | "is this touching files from other tasks?", etc. You can do that with LLMs, but some things might become cheaper / faster. And you can pretty much use it to check against an ever growing list of conventions. Yours or project specific.
- jeeeb 14d agoI don’t think this is a very good use case. You could do it better with a strong LLM and structured outputs. The problem is that you want the model to carefully reason about the goal and code. Zero shot classification with an approach like this isn’t going to do that. It’ll answer on first pass vibes.
- altmanaltman 14d agoWhy not just let the LLM write a test instead of a "verifier"?
- Schlagbohrer 14d agoThe git repo linked at top has some good examples for email classification for automatic email forwarding to specific departments, along with judging email tone and severity / priority, like for customer service emails. Edit: The Flipper One is planning to have an LLM acceleration co-processor, and be able to host up to a 4GB VRAM size LLM. One use case they envision in their planning is using the microphone along with text to speech to be able to say, "Create an .ini file for this system with these specs" and the small LLM can do that on-device (its a handheld device) and then the user can use/send/upload that file. Second Edit: I would love a mini LLM in KiCad or Altium that could take a component datasheet and produce a good footprint and schematic symbol for it.
- yogthos 14d agoClassifiers can be very useful for guiding the agentic loop which is basically a state machine. You have the agent propose a task, write some tests, write code, run tests, tests fail, write more code, go to acceptance, etc. So, a classifier can judge state transitions and decide what the agent should do next for example.
- monkeydust 14d agoBit of a Jev explosion going on. Is it because it's taking us back to a simpler time we understand better? Classification models have been around for a while.
- toasty228 14d ago[flagged]
- mugul 14d agoThanks to these projects, what was an innovative-but-closed piece of technology one week ago is now much more accessible. Whether they're in it for fame or not, I couldn't care less!
- toasty228 14d agoThere already was an alternative a year ago, with a published paper and open weight lmao... all the other projects are literal slop shat out by script kiddies 2 hours after the release of jev, it reminds me of the flappy bird era, depressing
- Tycho 14d agoIt’s because it’s practically useful and enabled things that were impractical previously.
- petesergeant 14d ago> and enabled things that were impractical previously I think that there are not _that_ many use-cases that have been opened up by this that tool-calling on other models didn't solve already. Really depends what benchmark you're looking at. This one against BANKING77[0] has many issues, but suggests it's really not far off DeepSeek 4.1 Flash. This one against BoolQ[1] shows marginal improvement over Qwen3.6. This one against MMLU-Pro[2] (same author as the previous) shows significant improvements over two Qwen models. So there's definitely _some_ alpha there, but I don't think it's the sea-change that the hype would suggest; that is to say, yes, some things that weren't practical before are now, but many things were already very practical with the existing tools. 0: https://sanand0.github.io/llmevals/jev/ https://sanand0.github.io/llmevals/jev/ 1: https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26d5d7b3b2184a524fc61779dbcc1/evals/results/boolq-2026-09-18/comparison/report.md https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26... 2: https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26d5d7b3b2184a524fc61779dbcc1/evals/results/qwen38-27b/report.md https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26...
- webprofusion 14d ago- https://github.com/logan-markewich/jeff https://github.com/logan-markewich/jeff - https://github.com/bespokelabsai/nimble https://github.com/bespokelabsai/nimble
- webprofusion 14d agoWhy does nobody ever ship these as a docker image?
- tacomagick 14d agoI guess you have AI to write your docker files and push your images now.
- Schlagbohrer 14d agoAre these forks? Different orgs doing the same thing as the OP?
- nullbio 14d agoI think a great use case for these will be when they have large context windows and are able to enforce styling rules for frontend development, and component creation rules for react. You can then ditch the styles guides and styling skills and create a decision tree for enforcing styling, so that you can't run into drift issues or duplication issues. That's where I'm wasting most of my time right now, constantly correcting all of the UX/UI issues that are created for every single feature.
- spockz 14d agoBack in the good old days we would prevent these ux/ui issues by rigorously enforcing the use of our own stylesheets and classes. Later that grew to only using the company ux components. This was very successful in keeping everything neat and tidy. The only drawback was creating and curating new elements and getting consensus. But otherwise it works wonders. Try constructing reusable components out of what you are doing instead of building everything up from basic building blocks. This also allows more concrete testing of individual parts and then if you want to change the look you can change it in one place and have it apply everywhere. Agentic development doesn’t mean “throw all what we learned out of the window”, the same practices that helped speed up and improve quality of work of humans also helps agents. In fact, the multiplier is even bigger. You will notice it in development speed and reduced cost due to avoiding churn.
- nullbio 14d agoI use reusable components but the issue is that the model rarely checks to see if a component already exists. Or it will use the wrong one. Or if the component strays from the reusable component it doesn't extend it in a generalized way and will special case inline styles, or it'll create a new component. Rarely does it intelligently figure out the correct course of action. I've also built up a suite of linter rules to catch the same mistakes the model makes over and over. Still, there are a lot of gaps. I think it's mostly because my codebase is massive at this point. It was easy when the codebase was small and didn't require context gathering to make good decisions.
- 14d ago
- raahelb 14d agoBecause these decision models do not have tool calling, the knowledge cutoff might become a problem. We'll either have to keep training continuously if we run locally or switch to the newer version every month or so when using a closed one like Jev
- giuscri 14d agoeven with knowledge cutoff set a second from now, you still want to provide as much info as you can if you’re using such tools for delegating decisions
- cedws 14d agoI noticed it has a pretty small context window of only 32k. For most tasks I guess it would be enough with ample context.
- jwr 14d agoI wonder how these would do filtering my spam. I have been using 27B-class models for a while now, and they are nearly perfect at determining what is spam and what isn't. The only disadvantage is computational cost.
- walrus01 14d agoTake a look at Thomson 1.0-small, which is a variant of qwen 3.6 35b post trained by Thomson Reuters for text analysis. It classifies text content very well.
- Mumps 14d agoAre you on the foundation research team for Thomson? (If so, hiya from B!) Why would you expect Thomson to be particularly good at spam clf? I figured your additional corpus was all news and legal?
- walrus01 13d agoI have no connection with Thomson Reuters other than as an end user of a GGUF of the LLM I mentioned. That said, from my personal experience with this specific LLM, it's a decent improvement over a "base" Qwen 3.6 35B A3B Q8, and it does a good job of analyzing and categorizing documents on relatively small resources. It'll run fine in llama-server in pure CPU only on a 64GB RAM system with plenty of room to spare, takes something like 47GB with RAM reserved in llama-server for cache and full context size.
- faangguyindia 14d agoOn Gemma 4 12B, I am getting 220 ms per move or QS. I used it to play the Snake game locally: prompt_eval=244 ms wall=245 ms schema_cache=hit generated=0 Move limit reached after 200 moves: score=16, length=19. So, if a 12B dense model can offer this latency on a local old PC, then definitely you can scale it up with more powerful machines and get even lower latency.
- akkad33 14d agoCan someone tell me what is the difference between Jev and a normal neural network that does classification ? My understanding is: it takes text input and it does one shot classification (no training data)
- crackalamoo 14d agoYes, this is essentially it. As a corollary, the output classes can be any set, rather than needing to be set before training.
- akkad33 14d agoCan someone do a ELI5A of how they achieve classification over any user defined list of items? Normal neural networks do a softmax over a known output set to get probabilities
- theodoretliu 14d agoI can think of two possible approaches 1. Jev limits to 255 distinct options. So they can preprocess your set of options and “tell” the LLM via input tokens 1 = red, 2 = blue, etc then jev need only output softmax over 255 states while benefiting from pretrain of other LLMs 2. You allow the forward pass to output over the total token state but mask over the logits to limit to the user options. Less plausible? bc tricky when input is multi token which they clearly support. My guess would be option 1. Didn’t read the kev repo here which would also explain
- andy12_ 14d agoYou can achieve open-vocabulary classification by making the final weights in the softmax come from a category encoder instead of being fixed learned weights. So instead of softmax(encode(input)*learned_weights) You have softmax(encode(input)*encode(categories)) I'm not sure if Jev does it this way, but it's how you get open-vocabulary zero-shot image classification with models like CLIP [1]. [1] https://openai.com/index/clip/ https://openai.com/index/clip/
- Eastmill 14d agoInteresting approach with Qwen3.5 for decision models. Curious how "tiny" they've made them while keeping LLM reliability for critical paths.
- rkeswick 14d agoInteresting to see a Jev-like approach applied to Qwen3.5. Always appreciated Jev's simplicity for quick decisions.
- hackernud3s 14d agoAlways?
- raahelb 14d agoThe bright side of Jev being so popular could be that many companies and individuals realize that their applications might work well with a System One model, and they decide to run an open-source (or fine-tuned) version on their own
- hn1rig3rak 14d ago[dead]
- monxer 14d agoWhy not name it Qev?
- hbarka 14d agoIf Jev is fundamentally trained using RLCD while you’re building on a Qwen model that was trained using RLHF, how can the resulting model be considered Jev-like?
- tietjens 14d agoAlso my question.
- mohsen1 14d agoI can't find it but saw that if you give Jev English alphabet as choices and ask it in a loop what model it is, it would say Qwen also tried myself: https://console.typesafe.ai/playground?share=shr_1690a3160f19c2e4851b6a2883bcf17a790 https://console.typesafe.ai/playground?share=shr_1690a3160f1...
- rrr_oh_man 14d agolol, this is hilarious
- fxwin 14d agofor some reason this is really funny to me. it's like the "black museum" black mirror episode where a consciousness in a toy animal can only communicate using very primitive predefined responses
- prodigycorp 14d agoPeople seem to turn their brain off when it comes to this type of cargo culting. This doesn’t mean much. Qwen often identifies itself as Claude. Does that make it Claude?
- dr_dshiv 14d agoSort of implies its basis, doesn’t it?
- prodigycorp 14d ago
- ingen0s 14d agoOh Jared is cool - he made After and Razzle - nice
- andy12_ 14d agoAll the people that are just writing an Jev-like API on top of a normal LLM are missing the point. What makes Jev special is the training data; it's how it's trained. The architecture is probably nothing special. Just a text encoder with parallel prediction branches. I have tried many of these open-source Jev-like models on some linguistic tasks and they are so bad compared to Jev.
- Tostino 14d agoIt won't be long until people produce a decent training data set generation pipeline. The number of people working on this is crazy. Something will coalesce.
- andy12_ 14d agoI hope so. And I would really like to try an actual Jev open source model. But it will make it more difficult to market it when someone releases something like that because of so many of these "open source Jev-like model".
- Tostino 14d agoI'm just sitting back for a few weeks / a couple months to let it shake out, let others put in all the work, and then see if people are still interested and finding use cases that this access model fits better than the usual chat completions endpoint people are used to.
- Schlagbohrer 14d agoWe are indeed in an era of riches (thanks to the $1 Trillion a year being spent on this tech) that it is improving so fast I can just sit back with my 3 year old hardware and newer, better, more amazing workflows keep becoming possible for me just due to model compression / optimization + workflow developments. If I can't get something working this season, I just wait until 3 months from now and there will be an easier to set up, less resource demanding, better working version I can have my local AI install for me. Pretty wild times.
- sinan-faizal 14d agowhat kinda of specs would it need to run?
- stackzero 14d agolooks high lev
- scotty79 14d agoDistilling Jev should be super easy and cheap.
- aetherspawn 14d agoI hope these get small and good enough to create “pet like” AIs for games. You know, like scream “follow me” at an NPC, STT stack translates it and feeds it to a local Jev-like model that then picks a number of things for the NPC to do.
- ranyume 14d agoI tried to use jev for this. I'll share what I learned for the interested. -- The setup was a simple map with different rooms. Each room had 1-3 doors. For the inputs: The AI had an array of "known places" empty at the start, the current position, the current doors with no information about where they lead to, and the list of past actions The goal / task for the AI was to explore all rooms and save them to known places. The AI needed to decide if to move or save the place at every turn. -- So I wasn't able to make the AI explore all of the rooms. The AI kinda always wanted to move to the first option when moving. Out of 6 rooms it was able to save 3. My hypothesis is that jev as it is now is really bad at making connections and understanding it's input. So for example, even if it had a list of previous actions, it wasn't able to reason about it and know where to go. For this to work I'd need to explicitly tell it where it did not go. So you could say that the model is also not good with uncertainty / ambiguous scenarios. edit: one last thing is that i replaced jev with an standard llm and it finished the goal no problem with the same information given edit 2: it also felt like the same tradeoffs between small model vs large model. With small models you need to be very conscious and careful with the input while large models are more forgiving. Maybe jev is a small model, and we just need a larger one.
- oscarfr 14d agoFound this benchmark for Jev-class models: https://benchmarkheaven.com/jev-models https://benchmarkheaven.com/jev-models There are already many Jev-like models in there. Edit: No affiliation. Just found it and thought others might find it interesting.
- jasonjmcghee 14d agoThe open source ones- I downloaded a number and tried them and compared to Jev. Anything that required knowledge / familiarity mmBERT and ModernBERT post-trains performed much worse. So it seems like they did some kind of useful expansive pre-training. Things that were Qwen or Gemma Diffusion did better at those kinds of tasks but were generally pretty inconsistent in terms of whether they could succeed repeatedly (and be stable + reliable) on the many types of tasks that are in the cookbook part of the Jev docs. If you ask Jev similar input + questions, it's pretty stable. And does a reasonable job on a lot of questions. This one public benchmark (the only I've seen) seems to give the open versions way too much credit. It wasn't my experience at all. It gave my a false wrong sense of what might be required to get it working for something at work to avoid needing a new subprocessor as - at least on Cloudflare / OpenRouter Jev is third-party not hosted.
- oscarfr 14d agoThanks for sharing your findings! We are working on running our own benchmark of Jev and some of the other models. Our use case is classification that runs in a UI. Currently LLMs have good accuracy, but are too slow (and expensive). Jev not being available through a cloud provider (Bedrock or similar) makes it more challenging for us to start testing and rolling it out.
- bglusman 14d agoI think I found Jev is now available through OpenRouter? Oh but maybe you mean somethign like Bedrock with zero-data-retention offered? Anyway, it was on OpenRouter the other day, lest anyone be confused by above, as I did some testing with it for my pareto frontier tool https://github.com/bglusman/model_skyline https://github.com/bglusman/model_skyline (which, apologies, is full of slop because its 100% AI maintained but, may or may not have some utility for guiding automated or manual decisions... it definitely showed that Jev was performing MUCH better than some of the open alternatives to it we tested anyway)
- merqurio 14d agoShouldn't Jev-like models be built on top of diffusion models ? like GSAI-ML/iLLaDA-8B-Instruct ? That showed us the best results at least
- epsilonic 14d agoHow is calibration of Jev or Jev-inspired models being evaluated?
- k__ 14d agoHalf-OT: Is Jev a decoder (e.g., BERT) or is it some kind of encoder (e.g., GPT) that just happens to be trimmed down to only outputting a handful of tokens for the answers and their probability?
- Alpha3031 14d agoYou've got encoder and decoder reversed. BERTs and other models that primarily convert text (or other input) into latent representations are encoders. Models that convert their internal representations back into outputs are the decoders (in the case of GPT et al., autoregressive decoders because they perform this decoding based on past tokens).
- npn 14d agoYou got it reversed. Bert is encoder only and gpt is decoder only.
- raybb 14d agoHow long does it typically take for something like this to become available on openrouter?
- yunusabd 14d agoHaven't tried it yet, but it looks like it's already live? https://openrouter.ai/typesafe/jev-1.13 https://openrouter.ai/typesafe/jev-1.13
- prometheus1992 14d agoHow are you forcing Qwen to answer in a structured way? I like this one better - https://github.com/deepanwadhwa/OpenDecision https://github.com/deepanwadhwa/OpenDecision
- prodigycorp 14d agoMan, I'm already burnt out on all this jev talk. The one thing jev has going for it is a dedicated company focused entirely on making the product good and keeping it maintained. I haven't been willing to jump on board with all these jev-shaped projects because their releases feel driven mostly by opportunism. I'm fine waiting a bit for the opportunists to shake out so we can see who is genuinely committed to bringing something valuable to the open-weight community. Jev is much better than the traditional ML crowd gives it credit for, but my enthusiasm hits a wall when it comes to their data policy. It is completely draconian. Whatever you feed into the system, they retain. The jev team needs to release a ZDR product, or their platform is dead on arrival. An open, jev-shaped model will win out solely on that basis.
- kerwioru9238492 14d agohttps://typesafe.ai/legal/privacy-policy https://typesafe.ai/legal/privacy-policy In their privacy policy they say We (1) will not train or fine tune any artificial intelligence or machine learning models on Input, and (2) will not disclose any Input to a third party other than our service providers.
- prodigycorp 14d agoYeah but they reserve the right to retain the data virtually indefinitely. These aren’t acceptable terms on a personal or corporate level. I’ve seen some fools brag about proxying their life through jev. Messages, emails, LLM calls, files.
- zambal 13d agoAccording to https://docs.typesafe.ai/legal https://docs.typesafe.ai/legal they do offer ZDR for enterprise customers.
- prodigycorp 13d agoI missed that. Typesafe, why is this hidden in your docs?!
- BeetleB 14d agoOK - I guess I'll ask here. As there have been a lot of Jev related submissions, can someone point me to a simple guide on how I can use it? For example, say I have a script/workflow where I use OpenRouter for LLM calls, and at some point I want to do a simple classification. Can I still use OpenRouter with some Jev model...?
- prodigycorp 14d agoBest just to use a skill. Here is typesafe’s skills.md https://docs.typesafe.ai/agent-skill https://docs.typesafe.ai/agent-skill (which I don’t think is well written, but it’s a start). Also read their docs, I think they’re interesting.
- BeetleB 13d agoI need it for running in standalone scripts. How do I get a skill to work there? I mean, I guess I can have my script call pi and offload it to that, but I just want everything contained in one script.
- algoth1 13d agoJev feels more and more like a glorified if/else if block
- nico 13d agoIf you only need classification, and you can provide some training data, you can ask Codex/Claude to build an embeddings + logistic classifier model for you For emails, I get 95% accuracy with this method, with only 50-100 examples for training Training the model takes less than 5 minutes on a CPU The resulting model is <1MB, and inference is sub 100ms Some other cool things about this approach: * the model doesn’t train on some “ideal” or general classification, instead it learns your preferences * the model runs on pretty much any mobile device and can be retrained online on the device * privacy, the whole training and inference is 100% local, no data goes anywhere (except whatever you feed codex/claude while building the model) Note: to do a more general test, I made a classifier for the Banking77 dataset. The model is <10MB, trains in <30s on CPU and gets 94.5% accuracy, which puts it in the top 5?models by accuracy for that set (the best one is at 94.86%, but it’s 350MB in size and takes hours to train on a GPU).
- rgbrgb 13d agolove this idea. did you try comparing to jev?
- nico 13d agoYes, I ran some benchmarks. This architecture seems to match or beat Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one) For the latter cases, you could probably enhance the architecture with a lightweight LLM, something like a Gemma model. Or even some basic MLP
- samuel 13d agoDo you realize people is using LLM's as classifiers, right? For lots of companies and developers reaching an API is feasible, while running a training pipeline, no matter how simple, is not. I know that they should still be gathering data for evaluation and they could use it to train a model instead. But they won't do it, for a variety of reasons. This is the same route but WAAAY faster and cheaper. And you can modify it like you do with code or prompts. It's really appealing, TBH.
- loclol101 13d agoUsing jev for data labeling would be interesting. I wonder how kev compares
- floatrock 13d agoApply it to this: https://minimallysufficient.com/posts/llm-classification-is-feature-extraction/ https://minimallysufficient.com/posts/llm-classification-is-...
- khazhoux 13d agoAll these Jev projects… great. But Jev was only just released a week ago. That’s the hard limit on how much effort has gone into all these OSS extensions and derivatives: one week. I don’t therefore see any value in adopting any of them, versus just vibe-coding my own if needed.